Multi-Source Data Integration Platform: The Context Layer Behind Governed AI
CTO & Co-Founder
CTO & Co-Founder at Nirmitee.io. Architects healthcare integrations across FHIR, SMART on FHIR, ABDM and NHCX, writing from production experience taking hospital software from sandbox to go-live.
At a glance
- Client: an Indian vaccine manufacturer, business development and competitive intelligence team.
- Scope: a multi-source data integration platform for market, pricing, procurement, regulatory, country and funding intelligence, with two AI assistants on top.
- Platform: one custom Frappe application with 25 modules, 210 data types, 28 workspaces and 18 approval workflows, with Power BI reports on the approved tables.
- Integrations: 54 live: 23 public datasets from 8 organisations, 18 funding agencies plus grants.gov, 4 news and document feeds, 8 platform services.
- AI: a research assistant that cites its sources, and a funding assistant that answers numeric questions by querying the database.
- Status: development from July 2024, in production since 2025, maintained by Nirmitee.
Summary
Nirmitee built and maintains a multi-source data integration platform that gives an Indian vaccine manufacturer one governed view of its global market. The platform collects prices, contract awards, shipments, product approvals, clinical trials, country statistics and funding calls from 54 live integrations. It reconciles every country, product and manufacturer name to one spelling, and holds commercial tables for human approval before they reach reporting.
The result is a context layer the business can rely on. Analysts start from current, reconciled data. Leadership reviews the market in role-based Power BI reports on approved tables. Two AI assistants answer from the same context layer, and every research answer names its source.
The business challenge
The manufacturer's strategy rested on public data that was scattered, inconsistent and slow to assemble. The business development team tracked global vaccine prices, tender awards, prequalification decisions, country demand and funding calls by hand. That picture fed a vaccine portfolio analysis and a registration and market database, and through them, decisions on capacity planning, pipeline portfolio review and investment strategy.
UNICEF publishes prices and awards as monthly PDFs. WHO publishes Excel files and web tables. Gavi, PAHO, the World Bank and India's drug regulator each use their own formats and spellings. The requirement specification noted that the data arrived as Excel, PDF, text and web pages, all of which had to be cleaned before any comparison. In May 2024 the team asked for a competitive intelligence database that would keep itself current, and described the project as being on an expedited path.
What was at stake
Decisions of this kind are expensive to get wrong. A capacity or pipeline decision made on an outdated price list, a missed prequalification notice or a misread tender award can commit a manufacturer to the wrong product, market or timeline. When one country appears under three spellings, totals are wrong before anyone begins the analysis.
There was also a capacity cost. Skilled analysts spent their time downloading and reconciling files rather than interpreting them, and every refresh repeated the work.
Trust was the third concern. The team asked that any automatic update be validated and released by an administrator "until we develop confidence on the database". They did not want a system that quietly changed the numbers their decisions rested on. That requirement shaped the design more than any other.
The approach
Nirmitee designed the platform around four principles, each taken from the client's brief.
- One source of truth. Every source lands in one database, moves through the same staged pipeline, and resolves to one master list of countries, products and manufacturers. Reports, alerts and assistants all read from that single store.
- Human judgement where it matters. Commercial tables wait for an administrator's approval. Official reference statistics load directly. The line between the two was drawn with the client and has moved as confidence grew.
- AI that shows its evidence. The assistants answer from the platform's own data and a curated document library. Research answers cite their sources, and figures come from database queries rather than from generated text.
- Built to survive changing sources. Public websites change without notice. Each collector uses the simplest reliable route, every run is recorded, and failures reach an administrator the same day, the discipline Nirmitee applies in integration monitoring and support.
It draws on Nirmitee's healthcare and life sciences data engineering practice, applied here to public market data rather than clinical records, and it is kept current under Nirmitee's software maintenance and support.
The context layer: why the AI is only as good as the data behind it
AI is useful only when it has the right context. A general-purpose model asked about vaccine prices or a country's move off donor support answers from its training data. The answer may sound confident, but it is not current, not specific to the manufacturer's products and not traceable to a source. That is not good enough for a capacity or investment decision.
The context layer is what closes that gap. In this platform it has three parts:
- Approved, reconciled data. Structured tables for prices, awards, shipments, approvals, coverage, schedules, trials, country statistics and funding calls, with every name mapped to one master spelling and commercial tables released by an administrator.
- A curated document library. WHO SAGE papers, Gavi strategy, market shaping and board documents, UNICEF supplier meeting documents, annual reports found on approved domains, and the team's own reports and meeting notes.
- An allow-list of trusted websites. Eight global health domains whose citations may appear in an answer.
It is built in seven steps. Collectors fetch each source on its own schedule. Cleaning turns PDFs, spreadsheets and web tables into consistent records. Name matching resolves every country, product and manufacturer. Approval releases commercial tables to the final store. Embedding indexes documents for search. Retrieval selects the passages and rows relevant to a question. Citation names the source of each claim in the answer.
The assistants depend on every step. Without the context layer, answers would be generic or impossible to verify. With it, they are grounded in current data, consistent with the reports leadership reads, and traceable to a named source. Analysts, reports and alerts read the same layer, which is why the data work came first. It is the same principle behind Nirmitee's retrieval pipelines for clinical data: retrieval quality is set by the data it retrieves from.
What each data source contributes
Each source answers a business question. The requirement documents grouped them into five purposes: market size, forecasted demand, tenders, pricing and product pipeline. Funding calls and news add research funding and early signals.
Pricing: UNICEF, WHO MI4A and PAHO price data. These sources show what public buyers pay for each vaccine, by supplier, presentation and year. They support tender pricing and price benchmarking, and the brief asked for price range analysis across procurement routes. WHO's Market Information for Access to Vaccines programme and the UNICEF Supply Division are the reference points for this market.
Tenders and supply: UNICEF contract awards and Gavi shipment reports. Contract awards show which manufacturers win which supply contracts, at what value and for what purpose. Shipment reports show planned and actual deliveries to each country. Together they form the tender history the team asked for and reveal competitors' positions in donor-funded markets.
Market access: Gavi country eligibility and transition status. Gavi classifies countries by financing stage, from initial self-financing to fully self-financing. A country moving through transition is moving from donor-supported procurement towards buying vaccines itself. The brief asked for the market to be viewed by procurement route, and this status is what defines the route.
Demand: immunisation coverage and national schedules. Coverage estimates show how many children each programme reaches. Schedules show which vaccines each country gives and at what age. The design combines birth cohorts with schedules to size demand in doses, country by country.
Market size and ability to pay: World Bank, UNDP and Transparency International. Population, births, mortality, GDP, GNI and income group describe each market's size and ability to pay. Development and corruption indices add country context.
Regulatory status and competitor pipeline: WHO prequalification, CDSCO minutes and clinical trials. WHO prequalification is required to supply UN agencies, so a new prequalification is a competitive event, and the team asked to be notified of each one. Minutes of India's Subject Expert Committee on vaccines show which domestic applications are progressing. The WHO clinical trial registry shows which candidates are in development and at which phase.
Funding: 18 funding agencies. Monthly collection surfaces open grant calls, with deadlines, eligibility and focus areas, for research and development that may qualify for grant funding. The funding assistant answers questions such as how many calls are open for a disease area.
Early signals: news, X, RSS and annual reports. Keyword-filtered news reaches each role as a daily digest across eight categories, including regulatory approvals, clinical pipeline, competitor activity, outbreaks, procurement policy, manufacturing and supply, intellectual property, and research funding. When a test showed that an outbreak notice on WHO's website had not appeared on WHO's X account, the team asked for the websites themselves to be monitored, and web news monitoring on approved domains was added.
What was delivered
Alongside the two assistants described below, the platform delivers six capabilities in production.
- A self-updating market database. A scheduler checks every 15 minutes which of the 54 integrations are due, and administrators can start any refresh by hand.
- Web data extraction from difficult sources. Browser automation with Selenium and Playwright, PDF data extraction, and an official API wherever one exists. In July 2026 World Bank demography moved from browser automation to the API, because the downloadable file had stopped at 2023.
- Entity resolution across sources. Approved mapping tables for countries, products and manufacturers, a review queue for unknown names, and a registry of values to skip.
- A human approval gate. Ten commercial and programme datasets reach reporting only after an administrator approves them.
- Role-based reporting. Around two dozen Power BI reports, including country, product and manufacturer views, pricing, tenders, prequalification and clinical trials, read the approved tables and open from the platform's analytics page. Report access follows role and confidentiality level, and the client reviewed that access model before go-live.
- Alerts and digests. Daily keyword digests by role, run status emails and alerts for unmapped names.
The application is a custom software build on the Frappe framework, with no ERPNext dependency. Frappe supplied data types as configuration, review workflows, role-based permissions, scheduled jobs, email and an audit trail on manual edits, so engineering time went into the sources. The build followed Nirmitee's product engineering practice from requirements through user acceptance testing and release. Heavy analytics stayed in Power BI, which the client already used.
Governance and trust: human in the loop where it matters
Governance is built into the pipeline from the start. It follows the staged design Nirmitee uses in its data pipeline development work. Each source moves through three stages: the raw capture as downloaded, a cleaned version and the final record. When a number looks wrong, the team traces it back to what the website published and reprocesses it without collecting it again. Record fingerprints ensure the same real-world row is stored once, even when a source republishes its full history.
The approval gate. Ten datasets remain gated: UNICEF prices, contract awards and shipments, WHO purchase prices and product list, the PAHO price list, immunisation coverage and schedules, Gavi country status and clinical trials. Their records wait in staging until an administrator approves them, and approval is refused while any name in the table is still unmapped. This is human-in-the-loop design applied to data rather than to model output. The gate was never meant to cover every source indefinitely. At the client's request in June 2026, World Bank statistics and country financials moved to direct loading, joining WHO prequalification and the CDSCO minutes. Funding calls, documents and news feeds were never gated; they feed search, alerts and the funding list.
Name matching. The platform checks every incoming name against the approved mapping tables. An unknown name goes to a review queue, and each batch is reported by email with a spreadsheet of what was found. An administrator maps it once, so future runs resolve it automatically, or marks it as junk, so future runs skip it.
Allowed sources and citations. Web citations shown to users are limited to eight domains: who.int, gavi.org, cepi.net, gatesfoundation.org, path.org, dcvmn.org, ghitfund.org and polioeradication.org. Every answer is saved with its sources, so a reader can check what the assistant relied on. These controls follow the same data governance principles of lineage and access control that Nirmitee applies to clinical data.
The AI layer: a research assistant with citations and a text-to-SQL agent
Research questions need documents and judgement. Counts, totals and deadlines must come exactly from the data. The platform therefore runs two assistants.
The research assistant answers questions about the market from the document library and a live web answer. A cited answer is built in six steps:
- On the first message of a conversation, the question is rewritten to spell out ambiguous acronyms, so "Gavi" means the vaccine alliance and "WHO" the World Health Organization.
- The question is embedded with Azure OpenAI and matched against the library in Azure Cosmos DB, which returns the three closest passages.
- Those passages go into a prompt that instructs the model to cite each claim as a library document or a named website.
- Perplexity drafts the answer from the passages and current web sources.
- Web citations outside the approved domain list are removed.
- The answer is shown with two reference sections, one for documents and one for websites, and saved with its sources.
Citations name the document or website. They do not claim a page number.
The funding assistant is a LangChain tool-calling agent on Azure OpenAI with two tools. When a question asks for a count, a total or a specific year, the agent writes a SQL query against the structured funding table, so the figure comes from the data itself. Other questions, such as what a funder currently prioritises, go to semantic search over the funding call documents. A name map lets users refer to about 21 agencies by common short forms.
Both are governed AI agents: narrow in scope, grounded in the platform's own data and explicit about sources. Nirmitee's AI solutions for healthcare and life sciences follow the same order of work: the context layer first, the assistant second. Teams that already have the data can go straight to Nirmitee's AI and machine learning services for the assistant itself.
Where this applies
The pattern fits any team that tracks many public websites and PDFs and needs clean, reviewable data with AI on top. These are applications of the pattern, not work delivered for this client.
- Payer and prior authorization policy intelligence. Coverage policies and prior authorization rules are published by many health plans as web pages and PDFs. The same collection, name matching and review gate can map them to one set of codes and plan names before they reach a rules engine or a prior authorization integration. See which health plans publish their prior authorization criteria.
- Payer operations data. Payer rules and remittance guidance that feed payer and clearinghouse integrations, reviewed before they change production behaviour.
- Pharma market access. Price lists, tender results, regulatory decisions and reimbursement lists across countries, matched to one product and manufacturer master.
- Tender and procurement monitoring. Procurement portals in many formats, with alerts routed by product line and a person approving what reaches the pipeline report.
- Competitive intelligence. Competitor approvals, trials, partnerships and news, filtered by keyword and role, with an assistant that cites its sources.
In each case the platform sits beside a client's healthcare interoperability estate, turning outside data into governed context.
Every integration, and what each one contributes
The platform runs 54 live integrations: 23 public datasets from eight organisations, 18 funding agencies plus the US federal grants.gov portal, which two of them share, four news and document feeds, and eight platform services. Legacy and retired connectors are not counted. The tables below are the full inventory.
A. Public datasets (23 from 8 organisations)
"After approval" means new records wait in staging until an administrator approves them. "Direct" means an official reference source that loads straight into the final table. "Knowledge base" means documents that feed the research assistant.
| # | Dataset | Organisation | How it is collected | Where it goes | What it contributes |
|---|---|---|---|---|---|
| 1 | Vaccine price data | UNICEF Supply Division | PDF table extraction; proxy fallback for protected pages; unchanged PDFs cached | After approval | Price benchmarks by vaccine, supplier and year |
| 2 | Contract awards | UNICEF Supply Division | PDF table extraction with proxy fallback; PDFs cached | After approval | Which manufacturers win which supply contracts |
| 3 | Gavi shipment reports | UNICEF Supply Division | PDF table extraction; each plan year replaced as a whole | After approval | Planned and actual deliveries by country and product |
| 4 | Supplier and technical meeting documents | UNICEF Supply Division | Browser automation, PDF download | Knowledge base | Supply outlook context for the research assistant |
| 5 | National immunisation coverage (WUENIC estimates) | UNICEF data service | API (SDMX) | After approval | Programme reach, an input to demand sizing |
| 6 | MI4A vaccine purchase price data | WHO | Excel download | After approval | Prices paid by self-procuring countries |
| 7 | MI4A vaccine product list with prequalification status | WHO | Excel download | After approval | Product and manufacturer master reference |
| 8 | Prequalified vaccines | WHO | CSV export | Direct | Which products can supply UN agencies; new approvals |
| 9 | National vaccination schedules | WHO | Excel download | After approval | Doses per child by country, for demand sizing |
| 10 | ICTRP clinical trial registry (vaccine trials) | WHO | Browser automation, CSV export, compared with the previous snapshot | After approval | Competitor candidates and trial phases |
| 11 | SAGE meeting documents | WHO | Browser automation, PDF download | Knowledge base | Immunisation policy recommendations |
| 12 | Revolving Fund vaccine price list | PAHO | PDF table extraction; each year's list replaced as a whole | After approval | Prices in the Americas procurement route |
| 13 | Country eligibility and transition status | Gavi | Browser automation, web scrape | After approval | Which countries are moving towards self-procurement |
| 14 | Country transition documents | Gavi | Browser automation, PDF download | Knowledge base | Detail behind each country's transition |
| 15 | Vaccine Investment Strategy documents | Gavi | Browser automation, PDF download | Knowledge base | Which vaccines Gavi plans to support |
| 16 | Market shaping documents | Gavi | Browser automation, PDF download | Knowledge base | Gavi's supply and market priorities |
| 17 | Board minutes, current and historical | Gavi | Browser automation, PDF download | Knowledge base | Funding and policy decisions as they are taken |
| 18 | Demography, 12 indicators (population by age band, births, deaths, infant, neonatal and under-five mortality, fertility, urban share) | World Bank | API; updated by country and year | Direct | Birth cohorts and market size |
| 19 | GDP and GNI, 5 indicators | World Bank | Browser automation, Excel download | Direct | Ability to pay |
| 20 | Income classification | World Bank | Excel download | Direct | Country segmentation by income group |
| 21 | Human Development Index | UNDP | Excel download | Direct | Country development status |
| 22 | Corruption Perceptions Index | Transparency International | Headless browser scrape (Playwright) | Direct | Country context for market prioritisation |
| 23 | Subject Expert Committee (vaccine) meeting minutes | CDSCO, India | Browser automation of a drop-down form, PDF table extraction | Direct | Domestic regulatory progress by applicant |
B. Funding agencies (18, collected monthly)
Together these sources contribute one current list of open grant and funding calls for vaccine and global health research. On the first day of each month a headless browser visits each funder's calls page, reads linked PDFs and compares what it finds with the calls already held. New and changed calls go into one structured funding table and are indexed for the funding assistant. For Unitaid and Wellcome, Azure OpenAI writes a short summary of eligibility and scope. Each run sends start, finish and error emails.
| # | Funding agency | Source collected |
|---|---|---|
| 1 | Gates Foundation | Grant opportunities, plus Grand Challenges |
| 2 | US Department of Defense | DARPA opportunities, plus grants.gov |
| 3 | USAID | grants.gov |
| 4 | NIH / NIAID | Funding opportunity listings |
| 5 | BARDA | Medical countermeasures opportunities |
| 6 | CEPI | Calls for proposals |
| 7 | Unitaid | Calls for proposals, with AI summary |
| 8 | Wellcome | Funding schemes, with AI summary |
| 9 | Open Philanthropy | How-to-apply page |
| 10 | PATH | Current requests for proposals |
| 11 | GHIT Fund | Open calls |
| 12 | RIGHT Foundation | Open calls, with PDF reading |
| 13 | UK Department of Health and Social Care | Opportunities via Innovate UK Business Connect |
| 14 | UK Foreign, Commonwealth and Development Office | International development funding on GOV.UK |
| 15 | EU4Health | EU Funding and Tenders portal |
| 16 | INSERM (ANRS) | Calls for proposals |
| 17 | BIRAC | Calls for proposals |
| 18 | ICMR | Calls for proposals |
C. News and document feeds (4)
| Feed | How it works | How often | What it contributes |
|---|---|---|---|
| X (Twitter) | X API, posts from a configured list of accounts | Weekly | Announcements for digests and home page news |
| RSS | Configured feeds, seeded with WHO News, Fierce Pharma and STAT News | Every minute | Industry and policy news for daily digests |
| Web news | Perplexity search for the latest items on each approved news domain | Every 4 hours | Items published on websites but not on social channels |
| Annual reports | Perplexity finds annual report PDFs on approved domains; the platform downloads them | Monthly | Organisation strategy and results for the research assistant |
Users also add material directly: library uploads, meeting notes whose attachments are indexed automatically, keyword lists and master data.
D. Platform services (8)
| Service | What it contributes |
|---|---|
| Azure OpenAI (chat models) | Reasoning for the funding assistant, funding call summaries, and text parsing inside some pipelines |
| Azure OpenAI (embeddings) | Turns documents and questions into vectors for search |
| Azure Cosmos DB with vector search | Stores and searches library and funding document vectors |
| Perplexity (Sonar) | Research answers with web citations, web news, annual report discovery |
| ScraperAPI | Fallback fetcher, used only when a direct request fails |
| Azure Data Lake Storage | Availability records from the uptime probe, shown as a monthly uptime report |
| Microsoft Power BI | Around two dozen reports on the approved final tables, opened from the analytics page |
| Email (SMTP) | Run status, unmapped-name alerts, role digests, feedback tickets and funder run reports |
For your engineering team
- Application: Frappe v15, Python, MariaDB, Redis queues; one custom app with 25 modules, 210 data types, 28 workspaces and 18 two-state workflows.
- Collection: requests and BeautifulSoup, Selenium, Playwright, pdfplumber, PyMuPDF and tabula; World Bank and UNICEF SDMX APIs; ScraperAPI as a fallback only.
- ETL pipeline: raw, cleaned and final stages; SHA-256 record fingerprints, business-key fingerprints for PDF sources, full-year replacement for snapshot sources; reprocessing from the cleaned stage.
- AI: LangChain, Azure OpenAI chat and embeddings, Azure Cosmos DB vector search, Perplexity Sonar; PDFs chunked at 1,000 characters with 200 overlap.
- Operations: Frappe scheduler, per-run email alerts, five-minute availability probe.
Working with Nirmitee
Nirmitee is ISO 27001:2022 certified. Public sources change without notice, so a platform like this is never finished at launch; Nirmitee was still adapting collectors and shipping change requests in September 2026.
A first conversation is a scoping call. Bring the list of sites, files and systems your team checks by hand, and the questions you want an assistant to answer. Nirmitee will show which sources can be automated, which need a person in the loop, and what the approval and citation rules should be. Book a scoping call.


