Skip to main content

The Voice of African Enterprise

Home Health HelpMum Africa Launches MamaBench to Improve AI Accuracy in Maternal and Child Healthcare
HealthNigeria

HelpMum Africa Launches MamaBench to Improve AI Accuracy in Maternal and Child Healthcare

Share
Share

Artificial intelligence is becoming more common in healthcare, but questions remain about how reliable these systems are when making clinical decisions. To help improve the safety and quality of AI in medicine, HelpMum Africa has launched HelpMum MamaBench, an open-source benchmark designed to test how well AI models reason through maternal and paediatric diagnoses when important clinical details change.

The health technology organisation, which focuses on reducing maternal and infant mortality, says the new benchmark goes beyond measuring whether an AI model gives the correct answer. Instead, it evaluates whether the model can consistently reach the right diagnosis when faced with small but clinically important changes in a patient’s condition. Alongside the benchmark, HelpMum Africa has also published a research paper presenting findings from tests carried out on some of today’s leading AI models.

Improving AI Reliability for Better Healthcare

Most existing medical AI benchmarks assess one clinical question at a time. While this approach measures accuracy, it may not reveal whether an AI system truly understands a medical case or is simply recognising familiar patterns.

HelpMum MamaBench was developed to address this challenge. The benchmark contains 434 clinical narratives organised into 217 matched pairs, covering 371 maternal and paediatric conditions. Every narrative was written from scratch by HelpMum Africa’s medical team using first-person, patient-reported language instead of adapting existing datasets.

Each pair consists of a standard clinical case and a counterfactual version where only a small but medically significant detail has changed. That single change is enough to alter the correct diagnosis, allowing researchers to test whether an AI system can identify the shift or continues to rely on pattern recognition.

According to Dr. Abiodun Adereni, Founder of HelpMum Africa, the organisation wanted to create a benchmark that gives an honest picture of AI performance. “We wanted a benchmark that told us the truth, even when the truth is uncomfortable,” he said.

He explained that an AI system may perform well in traditional evaluations but still fail to detect a life-threatening change in a mother’s or child’s condition. “An AI system that looks accurate on a leaderboard but can’t be trusted to catch a life-threatening shift in a mother’s or child’s condition isn’t ready for real clinical use. HelpMum MamaBench exists to make that gap visible, and to push the whole field toward closing it.”

The benchmark is expected to support researchers, developers and healthcare innovators working to build AI systems that are safer and more dependable for clinical use.

Research Reveals Important Gaps in AI Performance

To evaluate current AI capabilities, HelpMum Africa tested eight configurations across four leading AI models.

The study first measured how often each model correctly diagnosed the standard clinical cases before testing whether those same models could still produce the correct diagnosis after one important clinical detail was changed.

Across all models, the difference between standard-case accuracy and performance on altered cases ranged from 16 to 28 percentage points. In some cases, models that appeared to achieve around 80% accuracy under conventional testing dropped to just above 60% when their reasoning was challenged. One widely used model recorded 77% accuracy on standard cases but only 47% when tested on the counterfactual cases.

The research team also evaluated retrieval-augmented generation (RAG), a technique that allows AI systems to consult reference material before generating answers. However, the study found that conventional RAG did not improve reasoning because it often retrieved similar information for both the original and altered cases, making it difficult for the model to recognise the clinically significant difference.

To overcome this limitation, HelpMum Africa developed Evidence-Anchored RAG (EA-RAG), a new approach that changes how supporting evidence is retrieved. Instead of gathering broadly similar information, EA-RAG identifies the specific clinical details most likely to influence a diagnosis, verifies whether those details are covered in the retrieved evidence and fills any missing information before the AI generates its response.

When applied to one of the strongest-performing AI models in the study, EA-RAG reduced diagnostic failures on altered cases by more than 20% while maintaining almost the same level of accuracy on standard cases.

Despite these improvements, the research team acknowledged that further work remains. Even the best-performing model enhanced with EA-RAG still produced incorrect diagnoses in approximately one out of every five altered clinical cases.

Adewuyi Ayomide, Research Engineer at HelpMum Africa, said the project was driven by a simple but important question.

“This started as a question we couldn’t answer with existing benchmarks: does a model actually understand a diagnosis, or has it just memorized the shape of one?”

He said MamaBench was created to answer that question more effectively, while EA-RAG represents the team’s first practical step towards improving AI reasoning.

In line with HelpMum Africa’s commitment to open and standards-based health innovation, the MamaBench dataset has been released free of charge with unrestricted public access. The organisation said the accompanying evaluation code is in its final stages of preparation and will be released soon.

By making both the benchmark and research publicly available, HelpMum Africa hopes to encourage broader collaboration among researchers, healthcare providers and AI developers, ultimately contributing to the development of more reliable artificial intelligence tools that can support better maternal and child healthcare outcomes.

Source: Tech Cabal

Share
Related Articles

Itana Expands Digital Free Trade Zone with Mr Eazi’s Choplife

Itana has signed Choplife, the entertainment and technology venture founded by Nigerian...

Clea Launches Vendor Payments to Simplify International Trade for African Businesses

Nigeria-based fintech Clea has launched Vendor Payments, a new capability designed to...

AFC Leads US$2.5 Billion Investment in Dangote Refinery Expansion

Africa Finance Corporation (AFC) has led a group of strategic investors in...

US-Nigerian Fintech Cryptofy Digital Partners With ECC Nigeria to Access 29 European Countries

Cryptofy Digital, a Nigerian and U.S.-based blockchain and fintech company, has entered...