When reading the news, you may occasionally be bemused by the flip-flopping science and health stories: one week you’re told coffee is good for you, the next week that it’s bad. Red wine extends your life, then it doesn’t. Over time, new scientific studies seemingly contradict one another, leaving us unsure what evidence to believe.
As a science journalist, I’ve learnt that often the actual problem is that each study is reported in isolation, rather than weighed against everything else we know. And sometimes, the risks of harm are high. For example, in September 2025, I reported on the US Food and Drug Administration’s announcement that it would add a warning label to the painkiller acetaminophen (paracetamol), claiming that taking the drug during pregnancy increased a child’s risk of autism. It made global headlines. Yet looking closer at the science, it was clear that the FDA and its parent health department had cherrypicked a few studies, and downplayed other more robust findings. This was confirmed a couple of months later when researchers published a major review that better represented the full body of knowledge. ‘Existing evidence does not clearly link maternal paracetamol use during pregnancy with autism or ADHD in offspring,’ it concluded. By then, of course, countless women had been needlessly scared about a painkiller that is widely considered one of the safest to take during pregnancy.
In 2025, academics worldwide published about 7 million scholarly articles – that’s more than 19,000 each day. On the surface, that might seem like a number to celebrate, but it also poses a problem. As the volume of research balloons, it can be hard to discern what these millions of papers actually tell us. Within the firehose, there are rigorous methodologies and important findings, but also baffling contradictions, unconfirmed results and sloppy science. And some diverse forms of knowledge, such as lived experience and Indigenous insights, are rarely captured in scholarly articles and databases at all.
The causes are systemic. Scientists publish and promote one paper after another because that’s how they advance in their careers. Journalists breathlessly chase the latest, flashiest studies so their headlines get clicks online. Meanwhile, hardly any of us spend time trying to make sense of what the world already knows by carefully synthesising and taking stock of existing knowledge. Iain Chalmers, a doctor who co-founded the Cochrane Collaboration, an evidence-synthesis group based in London, once called this the ‘scandalous failure of science to cumulate evidence scientifically’.
This failure was a key motivation to write my book Beyond Belief: How Evidence Shows What Really Works (2026). Researching it, I discovered better ways to make sense of the world – but these methods are not as widely known as they should be. If they were, we might pause before believing news stories based on single studies, and instead recognise the real work: finding, sorting and synthesising evidence. This unglamorous labour already shapes our lives far more than any individual study or news headline will. We might also realise that, for many of the wicked problems we face, humans already possess much of the knowledge needed to solve them. All it needs is the ability to assemble it – and then act.
One of the biggest events in the history of evidence synthesis occurred in a hotel in San Francisco in 1976. There, at the annual meeting of the American Educational Research Association, the statistician Gene Glass revealed an ingenious technique that showed how to integrate findings from a large volume of individual studies. This technique would become so important to science that it has its own biography.
Glass’s work had emerged from his own struggles with mental health. A decade earlier, when he had finished his PhD in psychometrics and statistics, he was suffering from anxiety and neurosis. He began weekly psychotherapy sessions – which helped him hugely – and stayed in therapy for eight years. But his positive experience ran counter to the weight of academic opinion, which maintained that psychotherapy had little benefit. This stemmed in large part from the work of the psychologist Hans Eysenck, whose influential reviews of psychotherapy research concluded that it was worthless – or had a placebo effect at best.
Glass was irritated. Not only did this suggest that he’d flushed away money on ineffective therapy, but he felt Eysenck’s reviewing methods were flawed.
Eysenck – like other researchers at the time – tended to do what’s called ‘vote-counting’: counting up the number of studies showing a treatment has benefits, and the number that did not find it helped. The one with the most ‘votes’ wins. Intuitive though vote-counting is, it’s also misleading as a way of synthesising studies, partly because it ignores the size of the effects. One study may find that 55 people out of 100 improved after a particular treatment, and another that 85 out of 100 did. Even though the second study showed the treatment had a much stronger effect, both studies contribute one vote. Glass could see this was problematic – and was also astonished that Eysenck arbitrarily excluded hundreds of valid studies just because they were in theses or dissertations rather than published in academic journals.
The audience was thunderstruck. Here was a way to extract meaning from an apparent mess of different results
Glass set out to do a better job. He and his wife, a psychologist named Mary Lee Smith, systematically hunted down every scrap of research they could find that compared the effect of psychotherapy with a control group or another therapy – ending up with more than 370 studies.
Then Glass and Smith worked out a way to extract and combine the measurements of psychotherapy’s effect in each study. Even though one study measured the effects of the therapy on anxiety and another its impact on blood pressure, they devised a way to convert these results into one standard measure known as ‘effect size’ and then average them across all the studies. (This is somewhat like converting various currencies into dollars, so they can all be combined.) Glass called this new statistical method a meta-analysis – an analysis of analyses, just as metadata is data about data.
After toiling on this work for two years, Glass and Smith concluded that psychotherapy had a beneficial effect, and Glass presented the method at the San Francisco hotel where the education meeting was taking place. Debate continues to this day about whether psychotherapy works, when and for whom. (Eysenck called the work ‘an exercise in mega-silliness’.) But of the meta-analysis, the audience was thunderstruck. Here was a way to extract meaning from an apparent mess of different results.
The meta-analysis was not the only synthesis tool to emerge around this time. The work of Glass, Smith and other researchers also led to the systematic review, a rigorously structured method for gathering and evaluating evidence. In a systematic review, researchers scour databases worldwide of published and unpublished work for all studies that address a certain question. Then they whittle down a longlist of thousands of studies to the most relevant few, assess their reliability, extract the data, and combine the results. Many systematic reviews include a meta-analysis to pool results from the included studies and estimate the overall effect of a treatment or other intervention.
These tools to make sense of a body of evidence are one of the most important developments in science over the past few decades. They have the power to identify important conclusions that would never be possible from assessing each underlying study on its own. It’s the scientific equivalent of seeing the forest, not just the trees.
Without fanfare, these approaches have changed modern life – and particularly so when it comes to health decisions. In medicine, this happened partly thanks to a morning stroll by Iain Chalmers, an early champion of evidence synthesis, along the Wolvercote Mill Stream in Oxford, UK, in May 1991.
Chalmers had just finished a pioneering, decade-long project to synthesise all the evidence from clinical trials on treatments in pregnancy and childbirth. This work, which involved doing hundreds of systematic reviews, had shown that many standard medical practices – such as shaving women’s pubic hair during labour, and surgical episiotomies – were based on little evidence and some were harmful. The work from Chalmers and his colleagues helped change some of these practices.
Now Chalmers was thinking about expanding this work. Wouldn’t it be useful, he thought, to synthesise clinical trials in every area of medicine and healthcare? Then doctors and patients would know, based on evidence, effective ways of treating diabetes, cancer, heart disease and many other conditions. Surprisingly, many medical decisions at that time were based on conventional wisdom or the unsubstantiated opinions of senior doctors rather than on evidence from research.
Some years earlier, in 1979, a doctor called Archie Cochrane working in Cardiff, Wales, had challenged the medical community to do just this. ‘It is surely a great criticism of our profession that we have not organised a critical summary, by speciality or subspecialty, adapted periodically, of all relevant randomised controlled trials,’ Cochrane wrote. Chalmers was greatly influenced by Cochrane, and he was now ready to take on the task. He and his team started searching for clinical trials in online academic databases and scouring journals by hand in the library. The task was so big that they recruited volunteers to join the hunt, including elderly people’s groups and even the unlikely source of the Headington Bowls Club in Oxford. Soon, they had tens of thousands of trials.
Most people who have seen a Western doctor have unknowingly benefited from systematic reviews
But Chalmers knew that synthesising all the evidence on effective treatments in medicine was such a gargantuan task that it would take more people power than this. So, in 1993, he and a group of like-minded folk started the Cochrane Collaboration (now known as just Cochrane), dedicated to producing high-quality systematic reviews of evidence on the effectiveness of health treatments.
Within 10 years of starting, the Cochrane Collaboration had published around 2,000 systematic reviews and was helping to popularise this type of study. By 2010, researchers inside and outside the collaboration were publishing 11 systematic reviews on healthcare every day.
Today, the ‘Cochrane review’ is known for being one of the most rigorous and reliable syntheses of science, and systematic reviews are used to develop the clinical guidelines that doctors use to guide decisions. Most people who have seen a Western doctor have unknowingly benefited from systematic reviews. They have become an invisible bedrock of evidence on which medicine is based.
To understand what the world would look like without them, consider this notorious piece of medical advice that went unchallenged in the mid-20th century. In 1957, the paediatrician Benjamin Spock made a small change to the second edition of his parenting bestseller The Common Sense Book of Baby and Childcare (1946), reissued several more times over the years. He said that parents should put babies to sleep on their fronts rather than on their backs. By the 1980s and ’90s, studies were clearly showing that this was one of the most lethal pieces of unsubstantiated advice in the history of child health, as it contributed to a steep rise in sudden infant death syndrome (SIDS).
What makes this story even more tragic is that the link between sleeping position and SIDS might have been detected earlier – if researchers had synthesised the evidence from studies rather than looking at them one at a time. When researchers did this with a systematic review and meta-analysis published in 2005, they discovered that the link between front-sleeping and SIDS was clear in 1970, some 20 years before most parents were warned about the risks. They calculated that at least 50,000 infant deaths in Europe, Australasia and the United States might have been prevented if evidence had been synthesised and acted on earlier.
Medicine is not the only field to have built up a repository of evidence syntheses. Other disciplines – ranging from conservation to education – have mimicked the medical model.
The nonprofit International Initiative for Impact Evaluation holds more than 1,700 systematic reviews on what works in international development, such as improving sanitation and tackling malnourishment. CEEDER, a major environmental database, hosts more than 2,100 evidence reviews on what works to save species and protect the planet.
The Education Endowment Foundation, a nonprofit organisation in London dedicated to boosting learning outcomes for disadvantaged children, has constructed a widely admired ‘teaching and learning toolkit’ using systematic reviews of studies that have tested education approaches. Nearly 70 per cent of school leaders in England now use it to make spending decisions, and it has been adapted for use in more than 30 countries and translated into nine languages.
The IPCC’s syntheses have helped to inspire global action by showing that humans are causing climate change
One of the most effective strategies, it shows, is teaching students metacognition – the ability to learn how to learn. (For instance, a child who learns that writing out the steps helped her solve a maths problem is practising metacognition.) Introducing metacognition provides children with an impressive eight months of additional learning progress, on average, over a year, according to the evidence review. Repeating a school year, by contrast, sets back students’ learning by about two months.
More recently, climate scientists have been taking a leaf out of medicine’s book. They are used to synthesising evidence: the six massive scientific assessments that the Intergovernmental Panel on Climate Change (IPCC) has published since 1988 are probably the biggest and most influential evidence syntheses ever done – helping inspire global action by showing that humans are causing climate change. But the IPCC reports have so far said less about solutions, such as which climate policies could help the most.
A group called What Works Climate Solutions, which started in 2024, is now striving to build a Cochrane-like bank of evidence syntheses that show the most effective ways to cut emissions or adapt to global warming. This evidence bank will feed into the next scientific assessment of the IPCC, due to be published by late 2029, and should help governments choose which policies to adopt.
The soaring popularity of systematic reviews is visible in academic databases. The number published in medicine alone rose from around 1,400 in the year 2000 to more than 29,000 in 2019 – about 80 per day. A set of guidelines for performing them accurately, called PRISMA, has become one of the most cited papers of the 21st century – with as many as 138,000 citations, by one count.
Systematic reviews have their critics. One is that they conventionally emphasise quantitative studies – those with numerical data – and ignore qualitative research, such as that based on interviews and observations. There is a lively debate going on among evidence-synthesis experts about how best to integrate different streams of information.
In some fields, such as conservation, scientists have been reproached for privileging Western science over other forms of knowledge, such as the lived experience of local communities and Indigenous groups. Some Indigenous peoples in Canada, for example, have extensive knowledge of the behaviours of the migratory fish that they hunt. These types of evidence don’t fit in conventional reviews that just crunch numbers.
But there are ways to combine them. In 2021, Parks Canada and a group called Foundations of Success undertook an extraordinary analysis of evidence to decide whether to pursue an expensive and risky captive breeding programme for endangered caribou in Jasper National Park. As well as collecting a range of evidence from research, they asked people from Indigenous groups, zoos, and reindeer experts from Finland, to review the evidence and reach agreement about whether to proceed. (They did – and some adorable caribou calves were born in 2025.)
Over the past few decades, researchers have developed a suite of methods for combining evidence of different types. Some pool only qualitative studies, while other ‘mixed-methods’ syntheses combine qualitative and quantitative work. Some researchers who supply evidence to busy government policymakers favour the rapid review, which involves getting together the best evidence synthesis you can, to answer their question in the time available – which could be the next day. They prefer a good synthesis delivered on time rather than a perfect one that arrives too late.
Studies that claim to provide an overview of evidence convey authority, so scientists have a responsibility to get them right
Elsewhere, scientists have been working on another important innovation called living reviews. These are updated frequently and continuously as research is published rather than, as often happens, a static systematic review that quickly becomes out of date. Living reviews can be labour intensive and so are most worthwhile for important questions on which new studies are being published quickly. During the COVID-19 pandemic, a handful of groups soon rallied and developed living reviews to synthesise rapidly emerging studies on COVID therapies and vaccines.
The methods used for synthesising science have become so complex over the past 20 years that evidence synthesis has become a profession in its own right. There are entire journals dedicated to the craft, as well as academic centres, professorships and prizes. As a science journalist, I’ve interviewed evidence geeks who can talk for hours about the finer details of systematic reviews.
And sometimes tensions bubble up about the best way to synthesise a body of research. A purist argues that a gold-standard systematic review is the most rigorous way of doing it and that other methods risk missing studies or including poor-quality work and thus reaching the wrong conclusion. A pragmatist argues that a systematic review takes too long, and that often a quick-and-dirty aluminium-standard synthesis will do.
But, really, this is a storm in a teacup. Different methods for evidence synthesis suit different situations. What’s important is to be transparent about the methods used and any weaknesses they have. Studies that claim to provide a rigorous overview of evidence convey authority, so scientists have a heavy responsibility to get them right.
In a world where a news story or a social media post can spread in hours, one of the biggest problems in systematic reviewing is the time it takes. An analysis of more than 500 Cochrane reviews found that a review typically took nearly three years – and one took a glacial eight. That’s because doing a systematic review involves working through around 25 steps – carefully finding, filtering, and combining studies – of which many are even done in duplicate to avoid errors. This exacting process explains why many researchers view with horror the prospect of doing a systematic review.
Researchers have long used computers to help, but the rapid advances in generative AI are fuelling new excitement about accelerating and automating the task. Could AI help make sense of the overwhelmingly large scientific literature?
It has changed the way I think about research as a science journalist and in everyday life
Start-ups are racing to develop AI-powered systems that can find, sort and summarise research publications, but most have not yet been well tested and shown reliably to produce high-quality systematic reviews. Even so, it seems inevitable that they will improve and that before too long most of the grunt work involved in compiling evidence will routinely be done by machines, freeing up humans to check and interpret the results. That could mean that anyone would be able to assemble a reliable evidence synthesis tailored to their question, virtually at the push of a button – an appealing goal.
But it’s still some way off – and concerns about AI and evidence synthesis also abound. One is that AI is leading to a flood of inaccurate scientific reviews because researchers are using these tools to race through poor-quality efforts. Another is that people will bypass rigorous evidence syntheses entirely: because why bother, when you can just ask ChatGPT what the evidence says? (This is currently dangerous as AI tools can hallucinate and give an inaccurate picture, and most do not have access to the full corpus of the world’s research.)
One lesson from all this is to be aware that single studies can be outliers – something I’ve learned myself over the past few years. I’ve spoken with hundreds of researchers worldwide about the importance of synthesising evidence, and it has changed the way I think about research as a science journalist and in everyday life. Now, when I speak to scientists about their latest study, I try to ask how it fits with the wider body of research. When seeking evidence on a personal issue, such as a medical diagnosis, I tend to search for systematic reviews and other syntheses (while being aware that not all reviews are good quality).
We should resist the allure and glamour of new studies. It’s true that new research is important, because it leads to innovation and fresh ways to protect the planet and human lives. But if we truly want to understand the world and make it a better place, we also need to do the less glamorous, unappreciated work of collating existing knowledge – and figuring out what it all means.
Source: aeon.co




