The Replication Crisis

[Back to Reason, Science and Faith]

Introduction

'The replication crisis' is a name frequently given to a troubling feature of modern science.

For a long time, science (as we understand it today) was a hobby, an activity undertaken by wealthy amateurs.  They experienced all the usual human problems - such as pride, rivalry, jealousy, and desire for recognition and status, and so on - but they rarely had any incentive to falsify their results; Gregor Mendel is one obvious counter-example.

More Recently

More recently, science turned into a career, and scientists began to compete - not only for recognition, but also for funding.  There were significant rewards for those scientists who could publish frequently, and in respected journals.  The act of publishing became an end in itself, quite different from the aim of producng new and useful results.

Science also turned into a business - several businesses, in fact.  Publishing journals containing scientific papers became lucrative, and different publishers needed to compete with each other for both high quality papers and subscribers.

These two developments led to this modern problem: researchers have found that a significant portion of published scientific studies cannot be successfully reproduced. For example, the Center for Open Science's reproducibility project showed that 36% of the 100 studies they examined, all published in reputable journals in 2008, yielded significant results.

Consequently

It is clear that both funding restrictions and the need to publish quickly incentivise scientists to engage in 'questionable' research practices, such as running experiments with small sample sizes and p-hacking (see below).  Belief in the hypothesis being tested can lead to the publication of 'enhanced' results (better than the results actually obtained) in order to release the funding which will enable the experimental techniques to be refined and the desired results to be achieved.  Academic career structures are based on reputation, so new and significant results are rewarded, while methodological rigor and verification remain largely invisible and therefore unrewarded.  Consequently, simple greed and the desire for career development can each lead scientists to publish studies based on completely falsified results.

The economic pressures on journals reinforced these problems.  They benefit from publishing papers containing impressive new results, rather than null results or simple replication studies.  The cost of not publishing null results is not carried by the journals, but by the research community, because each research institute can have the same idea and spend significant time and money testing it, not knowing that this has been attempted several times before.

Techniques

While all areas of science are affected, these problems have been most obviously evident in the fields of medicine and psychology and medicine.  Drug treatment trials are almost entirely funded by the pharmaceutical companies hoping to sell the drugs, and a successful drug can generate vast amounts of profit, so there are strong incentives to produce the 'right' results.  It is easy to produce studies which over-state the effectiveness of a drug: simply undertake many small-scale studies, and only report the ones which report a higher than average effectiveness.  You then produce a meta-analysis of the research, which 'proves' the result.

Another frequently used technique goes by various names; 'p-hacking' is a common one.  'P' here refers to a statistical measure: it is the probability of obtaining the observed results (or better), given the assumption that the null hypothesis (the alternative to the hypothesis being tested) is true.  A figure of 0.05 (in other words, 1 in 20) is often regarded as statistically significant; this is an arbitrary figure, but it is often used.

There are various ways to manipulate experimental data to produce a statistically significant result without falsifying anything.

  • Firstly, you can test many different variables.  The more data you can collect, the better the chances that cross-referencing any two of them will give you a significant result - 2 variables can be compared 1 way, 3 variables 3 ways, 4 variables 6 ways, 5 variables 10 ways.  You only report on the combination of variables which gives the significant result, and present that result as the hypothesis you were testing.
  • Secondly, you can keep testing the data, and stop collecting it as soon as you get a statistically significant result.
  • Thirdly, you can adjust for experimental error, and exclude a few outlying results if this makes your data significant.
  • Finally, there are often multiple statistical tests which can be performed, or multiple assumptions about the data which affect how the statistics are calculated - for example, how many distinct populations are present in the subjects being studied?

The use of AI in modern research has made all this much easier to do.  The Replication Crisis is now so bad that organisations are automating the process of retracting untrustworthy and unethical papers, which has resulted in Springer Nature wrongly removing two studies by Max Planck.

Solution

But all this manipulation is very easy to counter: you simply establish a central repository of research and require both pre-registration of all experiments, and mandatory reporting of results.  The pre-registration would have to state the hypothesis, what is being measured, what statistical tests will be used, and what is the target population size.  The reporting would include both the study and the full set of data collected.

This will not address the problem of complete fabrication of results, but it would solve both the problem of p-hacking and the problem of repeating many experiments which do not produce interesting results.

See Also

 

E-mail me when people leave their comments –

You need to be a member of Just Human? to add comments!

Join Just Human?


Donate