REMAP-CAP Oseltamivir Results and Discussion
By Kert Viele, Ph.D.
REMAP-CAP recently announced results on the use of oseltamivir for critically ill influenza patients. Overall, about 14% of patients died on the control arm, as opposed to about 20% on each of the two oseltamivir arms (two different doses). The REMAP-CAP team announced the top level results and various analyses of the data, with the prespecified primary analysis reporting a 98% chance of harm.
There are many interesting aspects to these data. The trial was stopped early for potential harm to patients, not all sites randomized to control, and the covariate adjusted analyses differ a fair amount from the unadjusted analyses. This has resulted in several blogs being written criticizing the conclusions of REMAP-CAP. It’s worth diving into each of these issues and determine which factors are really driving the results and where we might be concerned. Note that, to my knowledge, none of the current takes conclude that oseltamivir should be used in this population, the central issue is how strong are the results for harm as opposed to simply “no benefit”.
An example blog was written by Dr. Josh Farkas. This blog brings up lots of important issues about these results that I think are worthwhile to discuss. It’s important to discern which aspects of the data and analyses drive different conclusions. The blog has some very strong opinions on Bayesian statistics. I don’t think the Bayesian aspects of REMAP-CAP are actually connected to the issues raised, but at the same time I think those issues are independently important. I have some possibly surprising agreement with some key points.
The blog is here:
https://emcrit.org/pulmcrit/remap-cap-oseltamivir/
We structure our response in the same order of the points in Dr. Farkas’s blog, with a synthesis at the end. Note I’m taking numbers from the initial press release for comparability to Dr. Farkas’s blog. A preprint is now available at:
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7172531
This preprint includes some extra follow-up and thus has slightly different estimates.
The first issue involve randomization to control. Many REMAP-CAP sites felt they did not have equipoise to randomize to control, and thus only randomized among the active doses of oseltamivir. Thus, not all sites are randomizing on the main “oseltamivir versus control” question. This creates two relevant datasets, the full dataset, containing all sites, as well as a restricted dataset using only the sites that randomized to control. REMAP-CAP used the full dataset as the primary analysis, and the restricted dataset as a prespecified sensitivity analysis.
I can see debate over which dataset should drive the primary analysis. I’m not as concerned about this in REMAP-CAP, however, because the two analyses agree. The primary gives a 98% posterior probability with OR point estimates of 2.19 and 2.33 in the two oseltamivir arms (greater than 1 is harmful), while the sensitivity gives a 96% posterior probability of harm with point estimates of 2.31 and 2.37. Regardless of which analysis you prefer, you get to the same place. For what it’s worth, I tend to value looking at all the available data, and here the mortality rates in the oseltamivir arms are equal with or without randomization to control. None of that guarantees there aren’t systematic differences, but it’s reassuring. And again, if you aren’t reassured, just use the sensitivity restricted to the sites randomizing to control.
There are good reasons REMAP-CAP ends up in this situation. Clearly REMAP-CAP can’t force sites to randomize to control if they don’t want to. In addition, REMAP-CAP is exploring multiple domains. The sites that don’t randomize to control are randomizing in other domains, and if the question had become “which dose of oseltamivir is best”, they would have provided a randomized comparison between active doses. Essentially, having these sites not randomize to control does not “contaminate” the other sites. They generate extra data. You can use it or not, and here it doesn’t matter which way you choose.
In summary, the sites not randomizing to control are not driving the result.
Covariate adjustment. REMAP-CAP had a prespecified covariate adjustment, as do many clinical trials (and discussed in guidance documents from all major regulatory agencies). This covariate adjustment is clearly having an impact. I’ll focus on the sensitivity analysis (sites that randomized to control) as that is clearly preferred by the blog and the two analyses are similar.
The blog notes that, unadjusted for covariates, the two-sided fisher exact test p-values for the two doses are 0.2 (10mg) and 0.4 (5mg). I ran the pooled (both doses together) fisher exact test as well and got two sided p=0.24. As I understand the objection made, the issue is the difference between 0.2 and 0.4 versus the 96% posterior probability. Why are these different?
A large part of this difference is simply one sided versus two-sided inferences. The posterior probability of harm is one-sided, computing the probability the odd ratios is greater than 1. Thus, they should be compared to the one-sided p-values, which would be 0.1 and 0.2.
From here we ask whether differences between the frequentist unadjusted tests and the Bayesian covariate adjusted tests are due to “frequentist/Bayesian” or due to “unadjusted/adjusted”. I think it’s the latter, with minimal impact of frequentist/Bayesian.
I did an unadjusted Bayesian analysis, computing the probability of harm using Beta(0.5,0.5) priors on the mortality rate in each arm. The posterior probabilities of harm are 90% for the 5mg dose (almost perfectly in line with a one sided p=0.1) and 83% for the 10mg dose (not quite as much agreement, but pretty close to a one sided p=0.2). The pooled analysis also shows around a 90% chance of harm. Thus, the unadjusted Bayesian analyses seem quite in line with the unadjusted frequentist analysis.
The full analysis includes covariates and combines (not pools) the information on the two doses (having both doses give similar results is stronger than only having one). This takes that 90% posterior probability without covariates to 96% with covariates. This isn’t that dramatic. Unfortunately, we don’t have enough information to compute an analogous frequentist test, but I would suspect we might get a one-sided p-value near 0.04, two-sided 0.08.
This doesn’t seem like an unusual move for a covariate adjustment. The point estimates do move a fair amount (there is a lot of variability attached) toward harm. It will be important in the manuscript to describe why this occurs, and of course that may result in further questions.
In summary, a large part of the apparent discrepancy between p=0.2,0.4 and posterior probabilities of 96-98% is that the p-values are two sided and the posterior probabilities are one sided. The remainder appears to be the fact there are two doses that confirm each other and the covariate adjustment. I don’t see any evidence the Bayesian piece matters here.
How do we monitor and describe probabilities of harm? I think this is one of the most important aspects of the blog. My read is that the blog is concerned about any claim of harm with a nonsignificant result (from the blog “the authors admit that these data are not statistically significant by standard benchmarks”).
Stopping a trial for lack of benefit or harm is not a mirror image of stopping for efficacy. For efficacy, we choose to continue trials until the evidence reaches a high bar (small one-sided p-values less than 0.025, posterior probabilities above 97.5%, perhaps even stronger to account for multiple interims, and even stronger when two trials are considered). When monitoring for lack of benefit or potential harm, it is not ethical to continue exposing patients to potential harm simply to change a posterior probability of harm from 70% to 97.5% (or from 90% to 97.5%, or whatever the values may be). Once trial success is out of reach, it is unethical to continue to randomize patients, particularly if those patients are being exposed to a credible risk of harm.
Thus, we will rarely obtain “statistically significant” evidence of harm. Focusing on the 96%, it’s not statistically significant. But it’s much different than 1% or 51%. There is much more evidence of harm than benefit, and that is worth noting. If our policy were to carefully monitor trials, stop for potential harm, and then always dichotomize results like 96% into “not significant = no harm”, we would significantly downplay significant risks. Absolutely note 96% is not 99.9%, but dichotomization alone does not properly convey the evidence. Note that it’s exceptionally hard to get this kind of language correct when looking at social media and press coverage.
From here there is a lengthy discussion about potential problems with Bayesian designs. As states above, I don’t see the Bayesians aspects of REMAP-CAP driving the issues mentioned so far. But I do see the discussion here as raising important general points.
The first involves a prediction that Bayesian inferences will eventually fail to be confirmed. I’d like to spend a little time on how such a confirmation would be conducted, and whether we are confirming the posterior probabilities or the decisions.
I’m going to assume away one of the hardest parts of any such confirmation, which is getting some sort of “gold standard” knowledge of the truth of these therapies. So I’m just going to assume that we have conducted a large number of Bayesian trials, and that, somehow, eventually the true parameters are revealed.
I consider it a great strength of Bayesian analysis that it makes easily testable assertations. If/when we know the parameters, we can ask straightforward questions like “of trials that claimed a 90% chance of benefit, are 90% truly beneficial” (or 80%, or 20%, or 99%, or whatever). We could look at posterior probabilities of particular levels of benefit (or harm) and see if those predictions are made at the correct frequencies. We could look at aspects of the distributions, for example, do we find that 50% of the real parameters are above the posterior medians. I won’t say this is impossible from a frequentist perspective, but I don’t think it’s as natural. I think Bayesians should welcome such an exercise, as long as it is done in good faith (for example, we may find that Bayesian inferences should be made with more skeptical priors, etc., that doesn’t mean Bayesian is “wrong”, but that it can be improved. I’m not clear on how such an exercise could find frequentist is right and Bayesian is wrong, given the common correspondence between frequentist methods and Bayesian methods using noninformative priors. For what it’s worth, I have heard talks from pharmaceutical statisticians who have tested their predictions for whether phase 3 trials would be successful, and who claim good correspondence. But to my knowledge this hasn’t been published. I get why (lots of private info) but it would be REALLY interesting to see in more detail.
A separate issue is what decision to make with the posterior distribution we have. We can be perfect in our probabilistic assessment and still make significant errors. For example, suppose I create a decision rule that says I will approve any therapy whose probability of effectiveness is over 1%. Even if my assessments of the probabilities are spot on, with this decision rule I will be approving a lot of ineffective and harmful therapies.
To me, getting the probabilities correct and getting the decision correct are connected, but different problems. You can do one correctly and mess up the other one. My read of the blog is that this is the central issue, that some Bayesians both simultaneously switch from frequentist to Bayesian and simultaneously pick thresholds which are too low.
On this issue I agree! I’ve seen some claims following this chain of reasoning “I got a one-sided p-value of 0.04 which wasn’t significant. When I do a Bayesian analysis I get a 96% posterior probability of efficacy. The 96% seems really high, so now approve my drug”.
There may be reasons to adjust the standards for approval from default levels (p<0.025, posterior probabilities greater than 97.5%, note when requiring two trials the program success levels are much stronger) to something else. High mortality with an unmet need may call for adjusting the level of evidence required for approval. But this should be done with a clear-eyed assessment of the risks and benefits involved in the decision, both for individual patients and for society, not simply because 0.04 and 96% sound different. Thus, I worry greatly that some Bayesians potentially risk claims of “lowering the bar” when a change in threshold is made “because I was Bayesian” as opposed to a utility based assessment of why the bar for success should change.
Specific to REMAP-CAP, there really is no decision of “harmful” versus “no benefit”. Both result in not giving oseltamivir in this population. As opposed to “gaslighting” (the blog’s accusation), the discussion I’ve seen has been trying to navigate the challenge of noting that there is a decent amount of evidence of harm, explicitly without claiming oseltamivir is conclusively harmful. This is hard to do perfectly in every communication. If the direction were reversed, and we were making a statement giving oseltamivir in this population, then I would not have stopped the trial with a 96% (or 98%, given the other issues, interims, etc.) posterior probability.
Finally, despite the negativity toward Bayesian, the article closes on a very Bayesian note, which I also agree with. Should we have used an informative prior? The mechanism by which oseltamivir would cause harm is not clear, and while I am not an expert in the clinical aspects of this area, my understanding is that a good “state of the field prior to REMAP-CAP” opinion would likely have limited prior probability on oseltamivir causing extensive harm. Such an informative prior would indeed pull the point estimate toward equality and away from harm and likely reduce the posterior probability of harm (this depends on the details of the prior). At the end of the day, this seems like a convincing argument that we should take the 96% posterior probability and/or the point estimate with a grain of salt. It is a common and straightforward Bayesian exercise to compute results with a range of priors from optimistic to skeptical.
I’m not sure what REMAP-CAP could have done on this issue. I would imagine that REMAP-CAP would have faced significant pushback for such an informative prior (I can see statements like “rigging the trial in favor of oseltamivir” even though I disagree with them). There is also a significant block of people arguing that trial results should be reported “on their own”, applying such prior distributions in supplementary analyses (I don’t fully agree with this either, but that’s another discussion). In any case, it’s just ironic that a particularly anti Bayesian article also argues the prior wasn’t informative enough. What’s a trialist to do?
So in summary…I’m not that worried about the sites not randomizing to control, as the results are confirmed by the sensitivity analysis. I also don’t see the frequentist/Bayesian results as that different, keeping in mind the p-values should be one sided for this comparison. I think the covariate adjustment is fairly modest here, but am interesting in seeing more detail about how and why the point estimates moved between the adjusted and unadjusted analyses. I wouldn’t like it if the trial had stopped for efficacy with the 96% posterior probability, but harm and efficacy aren’t symmetric and I think the authors are doing their best to convey “decent evidence of harm” without overstepping to claiming the results reach statistical significance.
As to the general Bayesian comments, I like the idea of trying to confirm Bayesian results, and think the ability to do so naturally is a strength of Bayesian inference. When doing so, we need to separate the idea of confirming the probability statements with the decisions made (we can make or miss either of these individually). I worry about Bayesians decreasing the threshold for success (not harm) without justifying this on clinical utility grounds. And I’m all in favor of taking into account prior information and feel it would matter here, lowering but not eliminating the likelihood and degree of harm.
The blog has brought out a number of important issues, and for that I thank the author, but was it necessary to accuse the trialists of unethical behavior (“gaslit”, “admit”, “sleight of hand”)? Our discourse is already so nasty, why bring this into science? We have important issues to discuss, why introduce them in ways that cheapen that discussion, encourages people to be tribal rather than nuanced, to be emotional rather than thoughtful? I don’t know everyone on the REMAP-CAP team, but I know a lot of them, some of them I’ve known almost since they were born. They are doing their best, and care deeply about getting things right. It’s important to disagree and to question, but disagreeing with someone does not make them immoral.