Wednesday, September 19, 2007

My O'Brien-Fleming design is not the same as your O'Brien-Fleming design

I know this discussion is a little technical, and nonstatisticians can probably skip this, but I hope that a statistician struggling with the O'Brien-Fleming design and its implementation in SAS/IML (notably the SEQ, SEQSHIFT, and SEQSCALE functions) can find this from a search engine and save hours of headache.

There are two ways of designing an O'Brien-Fleming design, a popular design for conducting interim analyses of clinical trials. The first method is to use an error (or alpha) spending function, which essentially gives you a "budget" of error you can spend at each interim analysis. The second is to realize that, if you are looking at cumulative sums in the trial, the O'Brien-Fleming design terminates if you cross a constant threshhold. In the popular design programs LD98 and PASS 2007, the spending function approach is used. In the book Analysis of Clinical Trials using SAS, (a book I highly recommend, by the way), the cumulative sum approach is used at the design stage (the spending function is used at the monitoring stage). When interim analyses are equally spaced, the two approaches give the same answer. When interim analyses are not equally spaced, the two approaches seem to give different answers. What's more, the spending function for O'Brien-Fleming as implemented in LD98 and PASS are different from what they show you in the books. They use:

4 - 4*PHI(z(1-alpha/4)/sqrt(tau))

for two-sided designs.

They don't tell you these things in school. Or in the books.

Update: Steve Simon's post on the topic has moved as of 11/21/2008. Please see the third comment below.

Monday, September 17, 2007

He makes the data sing their story

Hans Rosling gave a TED talk in 2006. If you love to work with data, you must watch this.



By the way, I am a statistician, and I love it. Yeah, this is all observational and "hypothesis-generating," as we like to say, but letting the data sing their story tells us where to concentrate our efforts.

Saturday, September 1, 2007

Bias in group sequential designs - site effect and Cochran-Mantel-Hanszel odds ratio


It is well known that estimating treatment effects from a group sequential design results in a bias. When you use the Cochran-Mantel-Haenszel statistic to estimate an odds ratio, the number of patients within each site affects the bias in the estimate of the odds ratio. I've presented the results of a simulation study, where I created a hypothetical trial and then resampled from this trial 1000 times. I calculated the approximate bias in the log odds ratio (i.e. log of the CMH odds ratio estimate) and plotted that versus the estimated log odds ratio. The line is cubic smoothing spline, made by the statement symbol i=sm75ps in SAS. The actual values are underprinted in light gray circles just to get some idea of the variability.

Sunday, August 19, 2007

Quick hits - the pharmacogenomics of Warfarin, and the statistical analysis of TGN1412

1. Terra Sigillata explains the pharmacogenomics of warfarin far better than I can.

2. Andrew Gelman discusses the statistical analysis of TGN1412, a clinical trial that resulted in a cytokine storm (which lead to multiple organ failure, cancer, gangrene, and amputation) for six very unfortunate volunteers and caused the fledgling biopharmaceutical company TeGenero to go bankrupt. I don't think anyone ever thought of doing a statistical analysis of the trial, simply because it was unnecessary. As it turns out, if you do a classical statistical analysis on the data, you get a result that isn't statistically significant! Basing a scientific conclusion on such an analysis is clearly absurd, and serves to highlight the limitations of statistics, or, rather, the way we think about and use statistics. The crux of the matter is that the adverse effects would have been significant if even one person had experienced a cytokine storm. Frequentist statistics can't pick up on that assumption, and Bayesian statistics would probably require an otherwise absurd prior to pick it up.

I guess there's a lot going on with statistics this week.

What does the First Ever Pharma Blogosphere Survey tell us

First, let me make a few comments. I find John Mack's Pharma Marketing blog useful. Marketing tends to be a black box for me. From my perspective, for the inputs you have guys who want to sell things, and for the outputs you have commercials and other promotional materials. I (partly by choice and partly by the way my brain works) understand very little about what happens between input and output. All I know about it is play my strengths and downplay my weaknesses. This is part of the reason I'm limiting myself to discussing statistical issues, at least on this blog.

However, when he came out with his First Ever Pharma Blogosphere Survey®©™, I was skeptical. In fact, I didn't pay much attention to it. But then he started making claims based on the survey, especially surrounding Peter Rost's new gig at BrandweekNRX. His predicted Brandweek would "flush its creditibility down the toilet" by hiring Rost, and cited his survey data to back up his case (he had other arguments as well, but, as noted above, I'm just covering what I know). And since I'm skeptical of his data, I'm skeptical of his analysis, and, therefore, his arguments, conclusions, and predictions based on the data. To his credit, however, he posts the raw data so at least we know he didn't use a graphics program to draw his bar graphs.
Rost's counterarguments are worthy of analysis as well. He notes that most people read the Pharma Marketing blog (the survey was conducted from its sister site Pharma Blogosphere), raising the question about which population Mack was really sampling. The correct answer, of course, is people who happened to read that blog entry around the day it was posted who cared enough to bother to take a web survey. I would agree that Mack's following probably make up a bulk of the survey.
But more important is the comparison of the survey to more objective data, such as site counters. (Note that site counter data isn't perfect, either, but it is more objective than web polls since the data collection does not require user interaction.) And it looks like that objective data doesn't match Mack's data.
Then you throw in the data from eDrugSearch, which has its own algorithm for ranking healthcare websites, but they seem a very out of line with the ranking algorithm from that of Technorati, which uses some modifications to the number of incoming links (I think to adjust for the fact that some blogs just all link to one another).
So, at any rate, you can be sure that Peter Rost will keep you abreast of his rankings, and for now they certainly do not seem to match Mack's predictions. And, while the eDrugSearch and Technorati rankings seem far from perfect, they do tend to agree on the upward trend in readership of BrandweekNRx and Rost's personal blog, at least for now. Mack's survey, and the predictions based on them, are the only data I've seen so far that have not agreed.
In the meantime, I say the proof is in the pudding. Read these sites, or, better yet, put them in an RSS reader so you can skim for the material you like. Discard the material you don't like. As for me, well, I like to keep abreast of the news in my industry because, well, it could affect my ability to feed my children. So far, Rost's blog breaks news that doesn't get picked up anywhere else, (as does Pharmalot and PharmGossip). Mack's blogs did, too, at least until he started getting obsessed with his subjective evaluation of Rost's content.

Web polls in blog entries - I don't trust them

I distrust web polls. While there are more trustworthy sources of polling such as Surveymonkey, these web surveying sites have to be backed up with essential the same type of operational techniques found in standard paper surveys. The web polls I distrust are ones that that bloggers put in their entries in their blog entries to poll their readers on their thoughts of certain issues. Sometimes they will even follow up with an entry saying "this isn't a scientific poll, but here are the results."

A small step up from this are the web surveys, such as John Mack's First Ever Pharma Blogsphere Survey®™©. They have a lot of the same problems as the simple web poll, and few of the controls necessary to ensure valid results. So I'll discuss simple one-off web polls and web surveys together.

Most of the problems and biases with these web polls aren't statistical; rather, they are operational. The data from these is so bad that no amount of statistics can rescue them. It's better not to even bring statistics into the equation here. Following are the operational biases I consider unavoidable and insurmountable:
  • Most web polls do not control whether one person can vote multiple times. Most services will now use cookies or IP addresses to block multiple votes from one person, but these services are imperfect at best. Changing an IP address is easy (just go to a different Starbucks, and cookies can be deleted). Cookies are easily deleted.
  • Wording questions in surveys is a tricky proposition, and millions of billable hours are spent agonizing over the wording. (Perhaps 75% of that is going a bit too far, but you get the point.) Very little time is generally spent wording the question of a web poll. The end result is that readers may not be answering the same question a blogger asks.
  • Forget random sampling, matching cases, identifying demographic information, or any of the classical statistical controls that are intended to isolate noise and false signal from true signal. Web poll samples are "People who happen to find the blog entry and care enough to click on a web poll." At best, the readers who feel strongly about an issue are the ones likely to click, while people who are feel less strongly (but might lean a certain way) will probably just glaze over.
  • Answers to web polls will typically be immediate reactions to the blog post, rather than thoughtful, considered answers. Internet life is fast-paced, and readers (in general) simply don't have the time to thoughtfully answer a web poll.
Web polls and surveys might be useful for guaging whether readers are interested in a particular topic posted by the blogger, and so they do have a use in guiding the future material in a blog. But beyond that, I can't trust them.

Next step: an analysis of the John Mack/Peter Rost kerfluffle. 



Shout outs

Realizations is a tiny blog, getting just a tiny bit of traffic. After all, I cover a rather narrow topic. Every once in a while, someone finds an article on here worth reading, and, even less often, they link to it.

So, shout outs (some long overdue) to the people who have linked:

Friday, August 17, 2007

Pharmacogenomics: "Fascinating science" and more

The FDA has issued two press releases in two days where pharmacogenomics has played the starring role:

Pharmacogenomics is the study of the interaction of genes and drugs. Most of the study, and certainly the most mature part of the field, has been on the study of drug metabolism, especially the cytochrome P450 enzymes, which are found in most kinds of life. The FDA's press releases are based on this science.

Pharmalot has reported on the mixed reactions to the warfarin news. One reaction was that "It's fascinating science, but not ready for prime time." Maybe not in general for all drugs, but the pharmacogenomics of warfarin has been studied for some time, and a greater understanding of the metabolism of this particular drug is critical to its correct application. Warfarin is a very effective drug, but it has two major problems:
1. The difference between the effective dose and a toxic dose is small enough to require close monitoring (i.e. it has a narrow therapeutic window)
2. It is implicated in a large number of adverse effects during ER visits (probably mostly for stroke or blood clots)

The codeine use update is even more urgent. Codeine is activated by the CYP2D6 enzyme, which has a wide variation in our population (gory detail at the Wikipedia link). In other words, the effects of codeine on people vary widely. The morphine that results from codeine metabolism is excreted in breast milk. If a nursing mother is one of the uncommon people who have an overabundance of CYP2D6, a lot of morphine can get excreted into breast milk and find its way into the baby. The results can be devastating. Fortunately, CYP2D6 tests have been approved by the FDA, and the price will probably start falling. Whether this science is ready for prime time or not (and CYP2D6 is probably the most studied of all the metabolism enzymes, so it probably is), it's fairly urgent to start applying this right away.

I applaud the FDA for taking these two steps toward applying pharmacogenomics to important problems. There may be issues down the road, but it's high time we started applying this important science.

Thursday, August 16, 2007

Good Clinical Trial Simulation Practices

I didn't realize they had gotten this far, and they did so 8 years ago! A group has put together a collection of good clinical trial simulation practices. While I only partly agree with the parsimony principle, I think the guiding principles are in general sound. I'd like to see this effort get wider recognition in the biostatistical community so that clinical trial simulations will get wider use. That can only help bring down drug development costs and promote deeper understanding of the compounds we are testing.

IRB abuses

Institutional Review Boards are a very important part of our human research. They are the major line of defense against research that degrades our humanity, and protects subjects in clinical research. Thank goodness they're there avoid a repeat of a nasty part of our history.

Unfortunately, as institutions do, IRBs have suffered from mission creep and a growing conservatism. It's a growing opinion that IRBs are overstepping their bounds and bogging down research and journalism that has no chance of harming human subjects. Via Statistical Modeling etc. I found IRBWatch, which details some examples of IRB abuses.