Friday, April 2, 2010

Adventures in graduate school

I was recently reflecting at, basic classes aside, I use information from mainly three graduate classes. Two of them were special topics classes, and one was a class that had finally evolved from a special topics class.

In one of the special topics class, we were given a choice of two topics: one field survey of Gaussian processes, which would have been useful but that was not so interesting to the professor, and local time (i.e. the amount of time a continuous process spends in the neighborhood of a point), which was much more specialized (and for which I did not meet the prerequisites) and much more interesting to the professor. I chose the local time because I figured if the professor was excited about it, I would be excited enough to learn what I needed to to understand the class. As a result, I have a much deeper understanding of time series and stochastic processes in general.

The second special topics class seemed to have a very specialized focus, pattern recognition. It covered the abstract Vapnik-Chervonenkis theory in detail, and we discussed rates of convergence, exponential inequalities on probabilities, and other hard-core theory. I could have easily forgotten that class, but the professor was excited about it, and because of it I am having a much easier time understanding data mining methods than I would have otherwise.

The third class, though it was not labeled a special topics class, was a statistical computing class where the professor shared his new research in addition to the basics. There I learned a lot about scatterplot smoothing, Fourier analysis, local polynomial and other nonparametric regression methods that I still use very often.

In each of these cases, I decided to forgo a basic or survey class for a special topics class. Because of the professor's enthusiasm toward the subject in each case, I was willing to go the extra mile and learn whatever prerequisite information I needed to understand the class. In each case as well, that willingness to go the extra mile and fill in the gaps has carried over to over a decade later where I have kept up my interest and am always looking to apply these methods to new cases, when appropriate.

I am currently taking the bootstrapping course at statistics.com and am happy to say that I am experiencing the same thing. (I was introduced to the bootstrap in fact in my computing class mentioned above but we never got beyond the basics due to time.) We are getting the basics and current research, and I'm already able to apply it to problems I have now.

Monday, March 15, 2010

Odds are, you need to read this

With my recent attacks on p-values and many common statistical practices, it's good to know at least someone agrees with me.

Odds are, it's wrong, ScienceNews

(via American Statistical Association Facebook page)

Saturday, March 13, 2010

Observation about recent regulatory backlash against group sequential trials

Maybe it's just me, but I'm noticing an increased backlash against group sequential trials from regulatory authorities in the last couple of years. The argument against these two trials seems to be twofold:


  1. Group sequential trials that stop early for efficacy tend to overstate the evidence for efficacy. While true, this can be corrected easily, and should be. Standard texts on group sequential trials, and software make the application of this correction easy.
  2. Trials that stop early tend to have too little evidence for safety.
The effect of this seems to be that the groups that need to use group sequential designs the most—the small companies who have to get funding for every dollar they spend on clinical trials—are being scared away from them especially in Phase 3.

The second point about safety is a major one, and one where the industry would do better to keep up with the methodology. Safety analysis is usually descriptive because hypothesis testing doesn't really work so well, because Type I errors (claiming a safety problem where there is none) is not as serious a problem as a Type II error (claiming no safety problem where there is one). Because safety issues can take many different forms (does the drug hurt the liver? heart? kidneys?) there is a massive multiple testing problem, and efforts to control the Type I error that we are used to are no longer conservative. There is the general notion that more evidence is better (and, to an extent, I would agree), but I think it is better to solve the hard problem and attempt to characterize how much evidence we have of the safety of a drug. We have started to do this with adverse events; for example, Berry and Berry have implemented a Bayesian analysis that I allude to in a previous blog post. Other efforts include using False Discovery Rates and other Bayesian models.

We are left with another difficult problem: how much of a safety issue are we willing to tolerate for the efficacy of a drug? Of course, it would be lovely if we could make a pill that cured our diseases and left everything else alone, but it's not going to happen. The fact of the matter is that during the review cycle regulatory agencies have to make the determination of whether safety risk is worth the efficacy, and I think it would be better to have that discussion up front. This kind of hard discussion before the submission of the application will help inform the design of clinical trials in Phase 3 and reduce the uncertainty in Phase 3 and the application and review process. Then we can talk with a better understanding about the role of sequential designs in Phase 3.

Saturday, March 6, 2010

When a t-test hides what a sign test exposes

John Cook recently posted a case where a statistical test "confirmed" a hypothesis but failed to confirm a more general hypothesis. Along those same lines, I had a recent case where I was comparing a cheap and a more expensive way of determining the potency of an assay. If they were equivalent, the sponsor could get by with the cheaper method. A t-test was not powerful enough to show a difference, but I noticed one method showed consistency lower potency than the other. I did a sign test (compared the number with lower results against the expected number based on a binomial) and got a significant result. I could not recommend the cheaper method.

Lessons learned:

- t-test is not necessarily more powerful than a sign test
- a t-test can "throw away" information
- dichotomizing data is often good and exchanges one type of information (qualitative) for loss of quantitative information

Friday, March 5, 2010

Another strike against p-values

Though I'm asked to produce the darn things every day, I have grown to detest p-values, mostly for the way that people wanted to engineer and overinterpret them. The fact that they do not follow the likelihood principle serve to provide an additional impetus to want to shove them overboard while no one is looking. Now, John Cook has brought up another reason: p-values are inconsistent (in the sense that they do not provide evidence for a set of hypotheses in the way that you would expect--I suspect if they were statistically inconsistent in the sense that no unbiased test could exist they would have been abandoned a while back).

Sunday, February 28, 2010

How to waste millions of dollars on clinical trials

When I started working for my current company (a clinical research organization), I was presented with a project that had discontinued due to lack of continued funding. The sponsor simply wanted to stop enrolling and get any conclusions from the study they could find. We presented some basic summary statistics, and then they presented me with an ethical dilemma: they found a couple of numbers that "looked good" and wanted me to generate a p-value that they could put into a press release so that they could get more funding. In statistics this is known as data dredging or p-value hunting. P-values generated this way are known to be extremely biased in favor of the drug being studied simply because the brain is very good at picking out patterns. Recently, however, I read that this company had a failed Phase 3 trial, which cost millions of dollars.

In a previous life, I worked with someone whose large Phase 2 trial failed on its primary endpoint. However, a secondary endpoint looked very good. They commissioned a Phase 3 study with thousands of patients to study and hopefully confirm the new endpoint. However, that study ended up failing as well, and I believe development of the drug was discontinued.

In my opinion, those failed studies could have been avoided. A Phase 2 study need not reach statistical significance (and certainly should not be designed so that it has to), but results in Phase 2 should be robust enough and strong enough to inspire confidence going into Phase 3. For example, estimated treatment effect should be clinical relevant, even if confidence intervals are wide enough to extend to 0. Related secondary endpoints should show similar trends. Relevant subgroups should show similar effects, and different clinical sites should show a solid trend.

I personally would prefer Bayesian methods which can quantify these concepts I just listed, and can even give a probability of success in a Phase 3 trial (with given enrollment) based on the treatment effect and variation present in the Phase 2 trial. However, these methods aren't necessary to apply the concepts above.

In both of the cases I listed above, the causes were extremely worthy, and products that are able to accomplish what the sponsors wanted would have been useful additions to medical practice. However, these products are probably now on the shelves, millions of dollars too late. The end of Phase 2 can be a very difficult soul-searching time, especially when a Phase 2 trial gives equivocal or negative results. It's better to shelve the compound or even run further small proof-of-concept studies than waste such large sums of money on failed large trials.

Sunday, February 7, 2010

Barnard’s exact test -- a test that ought to be used more

Barnard’s exact test – a powerful alternative for Fisher’s exact test (implemented in R) | R-statistics blog describes the use of Barnard's test, which I think is a more preferable test to Fisher's exact.

Barnard's exact test has one further advantage over Fisher's exact: Fisher's exact requires two fixed margins (e.g. the number of subjects in a treatment group and the number of subjects with a given adverse event), whereas in most places it is used only one of the margins is fixed (i.e. the number in a treatment group but not the number with a given adverse event).

The downside is that not too many software packages implement it. Specifically, SAS doesn't seem to implement it, so it doesn't get much use in the pharmaceutical industry. Having an implementation in R is a good start, so maybe more people will explore it and it will see more use.

Tuesday, February 2, 2010

Bradford Cross's "100 proof" project

While it sounds like the way to make the perfect whiskey, the 100 proof project is actually a way to remember the fundamentals of mathematics and how to write a proof. Bradford Cross is taking up the project just for this reason, and we can all benefit.

Every once in a while I pull out the functional analysis book just to brush up on things like the spectral theorem (a ghost of grad school past), but that doesn't have the impact that this project will have.

I found the project via John Cook's AnalysisFact Twitter feed.

Friday, December 25, 2009

Merry Christmas Using R | Keep on Fighting! | Yihui Xie

Merry Christmas Using R

Seasonal silliness with my favorite open source statistical package.