Medical News Blog Information

Showing posts with label prediction. Show all posts
Showing posts with label prediction. Show all posts

Google Flu Trends: What did you expect?

I posted this on Crawford Kilian's H5N1 blog in response to his positing yet another story whacking Google Flu Trends for its "failure".

In case you can't tell - I'm a little sick of the number of electrons being wasted on writing the same thing about this paper in Science. I know, there is no shortage of electrons. Still, I hope to see this same degree of ire elicited by and directed toward other places, corporations and States who have trouble providing data to the public within the expected realms of accuracy. I'd also hope for more focus on what and how we test now and how representative that is of what a virus is doing; or what we might be missing.


I think Olson et al said it well when noting GFT's earlier failure to predict the H1N1 2009 pandemic's influenza-like illness activity..
"Current internet search query data are no substitute for timely local clinical and laboratory surveillance, or national surveillance based on local data collection"
The post...

Okay. Google Flu Trends (GFT) was not 100% accurate. Wow. Who'd would have thunk it? Who could possibly have guessed this would happen? The disappointment is clearly widespread. A predictive computer-based system set up for devising regulatory guidelines, formulating vaccine formulations, ensuring suitable laboratory testing capacity and preparation or national surveillance guidelines failed. Wait. What? It wasn't setup for any of that! It�s really just a pretty thing you can go look at to get an estimate of flu activity near you; much easier to wade through than some country's public health efforts. Estimate. When did we expect an estimate to be perfect?

Come on people-interpreting-this-paper. GFT isn't a failure unless you were honestly expecting it to be 100% correct.

Of course it couldn't ever be that. THERE. WAS. NO. VIRUS. TESTING. Not done by GFT anyway. Some lab testing went into it apparently, but even that was a sliver of a slice of a shard. And if you know anything about respiratory virus testing, then you know that even the testing we do, represents only a tiny fraction of the amount of virus-positive cases out there, extrapolating from those. That testing even varies from place-to-place in type, quantity and extent of reporting. The choice of what to test (sampling) is itself biased in a number of ways, not the least of which is that we favour testing pretty sick people or those that feel crook enough to present to a Doctor. We�re comparing GFT�s �fail� to an estimate. You�re all comfortable with using that to lambaste GFT? You�re comfortable to call that a total fail?

"The folks at Google figured that, with all their massive data, they could outsmart anyone."

Really? Is that what the folks thought? Did Google really get bitten by the flu bug?; can Google truly not track the flu? Certainly catchy headlines one and all. I guess no-one would read something entitled "Google Flu Trend's estimates not in agreement with some national testing data which also represents only a portion of those who get infected". I can see where that might not be a real mouse-wheel turner.

GFT was and could only ever be a predictive system. Just like that shiny App you have on your phone that predicts the weather forecast. Let's drag "big weather" through the interwebs flailing it at every turn so we can suitably express our righteous indignation at its failure to predict the rain we wanted on the weekend. It failed! OMG! Now I have to water my lawn to stop it from drying up. But that's all I have to do. No-one died when the clouds held their watery payload. My child was no more or less safe because the weather bug bit the Bureau of Meteorology here in Queensland. I didn�t have to get a new lawn because it is now 24-hours drier.

Does GFT's overestimate of the number of predicted cases by 0.5-2 fold (depending on the story you read) really have a real-world impact on anyone? Seriously? Keep in mind that its estimates still followed the trend of flu activity pretty closely; they peaked when actual flu was peaking, just not (my other estimates) perfectly. But apparently someone 100% concordance between lab sampling and GFT estimate data.

GFT has been doing a perfectly good job given what it is and what it could ever hope to be in its current setup. Perhaps centralizing and plotting the WORLD'S lab-based data alongside Google �flu�-related search-result data would be a useful next step for GFT. Then we could make up our own
minds.

In the meantime, keep it in context people.


References...

Google Flu Trends: not so perfectly predictive?

I'm no expert at the algorithms that go into the search giant's Google Flu Trends (GFT) predictive website so take what follows as a very superficial opinion. It does not surprise me at all that a recent paper in Science [1], backing up previous chatter on this subject [2], finds GFT is is not very accurate. Specifically, it has been overestimating peak influenza levels compared to more traditional laboratory-confirmed cases (itself only a subset of all cases) and influenza-like illness presentations to Doctors (a non-specific method of trying to identify influenza from a swarm of other ILI-capable viruses). 

A note: this recent paper is more a look at big data and whether it deserves our complete trust yet (it doesn't, is the message) than it is an analysis of how best to predict influenza virus activity in the future.

It would be fantastic if we could have a predictive system that could work around the need for actual testing of sampled people and give us an informed guess as to what flu was doing, how long it would be doing it, how severely it would do it and when it might start and stop doing it...I just don't have a lot of faith in predictive things like this. Perhaps I've just entered into a grumpy middle-aged male phase of my life....but I think that if we want to find out what's happening, we don't need to look too much beyond simply (not so simple when it comes to lab capacity and funding of course) upping the level of testing and typing that we do.

Even now, the current situation in Queensland of a 2-fold increase in influenza virus notifications compared to the mean of the past 5-years does not really show up so clearely on GFT.


Given that so many variables will contribute to a person's choice to search for "flu" (or whatever related text GFT includes in its algorithms), it makes perfect sense that a website showing flu activity in your area that is based on that component of the results, will be an over-estimate, especially during the peak times of flu activity. Why then? Imagine the impact on search when the media is most active in trying to get your pageviews using headline banners with "killer flu" or "early flu season" in them. People don't just chat over the back fence in response to those headlines any more, they go looking to the internet to provide their answers, news and sometimes poorly communicated facts. This will not just indicate that the have the flu, it will reflect concern that they may get it at some point in the future.

GFT also taps "real" flu data from real testing labs and Doctors clinics. This means its performance is probably not "off the rails" wrong, just overly influenced by non-infectious factors at peak times.

Are the inflated results positively affecting flu vaccine uptake I wonder? That would be a good thing. Might even have an impact on the size of the peak season.

Of course, no one knows what the actual numbers of flu cases in the community are; because flu is not always a serious disease that leads us to get a sampled collected and tested. The serious disease gets to a hospital and does get tested. these get added to notifications. Sure, influenza virus can cause a more serious outcome than many other respiratory virus infections, especially in certain groups, including death on occasion. There are also many mild infections that fly "under the radar". Those numbers won't be accounted for anywhere except through modelling. Perhaps the overestimate isn't that much of an overestimate; very hard to actually know that.

It's that damn iceberg tip again. 

References...

  1. http://www.sciencemag.org/content/343/6176/1203.summary?rss=1
  2. http://www.nature.com/news/when-google-got-flu-wrong-1.12413

Avian influenza A(H7N9) virus re-emergence risk factors...

Hard to believe it was over 6-months ago that we heard so much about H7N9. Papers are still coming out thick and fast describing all manner of aspects of the virus, its impact, transmission and ways to intervene in its replication.

There have been no new cases reported since July and the tally remains at 136 (including Taiwan case) with 44 deaths.

In a recent article in Scientific Reports, Fang and colleagues from China look into risk factors. 

It's hard to know how broadly applicable these data can be given the massive area and population covered and the relatively few cases identified.

Nonetheless the authors main predictors of re-emergence of H7N9 infections in humans are:
  1. Poultry markets and their environments
  2. Human population density
  3. Irrigated lands (exposed to waterfowl; carried by waterfowl)
  4. Built-up areas (see #2)
  5. High humidity
  6. Temperature around 15�C (citing drop in cases with rise in temperatures)

Like Us

Blog Archive