A LAZY GREEDY APPROXIMATION TO THE NP-HARD PROBLEM OF CHOOSING WHICH DRUGS AN ADVERSE DRUG EVENT CORPUS ANNOTATES
Main Article Content
Keywords
Pharmacovigilance, adverse drug reactions, spontaneous reporting systems, text corpora, therapeutic classification
Abstract
Background: Automated detection of adverse drug reactions from text is evaluated on a few corpora holding whatever drugs their sources discussed, yet the figures are read as general extraction performance.
Objective: To measure the therapeutic composition of a widely used adverse drug event corpus, and how much of the reaction space it reaches.
Methods: The CADEC corpus (1,250 forum posts, 5,796 reaction spans) was mapped to active ingredients and Anatomical Therapeutic Chemical classes. 51 quarterly US Adverse Event Reporting System extracts, 2004 to 2016Q3, gave 30,297,215 drug records, names normalised by a validated induced dictionary. Coverage is the share of an unseen drug's reaction mass inside the union of each corpus drug's 20 most frequent terms.
Results: The corpus holds 12 products but two ingredients, atorvastatin (80.0% of posts) and diclofenac, spanning 3 of 14 anatomical main groups. Against 1,526 profiled ingredients it covered 17.7% of an average unseen drug's reaction mass, 99.6% of drugs below 50%; normalisation reached 94.6% accuracy end-to-end. Reaching 80% took 208 ingredients chosen deliberately against over 500 at random, while the most reported drugs did worse than chance beyond two; conclusions held across 12 configurations. Even for its own ingredients only a quarter to a third of concepts aligned.
Conclusion: The reach of an adverse drug event corpus is bounded by the number of active ingredients it contains. Here that bound is structural, not a poor choice of the two, though choosing badly costs several fold. Performance should be reported as drug-specific, and therapeutic breadth treated as a design parameter.
References
2.Lazarou J, Pomeranz BH, Corey PN. Incidence of adverse drug reactions in hospitalized patients: a meta-analysis of prospective studies. JAMA. 1998;279(15):1200-5. doi:10.1001/jama.279.15.1200
3.Hazell L, Shakir SA. Under-reporting of adverse drug reactions: a systematic review. Drug Saf. 2006;29(5):385-96. doi:10.2165/00002018-200629050-00003
4.Rawson NSB. Canada's adverse drug reaction reporting system: a failing grade. J Popul Ther Clin Pharmacol. 2015;22(2):e167-72.
5.Sarker A, Ginn R, Nikfarjam A, et al. Utilizing social media data for pharmacovigilance: a review. J Biomed Inform. 2015;54:202-12. doi:10.1016/j.jbi.2015.02.004
6.Golder S, Norman G, Loke YK. Systematic review on the prevalence, frequency and comparative value of adverse events data in social media. Br J Clin Pharmacol. 2015;80(4):878-88. doi:10.1111/bcp.12746
7.Karimi S, Metke-Jimenez A, Kemp M, Wang C. Cadec: a corpus of adverse drug event annotations. J Biomed Inform. 2015;55:73-81. doi:10.1016/j.jbi.2015.03.010
8.Gurulingappa H, Rajput AM, Roberts A, et al. Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports. J Biomed Inform. 2012;45(5):885-92. doi:10.1016/j.jbi.2012.04.008
9.Herrero-Zazo M, Segura-Bedmar I, Martínez P, Declerck T. The DDI corpus: an annotated corpus with pharmacological substances and drug-drug interactions. J Biomed Inform. 2013;46(5):914-20. doi:10.1016/j.jbi.2013.07.011
10.Li J, Sun Y, Johnson RJ, et al. BioCreative V CDR task corpus: a resource for chemical disease relation extraction. Database (Oxford). 2016;2016:baw068. doi:10.1093/database/baw068
11.Sakaeda T, Tamon A, Kadoyama K, Okuno Y. Data mining of the public version of the FDA Adverse Event Reporting System. Int J Med Sci. 2013;10(7):796-803. doi:10.7150/ijms.6048
12.Wong CK, Ho SS, Saini B, et al. Standardisation of the FAERS database: a systematic approach to manually recoding drug name variants. Pharmacoepidemiol Drug Saf. 2015;24(7):731-7. doi:10.1002/pds.3805
13.Nelson SJ, Zeng K, Kilbourne J, et al. Normalized names for clinical drugs: RxNorm at 6 years. J Am Med Inform Assoc. 2011;18(4):441-8. doi:10.1136/amiajnl-2011-000116
14.Brown EG, Wood L, Wood S. The medical dictionary for regulatory activities (MedDRA). Drug Saf. 1999;20(2):109-17. doi:10.2165/00002018-199920020-00002
15.Lin J. Divergence measures based on the Shannon entropy. IEEE Trans Inf Theory. 1991;37(1):145-51. doi:10.1109/18.61115
16.Karp RM. Reducibility among combinatorial problems. In: Miller RE, Thatcher JW, editors. Complexity of Computer Computations. New York: Plenum Press; 1972. p. 85-103. doi:10.1007/978-1-4684-2001-2_9
17.Feige U. A threshold of ln n for approximating set cover. J ACM. 1998;45(4):634-52. doi:10.1145/285055.285059
18.Nemhauser GL, Wolsey LA, Fisher ML. An analysis of approximations for maximizing submodular set functions—I. Math Program. 1978;14(1):265-94. doi:10.1007/BF01588971
19.Hochbaum DS, Pathria A. Analysis of the greedy approach in problems of maximum k-coverage. Nav Res Logist. 1998;45(6):615-27. doi:10.1002/(SICI)1520-6750(199809)45:6<615::AID-NAV5>3.0.CO;2-5
20.Minoux M. Accelerated greedy algorithms for maximizing submodular set functions. In: Stoer J, editor. Optimization Techniques, Part 2. Lecture Notes in Control and Information Sciences, vol 7. Berlin: Springer; 1978. p. 234-43. doi:10.1007/BFb0006528
21.Cormode G, Karloff H, Wirth A. Set cover algorithms for very large datasets. In: Proceedings of the 19th ACM International Conference on Information and Knowledge Management. New York: ACM; 2010. p. 479-88. doi:10.1145/1871437.1871501
22.Collins GS, Reitsma JB, Altman DG, Moons KG. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. Ann Intern Med. 2015;162(1):55-63. doi:10.7326/M14-0697
23.Ransohoff DF, Feinstein AR. Problems of spectrum and bias in evaluating the efficacy of diagnostic tests. N Engl J Med. 1978;299(17):926-30. doi:10.1056/NEJM197810262991705
24.Metke-Jimenez A, Karimi S. Concept identification and normalisation for adverse drug event discovery in medical forums. In: Proceedings of the First International Workshop on Biomedical Data Integration and Discovery (BMDID 2016), co-located with the 15th International Semantic Web Conference; 2016 Oct 17; Kobe, Japan. CEUR Workshop Proceedings, vol. 1709.
25.Limsopatham N, Collier N. Normalising medical concepts in social media texts by learning semantic representation. In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Berlin, Germany: Association for Computational Linguistics; 2016. p. 1014-23. doi:10.18653/v1/P16-1096
26.Cohen KB, Fox L, Ogren PV, Hunter L. Corpus design for biomedical natural language processing. In: Proceedings of the ACL-ISMB Workshop on Linking Biological Literature, Ontologies and Databases: Mining Biological Semantics. Detroit, MI: Association for Computational Linguistics; 2005. p. 38-45.
27.Chapman WW, Nadkarni PM, Hirschman L, et al. Overcoming barriers to NLP for clinical text: the role of shared tasks and the need for additional creative solutions. J Am Med Inform Assoc. 2011;18(5):540-3. doi:10.1136/amiajnl-2011-000465
28.Banda JM, Evans L, Vanguri RS, et al. A curated and standardized adverse drug event resource to accelerate drug safety research. Sci Data. 2016;3:160026. doi:10.1038/sdata.2016.26
29.Tatonetti NP, Ye PP, Daneshjou R, Altman RB. Data-driven prediction of drug effects and interactions. Sci Transl Med. 2012;4(125):125ra31. doi:10.1126/scitranslmed.3003377
30.Bate A, Evans SJ. Quantitative signal detection using spontaneous ADR reporting. Pharmacoepidemiol Drug Saf. 2009;18(6):427-36. doi:10.1002/pds.1742
31.Harpaz R, DuMouchel W, Shah NH, et al. Novel data-mining methodologies for adverse drug event discovery and analysis. Clin Pharmacol Ther. 2012;91(6):1010-21. doi:10.1038/clpt.2012.50
32.Hauben M, Reich L, DeMicco J, Kim K. 'Extreme duplication' in the US FDA Adverse Events Reporting System database. Drug Saf. 2007;30(6):551-4. doi:10.2165/00002018-200730060-00009.

