About me

I’m a Senior Research Software Engineer at IBM Research in the Healthcare & Life Sciences group, where I design and build multimodal AI systems and LLM-based agents for accelerating drug discovery. My current work focuses on aligning structured biological data—such as molecular representations and protein sequences—with large language models to enable richer reasoning, retrieval, and decision support for biomedical researchers working on target screening and therapeutic design.

Before focusing on drug discovery workflows, I worked extensively with Electronic Health Records (EHRs), tackling the full spectrum of clinical NLP problems—from question answering over EHR data and clinical summarization to large-scale benchmarking of domain-specific versus general-purpose language models. This included building clinical foundation models, evaluating their reasoning over longitudinal patient records, and contributing to multiple award-winning systems for clinical NLP challenges.

Prior to joining IBM Research in 2016, I completed my Master of Science in Computer Science (with a focus on Natural Language Processing) at Columbia University in Fall 2015.

My CV

Publications

  • BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning

    Ching-Huei Tsou, Michal Ozery-Flato, Ella Barkan, Diwakar Mahajan, Ben Shapira.
    arXiv, 2025. Paper

  • Towards accelerating small molecule drug discovery with pre-trained, late fusion multi-view models

    Joseph Morrone, Diwakar Mahajan, Hongyang Li, Shreyans Sethi, Elif Eyigoz, Bc Kwon, Dan Platt, Partha Suryanarayanan
    American Chemical Society (ACS) Fall Meeting, 2024. Paper

  • Clinical natural language processing for secondary uses

    Yanjun Gao, Diwakar Mahajan, Özlem Uzuner, Meliha Yetisgen.
    Editorial Journal of Biomedical Informatics, 2024. Paper

  • Special Issue on Clinical Natural Language Processing for Secondary Use Applications (2024)

    Yanjun Gao, Diwakar Mahajan, Özlem Uzuner, Meliha Yetisgen.
    Guest Editor Journal of Biomedical Informatics, 2024. Paper

  • Do we still need clinical language models?
    🏆 Best Paper Award
    Eric Lehman, Evan Hernandez, Diwakar Mahajan, Jonas Wulff, Micah J Smith, Zachary Ziegler, Daniel Nadler, Peter Szolovits, Alistair Johnson, Emily Alsentzer
    Conference on Health, Inference, and Learning (CHIL), 2023. Paper

  • MISMATCH: Fine-grained Evaluation of Machine-generated Text with Mismatch Error Types

    Keerthiram Murugesan, Sarathkrishna Swaminathan, Soham Dan, Subhajit Chaudhury, Chulaka Gunasekara, Maxwell Crouse, Diwakar Mahajan, Ibrahim Abdelaziz, Achille Fokoue, Pavan Kapanipathi, Salim Roukos, Alexander Gray
    Findings of the Association for Computational Linguistics (ACL), 2023. Paper

  • National NLP Clinical Challenges (n2c2) - Contextualized Medication Event Extraction Challenge

    Workshop Organizer
    Diwakar Mahajan, Ching-Huei Tsou, Jennifer J Liang, Özlem Uzuner
    National NLP Clinical Challenges (n2c2) - American Medical Informatics Association (AMIA), 2022. Link

  • Toward Understanding Clinical Context of Medication Change Events in Clinical Narratives

    Diwakar Mahajan, Jennifer J Liang, Ching-Huei Tsou
    American Medical Informatics Association (AMIA) Annual Symposium Proceedings, 2021. Paper

  • IBMResearch at MEDIQA 2021: toward improving factual correctness of radiology report abstractive summarization

    🏆 2nd Place Award - MEDIQA Challenge 2021
    Diwakar Mahajan, Ching-Huei Tsou, Jennifer J Liang
    Proceedings of the 20th Workshop on Biomedical Language Processing, Association for Computational Linguistics (ACL), 2021. Paper

  • emrKBQA: A Clinical Knowledge-Base Question Answering Dataset

    Preethi Raghavan, Diwakar Mahajan, Jennifer J Liang, Rachita Chandra, Peter Szolovits
    Proceedings of the 20th Workshop on Biomedical Language Processing, Association for Computational Linguistics (ACL), 2021. Paper

  • SemEval-2021 Task 9: Fact Verification and Evidence Finding for Tabular Data in Scientific Documents (SEM-TAB-FACTS)

    Workshop Organizer
    Nancy X. R. Wang, Diwakar Mahajan, Marina Danilevsky, Sara Rosenthal
    Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), Association for Computational Linguistics (ACL), 2021. Paper

  • AI-assisted tracking of worldwide non-pharmaceutical interventions for COVID-19

    🏆 IBM Outstanding Technical Achievement Award
    Parthasarathy Suryanarayanan, Ching-Huei Tsou, Ananya Poddar, Diwakar Mahajan, Bharath Dandala, Piyush Madan, Anshul Agrawal, Charles Wachira, Osebe Mogaka Samuel, Osnat Bar-Shira, Clifton Kipchirchir, Sharon Okwako, William Ogallo, Fred Otieno, Timothy Nyota, Fiona Matu, Vesna Resende Barros, Daniel Shats, Oren Kagan, Sekou Remy, Oliver Bent, Pooja Guhan, Shilpa Mahatma, Aisha Walcott-Bryant, Divya Pathak, Michal Rosen-Zvi
    Nature, Scientific Data, 2021. Paper

  • Identification of Semantically Similar Sentences in Clinical Notes: Iterative Intermediate Training Using Multi-Task Learning

    🏆 1st Place Award - National NLP Clinical Challenges (n2c2) 2019
    Diwakar Mahajan, Ananya Poddar, Jennifer J Liang, Yen-Ting Lin, John M Prager, Parthasarathy Suryanarayanan, Preethi Raghavan, Ching-Huei Tsou
    Journal of Medical Internet Research (JMIR) Medical Informatics, 2020. Link

  • A Hybrid Model for Drug-Drug Interaction Extraction from Structured Product Labeling Documents

    🏆 1st Place Award
    Diwakar Mahajan, Ananya Poddar, Yen-Ting Lin
    Text Analytics Conference, 2019. Link

  • IBM Research System at TAC 2018: Deep Learning architectures for Drug-Drug Interaction extraction from Structured Product Labels

    🏆 2nd Place Award
    Bharath Dandala, Diwakar Mahajan, Ananya Poddar
    Text Analytics Conference, 2018. Link