Trustworthy Multimodal AI
Understanding when multimodal models should make predictions, express uncertainty, or abstain — under missing, conflicting, noisy and distribution-shifted evidence.
My research centres on trustworthy multimodal AI — systems that remain dependable when evidence is incomplete, conflicting, noisy, or collected under different real-world conditions. I am particularly interested in when models should make predictions, express uncertainty, or abstain, with an emphasis on biomedical and healthcare applications where reliability matters more than benchmark accuracy alone. I pair this with data-centric machine learning: building open datasets, careful evaluation protocols, and reproducible baselines, because the bottleneck in applied AI is rarely the architecture — it is the quality and honesty of the data and evaluation around it.
Understanding when multimodal models should make predictions, express uncertainty, or abstain — under missing, conflicting, noisy and distribution-shifted evidence.
Applying machine learning to medical images and clinical context with an emphasis on uncertainty estimation, external validation and responsible decision support.
Areas of focus
Systems that stay dependable when evidence is incomplete, conflicting, noisy or shifted — knowing when to predict, express uncertainty, or abstain.
Machine learning for medical images and clinical context, prioritizing reliability, uncertainty estimation, external validation and responsible decision support.
Image classification, object detection and real-time vision systems for healthcare, accessibility and agriculture — from dataset development to deployment.
How dataset quality, label reliability, class balance and distribution shift affect model behaviour — building datasets and protocols for reproducible research.
Grounded language systems that retrieve evidence before generating — document intelligence, citation-backed RAG, conversational agents, tool use and evaluation.
Research Assistant under the ICSETEP Research and Development Grant (RDG) at Daffodil International University, on the sub-project “Development and Effective Application of AI-Based, Post-Quantum Cryptography-Enabled Counterfeit Medicine Identification Tools for Better Treatment Outcomes in Bangladesh” — funded by the Asian Development Bank (ADB) and the Government of Bangladesh.
Working toward advanced research in reliable multimodal machine learning — clinical-context missingness, uncertainty quantification, selective prediction and cross-hospital distribution shift.
Shakib Howlader, Md. Sabbir Ahamed, Mayen Uddin Mojumdar, Sheak Rashed Haider Noori, Shah Md Tanvir Siddiquee, Narayan Ranjan Chakraborty. “A comprehensive image dataset for the identification of eggplant leaf diseases and computer vision applications.” Data in Brief (Elsevier), Vol. 59, 111353, 2025.
This dataset comprises 4,089 high-resolution images of eggplant (Solanum melongena) leaves, systematically categorized into six distinct classes: healthy leaves and five disease types — insect pest disease, leaf spot disease, mosaic virus disease, white mold disease, and wilt disease. The images were captured using smartphone cameras against consistent white backgrounds under varying lighting conditions across multiple geographic locations, then subjected to thorough manual labelling and preprocessing to ensure accuracy and consistency. The resource is particularly suitable for applications in plant pathology, precision agriculture, and disease forecasting, where timely and accurate diagnosis is crucial. Freely available for academic research, the dataset aims to advance automated disease-detection systems and sustainable farming practices.
An openly available, manually labelled image dataset of eggplant leaves spanning healthy specimens and five disease types (insect pest, leaf spot, mosaic virus, white mold, wilt), captured across multiple locations for reproducible machine-learning research.
Designed an accessibility-focused system that recognizes ASL gestures from a live camera feed, assembles detected signs into text and synthesizes speech — combining YOLO-based detection, temporal smoothing and real-time inference.
Led the construction and open release of a 4,089-image, six-class eggplant-leaf dataset captured under varied field and lighting conditions — manually labelled and preprocessed for reproducible computer-vision research. Published in Elsevier's Data in Brief.
Contributing to AI systems that can operate responsibly in healthcare and other high-stakes settings — pairing multimodal modelling with rigorous evaluation of when systems should defer to humans.
Actively seeking research collaborations and supervision in trustworthy multimodal AI, healthcare AI, computer vision and grounded language systems.
I am actively seeking MSc supervision and research collaborations in computer vision, agricultural AI and healthcare AI.
Get in touch →