Datasets
Public datasets and benchmarks released with my papers.
-
[D1]EMNLP 2026
BanglaVerse: A Cultural Benchmark for Multilingual Vision-Language Models Across Bengali Dialects
Hugging Face · Vision-Language, Multilingual, Cultural Understanding · Released with EMNLP 2026
BanglaVerse benchmarks how multilingual vision–language models understand Bengali culture across historically linked languages and regional dialects. It pairs culturally grounded imagery with questions written in multiple dialects, making it possible to measure where VLMs succeed on surface-level recognition but fail on cultural reasoning. Released alongside the EMNLP 2026 Findings paper.
-
[D2]EACL 2026
MathMist: A Parallel Multilingual Benchmark for Mathematical Problem Solving and Reasoning
Hugging Face · Mathematical Reasoning, Multilingual, Benchmark · Released with EACL 2026
MathMist is a parallel multilingual benchmark for mathematical problem solving and reasoning. Because the problems are aligned across languages, it isolates the effect of language from the effect of mathematical difficulty, exposing reasoning gaps that appear only outside English. Released alongside the EACL 2026 Findings paper.
-
[D3]CVPRW 2026
VisText-Mosquito: A Unified Multimodal Dataset for Detection, Segmentation and Textual Reasoning
Mendeley Data · Multimodal, Object Detection, Segmentation · Released with CVPRW 2026
VisText-Mosquito unifies three tasks over mosquito breeding sites in a single dataset: object detection, segmentation, and natural-language reasoning about why a site is a breeding risk. Pairing pixel-level annotation with textual explanation lets models be evaluated on perception and justification together rather than in isolation.
-
[D4]Data in Brief
ElectroCom61: A Multiclass Dataset for Detection of Electronic Components
Mendeley Data · Computer Vision, Electronics, Multi-Class · Released with Data in Brief, Elsevier
ElectroCom61 is a curated dataset comprising 2,121 annotated images of 61 different electronic components collected from the Electronic Lab Support Room at United International University (UIU). It is designed to train and evaluate machine learning models for real-time component detection. Images were captured from multiple angles under diverse lighting and background conditions to reflect real-world variability. Each image was standardized through auto-orientation and resized to 640×640 pixels. The dataset is split into training (70%), validation (20%), and test (10%) sets to support robust evaluation.
-
[D5]ICLR 2024
MosquitoFusion: A Multiclass Dataset for Real-Time Mosquito Detection
Kaggle · Computer Vision, Object Detection, Mosquito Dataset · Released with ICLR 2024
The MosquitoFusion dataset features 1,204 expertly curated and annotated images aimed at advancing real-time mosquito detection systems. The data is split into training (87%), validation (8%), and test (5%) subsets. Preprocessing includes auto-orientation and resizing to 640×640 pixels, with a strict filter to exclude unannotated samples. The dataset is further enhanced using data augmentation techniques such as flipping, cropping, rotation, and grayscale conversion to boost robustness and generalizability. This high-quality dataset was accepted at ICLR 2024 (Tiny Papers Track) as an invited oral presentation in the “Notable” category.
-
[D6]IEEE Access
JailbreakTracer Corpus: Datasets for Toxic Prompt and Forbidden Question Classification
Kaggle · LLM Safety, Prompt Classification, Synthetic Data, Toxicity Detection · Released with IEEE Access
The JailbreakTracer Corpus contains two curated datasets for analyzing and classifying prompts that aim to bypass the safety mechanisms of large language models (LLMs). (1) Toxic Prompt Classification Dataset: Consists of 16,029 non-toxic and 1,952 toxic prompts initially. After synthetic augmentation using GPT, the dataset expanded to 37,333 toxic and 16,053 non-toxic prompts—addressing class imbalance and enabling robust training of LLM safety classifiers. (2) Forbidden Question Reasoning Dataset: Designed for reasoning-based classification of unethical prompts across 13 distinct categories, including Hate Speech, Malware, Economic Harm, Pornography, Legal Opinion, and more. Each class includes 8,250 samples for balanced multi-class modeling.
Contact
Happy to talk about research, collaborations, or anything in the papers above. Email is the fastest way to reach me.
Send a message
Contact details
- Email msayeedi212049@bscse.uiu.ac.bd
- Secondary email faiyazabdullah114708@gmail.com
- Phone +880 1704 054900
- Location Dhaka, Bangladesh