Hullathy Balakrishnan Barathi Ganesh
Ph.D. Student
Text Information Processing Lab
Kitami Institute of Technology

CV

I build bias-aware, edge-ready speech and language models for under-represented languages. Currently a Ph.D. candidate at the Kitami Institute of Technology , I am developing unified Multimodal, Multilingual, and Multitask (M3) transformers to ensure AI systems serve diverse linguistic identities with demographic fairness and structural equity.

On-going Research

May 2026 WMT-2026: Experimenting TextDecoder in Indic Machine Translation Shared Task. It focuses on low-resource Indic languages from diverse language families. The focus will be on North Eastern languages like Assamese (State: Assam), Bodo (State: Assam), Mizo (State: Mizoram), Khasi (State: Meghalaya), Manipuri (State: Manipur), Kokborok (State: Tripura) and Nyishi (State: Arunachal Pradesh).
March 2026 M3LM-Indic Corpus: Developing a 50K+ hour open corpus covering 23 languages, complete with emotion annotations and gender balancing, utilizing weak supervision for low-resource languages.
October 2025 M3LM-Indic: A bias-aware, unified Multimodal, Multilingual, and Multitask (M3) Transformer for English and all 22 scheduled Indian languages under India's Constitution. It addresses the low-resource multilingual AI scaling crisis through depth-first specialization and para-linguistic bias-aware pre-training.

Education

Oct 2025 - Present Ph.D. in Multimodal AI
Text Information Processing Lab, Department of Co-creative Engineering, Kitami Institute of Technology, Japan.
Sep 2013 - July 2015 M.Tech in Computational Engineering and Networking
Computational Engineering and Networking (CEN), Amrita School of Engineering, Coimbatore.
Sep 2009 - Apr 2013 B.E. in Electronics and Communication Engineering
Anna University

Professional Experience

Oct 2025 - Present Ph.D. in Multimodal AI
Text Information Processing Lab, Department of Co-creative Engineering, Kitami Institute of Technology, Japan.
Aug 2020 - Sep 2025 Head of Product Development
RBG.AI, Coimbatore, India.
Mar 2018 - Jun 2020 Chief Technology Officer & R&D Product Manager
Arnekt Solutions Pvt. Ltd., Pune, India
Dec 2016 - July 2017 Research Scientist Analyst
Accenture Innovation Center for Analytics, Artificial Intelligence, Accenture
Sep 2015 - Dec 2016 Assistant System Engineer
Digital Enterprise Services and Solutions, Artificial Intelligence Practice, Tata Consultancy Services
June 2015 - Aug 2015 Intern
Digital Enterprise Services and Solutions, Artificial Intelligence Practice, Tata Consultancy Services

Selected Publications

Google Scholar

Overcoming Orthographic Discrepancies via Algorithmic Translinear Pipelines and Phylogenetic Script Mapping
Barathi Ganesh HB , Michal Ptaszynski, Meenakshi, Jairam R
WMT EMNLP 2026
[abs] [pdf] [code]
Disentangled Speech Encoder: A Robust Encoder with Dynamic Adapter for Language Identification
Barathi Ganesh HB , Jairam R, Michal Ptaszynski, Reshma U, Jyothish Lal G, Premjith B
TidyLang Odyssey 2026
[abs] [pdf] [code]
AURA-ST: Acoustic-Unconstrained Residual Architecture for Speech Translation.
Barathi Ganesh HB , Michal Ptaszynski, Reshma U, Jairam R
IWSLT 2026
[abs] [pdf] [code]
A Systematic Survey of Modality Fusion Strategies in Multimodal Multilingual Multitask Speech-Text Transformers.
Barathi Ganesh HB , Michal Ptaszynski, Reshma Unnikrishnan, Jairam R, Meenakshi
Information Fusion
[abs] [pdf] [Materials]
SMM4H-HeaRD 2025: LLMs in Healthcare Applications.
Barathi Ganesh HB , Anthony Vijay M., Naren Kishor S., Sharmila B., Jairam R., Jyothish Lal G.
ICWSM 2025
[abs] [pdf] [code]

Teaching Experience

2017-2018, Odd Deep Learning and Probabilistic Graphical Models (16CN613), TA
2017-2018, Odd Deep Learning (17AL605), TA
2016-2017, Even Deep Learning for Natural Language Processing (16CN704), TA

Coursework

  • CN611: Computational Linear Algebra and its Applications
  • CN613: Computational Optimization Theory - Linear and Non-Linear Methods
  • CN624: Probability and Graphical Models
  • CN614: Advanced Data Structures and Algorithms
  • SLMA101: Distributed Algorithms and Optimization
  • CN733: Neural Networks and Deep Learning
  • CN709: Applied Computational Linguistics
  • CN736: Text Analytics
  • SLCN101: Deep Learning for Natural Language Processing
  • CN705: Speech Processing
  • CN622: Wavelet Theory and Pattern Classification

Certifications

  • Neural Networks and Deep Learning” acquired from deeplearning.ai through coursera.
  • Deep Learning Using Keras” acquired from udemy.
  • Maxima and Minima concepts : Applications of Derivatives” acquired from udemy.
  • Deep Learning Prerequisites: The Numpy Stack in Python” acquired from udemy.
  • Machine Learning Foundation” acquired from “Tata Consultancy Services” through Tata Consultancy Services internal certification program.
  • Statistics and R” acquired from Harvard through edx.
  • High-Dimensional Data Analysis” acquired from Harvard through edx.
  • Data Mining with Weka” acquired from the University of Waikato.
  • Hadoop Fundamentals” acquired from IBM.
  • Spark Fundamentals” acquired from IBM.
  • Introduction to R” acquired from DataCamp.
  • Introduction to Data Analysis using R” acquired from Big Data University.
  • Introduction to Python” acquired from DataCamp.

International Shared Tasks Participated

  • Machine Translation (EMNLP WMT-2026)
  • African & Celtic Speech-to-Text Translation track (IWSLT-2026)
  • Speaker-Controlled Language Recognition (Odyssey - 2026 TidyLang)
  • Audio Encoder Capability Challenge for Large Audio Language Models (InterSpeech-2026 AECC)
  • Cross-Lingual Speaker Verification (InterSpeech-2026 TidyVoice)
  • Multilingual Automatic Detection of Reclamation of Slurs in the LGBTQ+ Context (MultiPRIDE)
  • Fallacy Detection in Italian Social Media Texts(FadeIT)
  • Machine Translation (EMNLP WMT-2025)
  • Detection of insomnia in clinical notes (SMM4H-HeaRD 2025)
  • First Security and Privacy Analytics Anti-Phishing Shared Task (IWSPA-AP 2018)
  • Semantic Evaluation Exercises: Affect in Tweets (SemEval-2018)
  • PAN @ Conference and Labs of Evaluation Forum: Author Profiling (PAN-2017)
  • Forum for Information Retrieval and Evaluation (FIRE-2017): Gender Identification in Russian Texts (RusProfiling)
  • Forum for Information Retrieval and Evaluation (FIRE-2017): Information Retrieval from Legal Documents (IRLeD)
  • 2nd Social Media Mining for Health Application shared task (AMIA 2017)
  • Sentiment Analysis for Indian Languages: Code Mixed (SAIL 2017)
  • The 8th International Cybersecurity Data Mining Competition (CDMC-2017)
  • Semantic Evaluation Exercises: Semantic Textual Similarity (SemEval-2016)
  • PAN @ Conference and Labs of Evaluation Forum: Author Profiling (PAN-2016)
  • Forum for Information Retrieval and Evaluation (FIRE-2016): Mixed Script Information Retrieval (MSIR)
  • Forum for Information Retrieval and Evaluation (FIRE-2016): Consumer Health Information Search (CHIS)
  • Forum for Information Retrieval and Evaluation (FIRE-2016): Code Mixed Entity Extraction – Indian Languages (CMEE – IL)
  • Forum for Information Retrieval and Evaluation (FIRE-2014): Named Entity Recognition Indian Languages (NER)
  • Named Entity rEcognition and Linking: Named Entity Recognition and Linking (#Micropost2015 NEEL)

Invited Talks

  • Topic: “Data Preprocessing and Feature Engineering: The Key to Successful Models” in the Faculty Development Programme at Dr. B.R. Ambedkar Institute of Technology Sri Vijyapuram Andaman & Nicobar Islands. 20, Jan 2025.
  • Topic: “Jailbreaking LLM” at Amrita University, Coimbatore. 18, October 2024.
  • Topic: “Cyber Security in Industry 4.0” in TEQIP-III Sponsored Workshop on “Deep Learning for Big Data and Cyber Security Applications” at National Institute of Technology, Surathkal, July 2019.
  • Topic: “Workshop on Machine Learning” in National Conference on Recent Trends in Computing and Communications NCRTCC’18 at Adi Shankara Institute of Engineering and Technology, February 2018.
  • Topic: “Natural Language Processing with Deep Learning” in the Faculty Development Program (FDP) at Vidya Academy of Science & Technology, Thrissur. 20, Jan 2017.
  • Topic: “Deep Learning with Python” in the Faculty Development Program (FDP) at Mepco Schlenk Engineering College, Sivakasi. 16-17, Jan 2017.
  • Topic: “Dyanamic Mode Decomposition” from scratch . In the workshop on Data -Driven Modelling 2017 at Computational Engineering and Networking (CEN), Amrita School of Engineering, Coimbatore. 8-9, Jan 2017.
  • Topic: “Deep Learning for BioInformatics” in the workshop on “DeepChem 2017: Deep Learning & NLP for Computational Chemistry, Biology & Nano-materials”, Conducted by the Department of Computational Engineering and Networking, Amrita Vishwa Vidyapeetham University, December 22-24, 2017.
  • Topic: “Dynamic Mode Decomposition in Epidemiology” in the Workshop on “A Refresher experiential course on linear algebra and Optimization for Most Modern Signal processing and pattern classification”, Conducted by the Department of Computational Engineering and Networking, Amrita Vishwa Vidyapeetham University, 25-11,26-11 and 27-11, 2017.
  • Topic: “Natural Language Processing for Cyber Security” in the workshop on “AISec 2017 Workshop: Modern Artificial Intelligence (AI) and Natural Language Processing (NLP) Techniques for Cyber Security”, Conducted by the Department of Computational Engineering and Networking, Amrita Vishwa Vidyapeetham University, Saturday, October 28, 2017.
  • Topic: “A Journey to Text Analytics and Recent Trends in Text Analytics”, TEQIP sponsored Faculty Development Program on “Research Directions in Data and Text Mining”, organized by Government Engineering College, Idukki on January 2016.
  • Topic: “Advanced Analytics in Text and Data Mining”, Tata Consultancy Services sponsored Faculty Development Program on “Text and Data Mining” held at Tata Consultancy Services, Kochi on April 2016.
  • Topic: “Computation differs, not Optimization”, TEQIP Sponsored Faculty Development Program on “Contemporary Developments in Optimization Techniques and its Applications”, organized by TKM College of Engineering, Kollam in May 2016.
  • Topic: “Web Mining through Data and Text Mining”, Association Inauguration “ORBIT 2K16” conducted by Viswa Jyothi College of Engineering and Technology, Vazhakulam on August 2016.

Workshops and Shared Tasks Conducted

Skills

Languages

Python, R, Matlab

Frameworks & Libraries

NumPy, Pandas, SciPy, Scikit-learn, XGBoost, LightGBM, Speechbrain, NLTK, Spacy, Gensim, Pillow, Librosa, Soundfile, OpenCV, Pytorch, Transformers, PEFT, Accelerate, Fairseq2, Allenai, Speechbrain, OpenAI, Langchain, LLMLite, vLLM, ONNX, Ray, MLflow, Flask, FastAPI

Database

MySQL, Postgres, MongoDB, FAISS, Quadrant

Visualization

Matplotlib, Seaborn, Plotly, Apache Superset

Cloud

AWS (S3, SageMaker, EC2), Azure (AI Services, VMs), GCP (Vertex AI, BigQuery)

Annotation & Data Labeling

Label Studio, PWALS

OS

Linux, Windows, Mac

Documentation Tools

Libreoffice, Microsoft Office, and Latex.


Last updated on 2026-07-12