Skip to main content
Moazzam Shoukat
AI Researcher • Immigration Expert
ACADEMIC SCHOLARSHIP & PAPERS

Research &
Publications.

Peer-reviewed survey papers, journal articles, and investigative studies in speech transformers, large audio foundation models, and non-invasive acoustic diagnostics.

INFOGRAPHIC • RESEARCH ARCHITECTURE

Moazzam Shoukat's AI & Speech Research Taxonomy

Mapping 7 peer-reviewed publications across foundation audio architectures, healthcare acoustic sensing, and affective computing.

Google Scholar Profile→
🔬2 Landmark Surveys

Transformers in Speech

Computer Science Review (2025)

Overcoming self-attention quadratic complexity O(T^2) in high-sample audio through conformers, subsampling, and chunked inference.

ConformersSelf-Attentionwav2vec 2.0
🔊Survey & Outlook

Large Audio Models (LAMs)

Multimodal Foundation AI (2023)

Tokenizing continuous soundscapes via neural audio codecs (EnCodec) to power real-time, low-latency generative speech agents.

Neural CodecsZero-Shot SynthesisAudio LLMs
🩺Medical Acoustic Systems

Internet of Audio Things (IoAuT)

Acoustic Healthcare Sensing (2024)

Extracting non-invasive biometric telemetry from respiratory and cardiac acoustics using edge spectrogram neural classifiers.

BiomarkersEdge AudioPreventative Care
🧠2 Affective Computing Papers

Affective Computing & Metaverse

Emotion AI Frontiers (2024 - 2026)

Multimodal cross-lingual speech emotion recognition enabling empathetic conversational avatars and synthetic voice defense.

Cross-Lingual SERVirtual AvatarsGenAI Safety
PRIMARY RESEARCH DOMAINS
🔬Speech Processing
🔬Large Audio Models
🔬Affective Computing
🔬Metaverse & Emotion AI
🔬AI for Healthcare & Audio IoT
🔬Natural Language Processing (NLP)

Published Research Papers (7)

Peer-Reviewed & Preprints
Paper 01Survey Paper
Computer Science Review (Elsevier) / arXiv:2303.11607 • 2023

Transformers in speech processing: A survey

Comprehensive state-of-the-art review covering transformer architectures across acoustic modeling, automatic speech recognition (ASR), speaker verification, and text-to-speech synthesis.

#Transformers#Speech Processing#ASR#Acoustic Modeling
Paper 02Foundation AI
arXiv:2308.12792 (Audio Foundation AI Survey) • 2023

Sparks of large audio models: A survey and outlook

Investigating emerging foundation audio models, multimodal audio-language representations, zero-shot acoustic comprehension, and the future trajectory of generative audio AI.

#Large Audio Models#Foundation AI#Generative Audio#Multimodal
Paper 03IEEE Journal
IEEE Open Journal of the Computer Society (Vol. 5) • 2024

Affective computing and the road to an emotionally intelligent metaverse

Explores real-time affective computing interfaces, emotional sentiment extraction from acoustic and physiological cues, and emotionally resonant avatars in virtual worlds.

#Affective Computing#Metaverse#Emotion Recognition#IEEE
Paper 04Flagship Journal
Computer Science Review (Elsevier) • 2025

Transformers in speech processing: Overcoming challenges and paving the future

In-depth analysis of computational efficiency hurdles, low-latency streaming constraints, cross-lingual adaptations, and next-generation transformer paradigms for audio.

#Computer Science Review#Speech Transformers#Low-Latency#Edge Audio
Paper 05Digital Health AI
IEEE Open Journal of the Computer Society (Vol. 5) • 2024

Medicine's New Rhythm: Harnessing Acoustic Sensing via the Internet of Audio Things for Healthcare

Groundbreaking evaluation of acoustic sensing, cough/respiratory sound analytics, wearable audio IoT devices, and privacy-preserving clinical diagnostics.

#Internet of Audio Things#Healthcare AI#Acoustic Sensing#IEEE
Paper 06Cross-Lingual AI
IEEE SNAMS 2023 • 2023

Breaking barriers: Can multilingual foundation models bridge the gap in cross-language speech emotion recognition?

Empirical investigation demonstrating how pretrained multilingual foundation representations bridge the acoustic-emotional divide across low-resource and cross-cultural dialects.

#Speech Emotion Recognition#Multilingual Models#Cross-Lingual Transfer#IEEE
Paper 07GenAI & Ethics
IEEE Computational Intelligence Magazine • 2026

Affective Computing in the Age of GenAI: Balancing Risks and Rewards

Critical inquiry into generative emotional synthesis, ethical boundaries of affective empathy machines, psychological influence, and governance frameworks.

#Affective Computing#Generative AI#AI Ethics#Emotion Synthesis

Interested in Research Collaboration?

Open to peer-reviewing, speech AI benchmarks, and acoustic healthcare explorations.

Contact Moazzam →