Abstract
The rapid advancement of Artificial Intelligence (AI) has paved the way for sophisticated Human–Computer Interaction (HCI) systems that leverage multiple input modalities to enhance usability and accessibility. This paper proposes a Multimodal AI System integrating real-time hand gesture recognition, natural voice command processing, and customizable manual shortcuts to provide an adaptive and resilient interface for diverse user environments. Using standard webcams and microphones, the system applies computer vision techniques with Mediapipe and Convolutional Neural Networks for accurate gesture detection, and employs speech recognition algorithms for voice command interpretation. Manual shortcut mapping offers a reliable fallback to ensure uninterrupted control. By fusing these modalities at both feature and decision levels, the system achieves robust performance across variable lighting and acoustic conditions. Extensive evaluations demonstrate the system’s efficacy, highlighting improvements in accessibility for differently-abled users and enhanced interaction fluidity in AR/VR, assistive technologies, and productivity contexts. The proposed approach exemplifies the practical integration of heterogeneous data streams within multimodal AI, reflecting contemporary trends in intelligent interactive systems.
Keywords
Natural Language Processing
Artificial Intelligence
Human–Computer Interaction
Convolutional Neural Networks
Gesture Recognition Module
Authors
How to Cite this Article
Dr.B.Venkatesan, K.Rajesh (2025).
"MULTIMODAL AI SYSTEM WITH HAND GESTURE, VOICE, AND MANUAL SHORTCUT CONTROL".
International Journal of Contemporary Research in Computer Science and Technology,
8(1), pp. 41-43.