Authors A.Anbarasan ArasuDepartment of Computer Science Engineering, Acharya Institute of Technology, Bangalore, IndiaP.K.RamachandranDepartment of Computer Science Engineering, Acharya Institute of Technology, Bangalore, IndiaG.Arun BabuDepartment of Computer Science Engineering, Acharya Institute of Technology, Bangalore, IndiaR.Sathish KumarDepartment of Computer Science Engineering, Acharya Institute of Technology, Bangalore, India Abstract The primary objective of the proposed sign language recognition system is to automatically interpret hand gestures into meaningful linguistic representations by analyzing their shape, spatial orientation, motion trajectory, and relative position. The work focuses on the development of a real-time, vision-based interpreter capable of recognizing and classifying alphabet gestures from American Sign Language (ASL). The system is designed as a practical, deployable assistive solution suitable for real-world human–computer interaction scenarios. The implementation employs MediaPipe Hands, a high-precision hand-tracking framework, to detect and extract 21 three-dimensional hand landmarks from live video streams captured through a standard webcam. These landmark coordinates serve as discriminative spatial features representing finger articulation and palm geometry. Instead of directly processing raw image frames, the extracted landmark vectors are normalized and structured into feature arrays to reduce computational complexity and improve invariance to scale, rotation, and background variations. For gesture classification, a Convolutional Neural Network (CNN)-based architecture is utilized to learn hierarchical feature representations from the processed landmark data. The dataset was constructed by recording multiple gesture samples under diverse environmental conditions, including varying illumination levels and backgrounds, to enhance generalization capability. The collected data underwent preprocessing steps such as noise filtering, coordinate normalization, temporal segmentation, data augmentation, and labeling to ensure robustness and improved training efficiency. The model was trained using supervised learning techniques with categorical cross-entropy loss and optimized through adaptive gradient-based optimization algorithms. To enhance usability, a text-to-speech (TTS) module is integrated into the system pipeline. Upon successful classification of a gesture, the predicted alphabet character is immediately converted into synthesized speech output, enabling seamless verbal communication. This multimodal feedback mechanism significantly improves accessibility for individuals with hearing or speech impairments. Keywords vision-based interpreter American Sign Language Convolutional Neural Network ASL CNN Citation of this Article A.Anbarasan Arasu, P.K.Ramachandran, G.Arun Babu, & R.Sathish Kumar. (2025). Multimodal Sign Language Recognition Using Convolutional Neural Networks. Journal of Artificial Intelligence and Emerging Technologies (JAIET). 2(1), 16-19. Article DOI: https://doi.org/10.47001/JAIET/2025.201004 Licence Copyright (c) 2026 Journal of Artificial Intelligence and Emerging Technologies. This work is licensed under a Creative Commons Attribution Non Commercial 4.0 International Licence. References Starner, T., Weaver, J., & Pentland, A. (1998). Real-time American Sign Language recognition using desk and wearable computer-based video. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(12), 1371–1375.T. Starner and A. Pentland, “Real-Time American Sign Language Recognition Using Desk and Wearable Computer Based Video,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 20, no. 12, pp. 1371–1375, 1998.A.Kumar and S. Kumar, “Indian Sign Language Recog- nition Using Hybrid Features and Deep Neural Networks,” Springer Advances in Intelligent Systems and Computing, 2021, pp. 235–246.F. Ronchetti, E. Quiroga, C. Estrebou, L. Lanzarini, and A. Rosete, “LSTM-Based Continuous Sign Language Recognition Using Skeleton Data,” J. Comput. Sci., vol. 28, pp. 20–33, 2019.H. Oyedotun and A. Khashman, “Deep Learning for Gesture Recognition: A Review,” Neural Comput. Appl., vol. 31, no. 3, pp. 817–828, 2019.M. Saarinen, “Hand Shape Classification Using Con- volutional Neural Networks,” Proc. IEEE Int. Conf. Image Processing, 2019, pp. 1845–1849.J. Camgoz, S. Hadfield, O. Koller, and R. Bowden, “Neural Sign Language Translation,” Proc. IEEE Conf. Com- puter Vision and Pattern Recognition, 2018, pp. 7784–7793.Pigou, L., Dieleman, S., Kindermans, P., & Schrauwen, B. (2015). Sign language recognition using convolutional neural networks. European Conference on Computer Vision Workshops.Molchanov, P., Gupta, S., Kim, K., & Kautz, J. (2015). Hand gesture recognition with 3D convolutional neural networks. IEEE Conference on Computer Vision and Pattern Recognition Workshops.Zhang, C., Tian, Y., & Xu, C. (2016). Real-time sign language recognition based on deep learning. Pattern Recognition Letters, 83, 1–8.Camgoz, N. C., Koller, O., Hadfield, S., & Bowden, R. (2017). SubUNets: End-to-end hand shape and continuous sign language recognition. IEEE International Conference on Computer Vision.Wadhawan, A., & Kumar, P. (2020). Sign language recognition systems: A decade systematic literature review. Archives of Computational Methods in Engineering, 27, 785–813.Howard, A. et al. (2019). Searching for MobileNetV3. IEEE International Conference on Computer Vision.Lugaresi, C. et al. (2019). MediaPipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172.