Authors Vijayalakshmi MDepartment of Computer Science Engineering, MVJ College of Engineering, Bangalore, IndiaRajkumar CDepartment of Computer Science Engineering, MVJ College of Engineering, Bangalore, India Abstract Sign language recognition (SLR) plays a crucial role in bridging the communication gap between the deaf and hearing communities. However, accurate interpretation of sign gestures remains challenging due to variations in hand shape, movement dynamics, occlusion, lighting conditions, and signer-specific differences. This research proposes a deep learning–based framework that integrates the fusion of visual and motion modalities to enhance recognition performance. The visual modality captures spatial features such as hand configuration and facial expressions using convolutional neural networks (CNNs), while the motion modality extracts temporal dynamics through optical flow representations and sequential modeling techniques such as Long Short-Term Memory (LSTM) networks. A multimodal fusion strategy is implemented at the feature level to combine spatial and temporal information effectively, enabling robust gesture classification. The proposed system is evaluated on benchmark sign language datasets, demonstrating improved accuracy, robustness, and generalization compared to unimodal approaches. Experimental results indicate that multimodal fusion significantly enhances recognition performance, particularly for dynamic and continuous gestures. The framework offers a scalable and real-time solution suitable for assistive communication systems, human–computer interaction, and inclusive educational technologies. This study highlights the effectiveness of deep multimodal integration in advancing intelligent sign language interpretation systems. Keywords vision-based interpreter American Sign Language Convolutional Neural Network ASL CNN Citation of this Article Vijayalakshmi M, & Rajkumar C. (2025). Fusion of Visual and Motion Modalities for Sign Language Recognition Using Deep Learning. Journal of Artificial Intelligence and Emerging Technologies (JAIET). 2(8), 26-29. Article DOI: https://doi.org/10.47001/JAIET/2025.208004 Licence Copyright (c) 2026 Journal of Artificial Intelligence and Emerging Technologies. This work is licensed under a Creative Commons Attribution Non Commercial 4.0 International Licence. References Starner, T., Weaver, J., & Pentland, A. (1998). Real-time American Sign Language recognition using desk and wearable computer-based video. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(12), 1371–1375.T. Starner and A. Pentland, “Real-Time American Sign Language Recognition Using Desk and Wearable Computer Based Video,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 20, no. 12, pp. 1371–1375, 1998.A.Kumar and S. Kumar, “Indian Sign Language Recog- nition Using Hybrid Features and Deep Neural Networks,” Springer Advances in Intelligent Systems and Computing, 2021, pp. 235–246.F. Ronchetti, E. Quiroga, C. Estrebou, L. Lanzarini, and A. Rosete, “LSTM-Based Continuous Sign Language Recognition Using Skeleton Data,” J. Comput. Sci., vol. 28, pp. 20–33, 2019.H. Oyedotun and A. Khashman, “Deep Learning for Gesture Recognition: A Review,” Neural Comput. Appl., vol. 31, no. 3, pp. 817–828, 2019.M. Saarinen, “Hand Shape Classification Using Con- volutional Neural Networks,” Proc. IEEE Int. Conf. Image Processing, 2019, pp. 1845–1849.J. Camgoz, S. Hadfield, O. Koller, and R. Bowden, “Neural Sign Language Translation,” Proc. IEEE Conf. Com- puter Vision and Pattern Recognition, 2018, pp. 7784–7793.Pigou, L., Dieleman, S., Kindermans, P., & Schrauwen, B. (2015). Sign language recognition using convolutional neural networks. European Conference on Computer Vision Workshops.Molchanov, P., Gupta, S., Kim, K., & Kautz, J. (2015). Hand gesture recognition with 3D convolutional neural networks. IEEE Conference on Computer Vision and Pattern Recognition Workshops.Zhang, C., Tian, Y., & Xu, C. (2016). Real-time sign language recognition based on deep learning. Pattern Recognition Letters, 83, 1–8.Camgoz, N. C., Koller, O., Hadfield, S., & Bowden, R. (2017). SubUNets: End-to-end hand shape and continuous sign language recognition. IEEE International Conference on Computer Vision.Wadhawan, A., & Kumar, P. (2020). Sign language recognition systems: A decade systematic literature review. Archives of Computational Methods in Engineering, 27, 785–813.Howard, A. et al. (2019). Searching for MobileNetV3. IEEE International Conference on Computer Vision.Lugaresi, C. et al. (2019). MediaPipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172.