Hierarchical Temporal Modeling with Cross Stream Gating and Hybrid Neural Networks for Regional Sign Language Recognition
DOI:
https://doi.org/10.53799/p8ttb781Keywords:
Dual-stream architecture, Kannada sign language recognition, Mediapipe landmarks, Real-time deploymentAbstract
SLR systems are the efficient and powerful devices to help deaf and hard-of-hearing individuals. The proposed architecture design is based on a mixed temporal encoder which includes different elements, such as separable Conv1D, bidirectional LSTM, and windowed self-attention using crossstream gating to connect 3D landmark kinematics by MediaPipe with optical flow by Farneback. The proposed architecture achieved 89.53% validation accuracy and 90.0% accuracy on an independent held-out test set, with an inference latency of 48 ms on the KSL dataset consisting of 1,287 videos spanning 11 gesture classes. This paper makes two significant contributions to applied research since it develops KSL research and creates an architecture that can be used in other languages in different countries with low financial and technological resources.
References
[1] B. Alsharif, Y. Ouakrim, N. D. Alotaibi, and A. S. Mahmoud, “Deep learning technology to recognize American sign language alphabet,” Sensors, vol. 23, no. 18, p. 7970, Sep. 2023, 10.3390/s23187970.
[2] M. Alshewimy, “Efficient deep learning models based on tension techniques for sign language recognition,” Intell. Syst. Appl., vol. 20, p. 200284, Dec. 2023, 10.1016/j.iswa.2023.200284.
[3] S. U. Amin, E. Kavak, B. Kim, J. Kang, and S. Park, “Deep learning based active learning technique for data annotation and improve the overall performance of classification models,” Expert Syst. Appl., vol. 228, p. 120391, Oct. 2023, 10.1016/j.eswa.2023.120391.
[4] S. U. Amin, E. Kavak, J. Kang, and S. Park, “An efficient attention-based strategy for anomaly detection in surveillance video,” Comput. Syst. Sci. Eng., vol. 46, no. 3, pp. 3939–3958, 2024, 10.32604/csse.2024.049058.
[5] S. U. Amin, B. Kim, Y. Jung, S. Seo, and S. Park, “Video anomaly detection utilizing efficient spatiotemporal feature fusion with 3D convolutions and long short-term memory modules,” Adv. Intell. Syst., vol. 7, no. 2, p. 2400065, Feb. 2025, 10.1002/aisy.202400065.
[6] M. Aziz and A. Othman, “Evolution and trends in sign language avatar systems: Unveiling a 40-year journey via systematic review,” Multimodal Technol. Interact., vol. 7, no. 10, p. 97, Oct. 2023, 10.3390/mti7100097.
[7] N. Bansal and A. Jain, “Word recognition from Indian sign language in videos using dual feature descriptor and GMT-MASKRCNN recognition technique,” Multimedia Tools Appl., vol. 84, no. 3, pp. 2565–2597, Jan. 2025, 10.1007/s11042-024-19317-2.
[8] I. Garibay, “A survey on sign language literature,” Mach. Learn. Appl., vol. 14, p. 100504, Dec. 2023, 10.1016/j.mlwa.2023.100504.
[9] B. Hdioud and M. Tirari, “A deep learning-based approach for recognition of Arabic sign language letters,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 4, pp. 452–459, Apr. 2023, 10.14569/IJACSA.2023.0140453.
[10] X. Hong-qin and Z. Yuan-yuan, “Advanced gesture recognition method based on fractional Fourier transform and relevance vector machine for smart home appliances,” Comput. Animation Virtual Worlds, vol. 36, no. 1, p. e70011, Jan. 2025, 10.1002/cav.70011.
[11] Z. Huang, W. Xue, Y. Zhou, L. Hu, and X. Lu, “Dual-stage temporal perception network for continuous sign language recognition,” Vis. Comput., vol. 41, no. 6, pp. 1971–1986, Apr. 2025, 10.1007/s00371-024-03414-2.
[12] A. Ji, L. Cao, and S. Wang, “Dataglove for sign language recognition of people with hearing and speech impairment via wearable inertial sensors,” Sensors, vol. 23, no. 15, p. 6693, Jul. 2023, 10.3390/s23156693.
[13] A. S. M. Miah, M. A. M. Hasan, J. Shin, Y. Okuyama, and Y. Tomioka, “Spatial-temporal attention with graph and general neural network-based sign language recognition,” Pattern Anal. Appl., vol. 27, no. 1, p. 37, Mar. 2024, 10.1007/s10044-024-01240-9.
[14] S. Mohsin, M. Mohsin, and M. Alshara, “American sign language recognition based on transfer learning algorithms,” Int. J. Intell. Syst. Appl. Eng., vol. 12, no. 5s, pp. 390–399, 2023, 10.18201/ijisae.2023.350.
[15] F. M. Najib, “A multi-lingual sign language recognition system using machine learning,” Multimedia Tools Appl., vol. 83, no. 5, pp. 1345–1362, Jan. 2024, 10.1007/s11042-023-15457-z.
[16] I. Olmos-Pineda, “Word level sign language recognition via handcrafted features,” IEEE Latin Amer. Trans., vol. 21, no. 7, pp. 839–848, Jul. 2023, 10.1109/TLA.2023.10279947.
[17] M. Papatsimouli, P. Sarigiannidis, and G. F. Fragulis, “A survey of advancements in real-time sign language translators: Integration with IoT technology,” Technologies, vol. 11, no. 4, p. 83, Aug. 2023, 10.3390/technologies11040083.
[18] S. Paul, M. M. Rahaman, M. A. Hossain, M. R. Islam, and I. H. Sarker, “An Adam-based CNN and LSTM approach for sign language recognition in real time for deaf people,” Bull. Electr. Eng. Inform., vol. 13, no. 1, pp. 499–509, Feb. 2024, 10.11591/eei.v13i1.6211.
[19] K. K. Podder, M. E. H. Chowdhury, A. M. Tahir, Z. B. Mahbub, A. Khandakar, and M. S. Hossain, “Signer-independent Arabic sign language recognition system using deep learning model,” Sensors, vol. 23, no. 16, p. 7156, Aug. 2023, 10.3390/s23167156.
[20] S. Renjith and R. Manazhy, “Sign language: A systematic review on classification and recognition,” Multimedia Tools Appl., vol. 83, no. 24, pp. 77077–77127, Aug. 2024, 10.1007/s11042-024-19563-4.
[21] J. Shin, A. S. M. Miah, M. A. M. Hasan, Y. Tomioka, and Y. Okuyama, “Korean sign language recognition using transformer-based deep neural network,” Appl. Sci., vol. 13, no. 5, p. 3029, Mar. 2023, 10.3390/app13053029.
[22] R. P. Singh and L. D. Singh, “Dyhand: Dynamic hand gesture recognition using BiLSTM and soft attention methods,” Vis. Comput., vol. 41, no. 1, pp. 41–51, Jan. 2025, 10.1007/s00371-024-03556-3.
[23] R. Sreemathy, M. Turuk, S. Nema, P. Nilesh, and D. Singh, “Continuous word level sign language recognition using an expert system based on machine learning,” Int. J. Cogn. Comput. Eng., vol. 4, pp. 170–178, Jun. 2023, 10.1016/j.ijcce.2023.05.001.
[24] A. Wali, A. Zafar, S. D. Ali, K. Javed, U. Tariq, and I. H. Lee, “Recent progress in sign language recognition: A review,” Mach. Vis. Appl., vol. 35, no. 1, p. 127, Jan. 2024, 10.1007/s00138-024-01628-x.
[25] L. T. Woods and Z. A. Rana, “Modeling sign language with encoder-only transformers and human pose estimation keypoint data,” Mathematics, vol. 11, no. 9, p. 2129, May 2023, 10.3390/math11092129.
[26] S. Xiong, C. Zou, J. Yun, L. Li, and X. Lu, “Continuous sign language recognition enhanced by dynamic attention and maximum backtracking probability decoding,” Signal Image Video Process., vol. 19, no. 1, pp. 141–150, Jan. 2025, 10.1007/s11760-024-03531-8.
[27] S. Xue, L. Gao, L. Wan, and W. Feng, “Multi-scale context-aware network for continuous sign language recognition,” Virtual Real. Intell. Hardw., vol. 6, no. 4, pp. 323–337, Aug. 2024, 10.1016/j.vrih.2023.09.001.
[28] J. Yao, J. Chen, L. Niu, and B. Sheng, “Scene-aware human pose generation using Transformer,” in Proc. 31st ACM Int. Conf. Multimedia, Ottawa, ON, Canada, 2023, pp. 2847–2855, 10.1145/3581783.3611768.
[29] H. Zhang, Y. Zuo, T. Xu, F. Han, and L. Zuo, “Multipath attention and adaptive gating network for video action recognition,” Neural Process. Lett., vol. 56, no. 1, p. 124, Feb. 2024, 10.1007/s11063-024-11570-8.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 AIUB Journal of Science and Engineering

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
AJSE contents are under the terms of the Creative Commons Attribution License. This permits anyone to copy, distribute, transmit and adapt the work non-commercially provided the original work and source is appropriately cited.