Modeling Temporal Relationships in Audio Signals Using MFCC And GRU to Enhance Voice Recognition Accuracy
Keywords:
MFCC-GRU-Deep learning.Abstract
Voice recognition is a prominent research trend in signal processing, given its growing role in intelligent interaction applications and biometric systems. This research presents a methodological framework that combines signal spectral characterization using Mel-Frequency Cepstral Coefficients (MFCC) with deep learning models based on Gated Recurrent Units (GRUs). GRUs are advanced models for temporal data analysis due to their ability to represent sequential dependencies with high efficiency and lower computational complexity compared to traditional recurrent models. The methodology begins by converting the audio signal into a compressed spectral representation that reflects its frequency structure. MFCCs provide an accurate description of the underlying acoustic characteristics. GRUs then utilize this representation to extract temporal variations within the signal, relying on a gating mechanism that enables them to retain relevant information over time and process long sequences without loss of context.
Experimental results showed that integrating MFCC with GRU networks significantly improved the performance of voice recognition systems, particularly in environments with time fluctuations or high noise levels. The proposed framework also demonstrated its ability to adapt to speaker differences and diverse acoustic environments.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Latakia University (formerly Tishreen) Journal for Research and Scientific Studies - Basic Sciences Series

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.