DFCRN: A Multi-Scale Complex-Domain Speech Enhancement Model for Coal Mine Scenarios
-
Abstract
Conventional convolutional recurrent networks have achieved notable progress with spectral representations; however, they are limited by insufficient receptive fields and inadequate modeling of long-range inter-frame dependencies. Moreover, traditional feedforward sequence memory networks (FSMNs) exhibit limited nonlinear expressive power in the complex domain, making them ineffective in coal-mine environments characterized by strong mechanical noise, multi-source non-stationary interference, and transmission distortions. The speaker's voice and its content cannot be recognized effectively, thus affecting safety production. To address these challenges, we propose a Dilated Frequency Convolutional Recurrent Network (DFCRN) Model. Specifically, DFCRN introduces, in the complex domain, a Frequency-Dilated Convolutional Encoder–Decoder (FDC-ED) to expand the frequency receptive field in a multi-scale manner, and employs stacked Complex-Dilated Feedforward Sequence Memory Networks (CD-FSMNs) to model temporal and spectral memory in parallel, thereby preserving phase information and improving the ability to recover from nonlinear distortions. Experiments conducted on the general benchmark VoiceBank+DEMAND, DNS-2020, and a newly constructed coal-mine dispatch speech dataset (Coal-single) demonstrate that DFCRN consistently and significantly outperforms representative models (e.g., GTCRN, AnyEnhance_lightweight) across both objective and subjective metrics, including CSIG/COVL, FWSEGSNR, SRMR, and several MOS-like no-reference evaluations. In particular, under real coal-mine noise conditions, DFCRN yields pronounced improvements in speech clarity and intelligibility.
-
-