CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement

Leyan Yang, Dahan Wang, Xiaobin Rong, Jiadong Zhao, Jing Lu

Nanjing University | NJU-Horizon Intelligent Audio Lab, Horizon Robotics

Code ArXiv

Model Complexity Breakdown

Detailed breakdown of parameters and MACs/s for core modules in CoFi-Lite.

Module Params (k) Params (%) MACs/s (M) MACs (%)
Coarse Encoder 7.04 8.47% 1.91 14.84%
Fine Encoder 2.19 2.63% 0.65 5.05%
CPF 63.91 76.89% 4.06 31.55%
Four Inter-RNNs 2.38 2.86% 2.15 16.71%
Coarse Decoder 6.24 7.51% 1.76 13.67%
Fine Decoder 1.36 1.64% 0.38 2.95%
Other Matrix Operations 0.00 0.00% 1.96 15.23%
Total 83.12 100.00% 12.87 100.00%

Comparison with Baselines

Comparison with GTCRN and AdaptCRN on the simulated DNS3 test set.

Noisy speech sample 1 Enhanced speech by GTCRN [1] Enhanced speech by CoFi-Lite
Spectrogram Spectrogram Spectrogram
Enhanced speech by AdaptCRN [2] Enhanced speech by CoFi-Lite (Large) Clean speech
Spectrogram Spectrogram Spectrogram
Noisy speech sample 2 Enhanced speech by GTCRN [1] Enhanced speech by CoFi-Lite
Spectrogram Spectrogram Spectrogram
Enhanced speech by AdaptCRN [2] Enhanced speech by CoFi-Lite (Large) Clean speech
Spectrogram Spectrogram Spectrogram
Noisy speech sample 3 Enhanced speech by GTCRN [1] Enhanced speech by CoFi-Lite
Spectrogram Spectrogram Spectrogram
Enhanced speech by AdaptCRN [2] Enhanced speech by CoFi-Lite (Large) Clean speech
Spectrogram Spectrogram Spectrogram

Ablation Study

Results of ablation experiments on key design choices

ID1: Only coarse path is retained.

ID2: Only fine path is retained.

ID3: Coarse path + fine path, but without the CPF module.

ID4: Our proposed CoFi-Lite

We mark the frequency division (2 kHz as the boundary) with a green line and enclose the areas showing significant differences in yellow boxes.

From these three sets of examples, we can observe the following:

All speech signals enhanced by ID1 suffer from insufficient suppression of low-frequency noise. This is consistent with our claims in the Introduction and motivates the design of the fine path.

Since ID2 relies solely on the fine path, the high-frequency band remains unprocessed.

By integrating the fine path, ID3 achieves noticeable improvements in recovering the low-frequency band, yet it has not reached its optimal performance.

The addition of the CPF module in ID4 further boosts the low-frequency noise suppression. The increased global scores in our manuscript also demonstrate the CPF's effectiveness in cross-path feature fusion.

Noisy speech sample 1 Enhanced speech by ID1 Enhanced speech by ID2
Spectrogram Spectrogram Spectrogram
Enhanced speech by ID3 Enhanced speech by ID4 Clean speech
Spectrogram Spectrogram Spectrogram
Noisy speech sample 2 Enhanced speech by ID1 Enhanced speech by ID2
Spectrogram Spectrogram Spectrogram
Enhanced speech by ID3 Enhanced speech by ID4 Clean speech
Spectrogram Spectrogram Spectrogram
Noisy speech sample 3 Enhanced speech by ID1 Enhanced speech by ID2
Spectrogram Spectrogram Spectrogram
Enhanced speech by ID3 Enhanced speech by ID4 Clean speech
Spectrogram Spectrogram Spectrogram

[1] X. Rong, T. Sun, X. Zhang, Y. Hu, C. Zhu and J. Lu, "GTCRN: A Speech Enhancement Model Requiring Ultralow Computational Resources," ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, Republic of, 2024, pp. 971-975, doi: 10.1109/ICASSP48485.2024.10448310.

[2] D. Wang, X. Rong, S. Sun, Y. Hu, C. Zhu and J. Lu, "Adaptive Convolution for CNN-Based Speech Enhancement Models," in IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 4400-4413, 2025, doi: 10.1109/TASLPRO.2025.3623897.