Model Complexity Breakdown
Detailed breakdown of parameters and MACs/s for core modules in CoFi-Lite.
| Module | Params (k) | Params (%) | MACs/s (M) | MACs (%) |
|---|---|---|---|---|
| Coarse Encoder | 7.04 | 8.47% | 1.91 | 14.84% |
| Fine Encoder | 2.19 | 2.63% | 0.65 | 5.05% |
| CPF | 63.91 | 76.89% | 4.06 | 31.55% |
| Four Inter-RNNs | 2.38 | 2.86% | 2.15 | 16.71% |
| Coarse Decoder | 6.24 | 7.51% | 1.76 | 13.67% |
| Fine Decoder | 1.36 | 1.64% | 0.38 | 2.95% |
| Other Matrix Operations | 0.00 | 0.00% | 1.96 | 15.23% |
| Total | 83.12 | 100.00% | 12.87 | 100.00% |
Comparison with Baselines
Comparison with GTCRN and AdaptCRN on the simulated DNS3 test set.
| Noisy speech sample 1 | Enhanced speech by GTCRN [1] | Enhanced speech by CoFi-Lite |
|---|---|---|
|
|
|
| Enhanced speech by AdaptCRN [2] | Enhanced speech by CoFi-Lite (Large) | Clean speech |
|
|
|
| Noisy speech sample 2 | Enhanced speech by GTCRN [1] | Enhanced speech by CoFi-Lite |
|
|
|
| Enhanced speech by AdaptCRN [2] | Enhanced speech by CoFi-Lite (Large) | Clean speech |
|
|
|
| Noisy speech sample 3 | Enhanced speech by GTCRN [1] | Enhanced speech by CoFi-Lite |
|
|
|
| Enhanced speech by AdaptCRN [2] | Enhanced speech by CoFi-Lite (Large) | Clean speech |
|
|
|
Ablation Study
Results of ablation experiments on key design choices
ID1: Only coarse path is retained.
ID2: Only fine path is retained.
ID3: Coarse path + fine path, but without the CPF module.
ID4: Our proposed CoFi-Lite
We mark the frequency division (2 kHz as the boundary) with a green line and enclose the areas showing significant differences in yellow boxes.
From these three sets of examples, we can observe the following:
All speech signals enhanced by ID1 suffer from insufficient suppression of low-frequency noise. This is consistent with our claims in the Introduction and motivates the design of the fine path.
Since ID2 relies solely on the fine path, the high-frequency band remains unprocessed.
By integrating the fine path, ID3 achieves noticeable improvements in recovering the low-frequency band, yet it has not reached its optimal performance.
The addition of the CPF module in ID4 further boosts the low-frequency noise suppression. The increased global scores in our manuscript also demonstrate the CPF's effectiveness in cross-path feature fusion.
| Noisy speech sample 1 | Enhanced speech by ID1 | Enhanced speech by ID2 |
|---|---|---|
|
|
|
| Enhanced speech by ID3 | Enhanced speech by ID4 | Clean speech |
|
|
|
| Noisy speech sample 2 | Enhanced speech by ID1 | Enhanced speech by ID2 |
|
|
|
| Enhanced speech by ID3 | Enhanced speech by ID4 | Clean speech |
|
|
|
| Noisy speech sample 3 | Enhanced speech by ID1 | Enhanced speech by ID2 |
|
|
|
| Enhanced speech by ID3 | Enhanced speech by ID4 | Clean speech |
|
|
|
[1] X. Rong, T. Sun, X. Zhang, Y. Hu, C. Zhu and J. Lu, "GTCRN: A Speech Enhancement Model Requiring Ultralow Computational Resources," ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, Republic of, 2024, pp. 971-975, doi: 10.1109/ICASSP48485.2024.10448310.
[2] D. Wang, X. Rong, S. Sun, Y. Hu, C. Zhu and J. Lu, "Adaptive Convolution for CNN-Based Speech Enhancement Models," in IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 4400-4413, 2025, doi: 10.1109/TASLPRO.2025.3623897.