We present DriftSE: Speech Enhancement with Generative Drifting, a novel generative framework that formulates speech enhancement as a distributional equilibrium problem. Rather than relying on iterative sampling or trajectory-based generation, DriftSE natively achieves one-step inference (1 NFE) by evolving the pushforward distribution of a mapping function to directly match the clean speech distribution via a Drifting Field. We introduce Dual Latent Drifting, combining speech semantic encoders (WavLM, HuBERT, DistilHuBERT) with acoustic encoders (BEATs, PANNs) and WavCube-pro for joint semantic-acoustic latents to provide rich, complementary training signals capturing phonetic, acoustic, and structural details. DriftSE natively supports training on fully unpaired noisy/clean speech data, enabling cross-dataset and cross-gender generalization without paired supervision. Extensive experiments on EARS-WHAM (speech denoising) and EARS-REVERB (speech dereverberation) test sets demonstrate that DriftSE achieves state-of-the-art performance across causal and non-causal architectures, outperforming multi-step diffusion and other strong baselines in a single step.
Noisy Input
Clean Reference
SGMSE+(30 steps)
ROSE-CD
DriftSE (TF-GridNet, Non-causal, WavCube)
DriftSE (TF-GridNet, Causal, WavCube)
DriftSE (TF-GridNet, Non-causal, WavLM+PANNs)
DriftSE (TF-GridNet, Causal, WavLM+PANNs)
Noisy Input
Clean Reference
SGMSE+(30 steps)
ROSE-CD
DriftSE (TF-GridNet, Non-causal, WavCube)
DriftSE (TF-GridNet, Causal, WavCube)
DriftSE (TF-GridNet, Non-causal, WavLM+PANNs)
DriftSE (TF-GridNet, Causal, WavLM+PANNs)
Noisy Input
Clean Reference
SGMSE+(30 steps)
ROSE-CD
DriftSE (TF-GridNet, Non-causal, WavCube)
DriftSE (TF-GridNet, Causal, WavCube)
DriftSE (TF-GridNet, Non-causal, WavLM+PANNs)
DriftSE (TF-GridNet, Causal, WavLM+PANNs)
Noisy Input
DriftSE (paired)
DriftSE (unpaired, map to DNS)
Noisy Input
DriftSE (paired)
DriftSE (unpaired, map to DNS)
Noisy Input
DriftSE (paired)
DriftSE (unpaired, map to DNS)
@inproceedings{xu2026driftse,
author = {Liang Xu and Diego Caviedes-Nozal and W. Bastiaan Kleijn and Longfei Felix Yan and Rasmus Kongsgaard Olsson},
title = {Speech Enhancement Based on Drifting Models},
booktitle = {Proc. Interspeech 2026},
year = {2026}
}
@article{xu2026driftsespeechenhancementgenerative,
author = {Xu, Liang and Caviedes-Nozal, Diego and Kleijn, W. Bastiaan and Yan, Longfei Felix and Olsson, Rasmus Kongsgaard},
title = {Speech Enhancement with Generative Drifting},
journal = {IEEE/ACM Transactions on Audio, Speech, and Language Processing},
year = {2026},
note = {Submitted},
}