DriftSE: Speech Enhancement with Generative Drifting

Submitted to IEEE/ACM TASLP
1School of Engineering and Computer Science, Victoria University of Wellington, New Zealand
2GN Audio A/S, Denmark

We present DriftSE: Speech Enhancement with Generative Drifting, a novel generative framework that formulates speech enhancement as a distributional equilibrium problem. Rather than relying on iterative sampling or trajectory-based generation, DriftSE natively achieves one-step inference (1 NFE) by evolving the pushforward distribution of a mapping function to directly match the clean speech distribution via a Drifting Field. We introduce Dual Latent Drifting, combining speech semantic encoders (WavLM, HuBERT, DistilHuBERT) with acoustic encoders (BEATs, PANNs) and WavCube-pro for joint semantic-acoustic latents to provide rich, complementary training signals capturing phonetic, acoustic, and structural details. DriftSE natively supports training on fully unpaired noisy/clean speech data, enabling cross-dataset and cross-gender generalization without paired supervision. Extensive experiments on EARS-WHAM (speech denoising) and EARS-REVERB (speech dereverberation) test sets demonstrate that DriftSE achieves state-of-the-art performance across causal and non-causal architectures, outperforming multi-step diffusion and other strong baselines in a single step.

Overview of the proposed DriftSE

Overview of the proposed dual-latent DriftSE framework.

Summary of Results Using DriftSE

Evaluation metrics on EARS-WHAM and EARS-REVERB. Causal indicates a causal architecture, Para denotes parameters in millions, and GMACs denotes total operations (per-step × NFE). Best within a group in bold.
Utterance-level PCA visualizations of DriftSE and DriftSEU for denoising (left) and dereverberation (right). Both models align closely in the acoustic PANNs space, but the unpaired variant diverges significantly in the semantic WavLM space. ★ denotes cluster centroids, and |Δμ| quantifies the distance between generated and clean centroids.

EARS-REVERB (speech dereverberation) Samples

Sample 1

Noisy Input

Clean Reference

SGMSE+(30 steps)

ROSE-CD

DriftSE (TF-GridNet, Non-causal, WavCube)

DriftSE (TF-GridNet, Causal, WavCube)

DriftSE (TF-GridNet, Non-causal, WavLM+PANNs)

DriftSE (TF-GridNet, Causal, WavLM+PANNs)

Sample 2

Noisy Input

Clean Reference

SGMSE+(30 steps)

ROSE-CD

DriftSE (TF-GridNet, Non-causal, WavCube)

DriftSE (TF-GridNet, Causal, WavCube)

DriftSE (TF-GridNet, Non-causal, WavLM+PANNs)

DriftSE (TF-GridNet, Causal, WavLM+PANNs)

Sample 3

Noisy Input

Clean Reference

SGMSE+(30 steps)

ROSE-CD

DriftSE (TF-GridNet, Non-causal, WavCube)

DriftSE (TF-GridNet, Causal, WavCube)

DriftSE (TF-GridNet, Non-causal, WavLM+PANNs)

DriftSE (TF-GridNet, Causal, WavLM+PANNs)

Unpaired Samples

Sample 1

Noisy Input

DriftSE (paired)

DriftSE (unpaired, map to DNS)

Sample 2

Noisy Input

DriftSE (paired)

DriftSE (unpaired, map to DNS)

Sample 3

Noisy Input

DriftSE (paired)

DriftSE (unpaired, map to DNS)

BibTeX

@inproceedings{xu2026driftse,
  author    = {Liang Xu and Diego Caviedes-Nozal and W. Bastiaan Kleijn and Longfei Felix Yan and Rasmus Kongsgaard Olsson},
  title     = {Speech Enhancement Based on Drifting Models},
  booktitle = {Proc. Interspeech 2026},
  year      = {2026}
}

@article{xu2026driftsespeechenhancementgenerative,
  author  = {Xu, Liang and Caviedes-Nozal, Diego and Kleijn, W. Bastiaan and Yan, Longfei Felix and Olsson, Rasmus Kongsgaard},
  title   = {Speech Enhancement with Generative Drifting},
  journal = {IEEE/ACM Transactions on Audio, Speech, and Language Processing},
  year    = {2026},
  note    = {Submitted},
}