The application provides a
speech enhancement method and device based on a hidden space Schrodinger bridge, equipment and a medium. The method comprises the following steps: obtaining degraded speech; encoding the degraded speech into degraded latent variables in a hidden space through a target
encoder, the target
encoder being used to converge different types of degraded speech to a distribution close to corresponding clean latent variables in the hidden space, and the target
encoder having the function of maintaining the energy characteristics of the power spectrum of an
audio signal;
processing the degraded latent variables through a target generation network based on a Schrodinger bridge to obtain clean latent variables corresponding to the degraded latent variables; and decoding the clean latent variables through a target decoder to obtain clean speech corresponding to the degraded speech. The application can solve the problem that the related art cannot realize the structured reconstruction of missing spectral information in a high-
noise and low
signal-to-
noise ratio environment, and effectively improve the
speech enhancement effect.