The invention provides a voice enhancement method, a neural network training method, devices, equipment and a medium, and relates to the technical field of voice
processing, voice environment complexity analysis is carried out to determine a
noise estimation value corresponding to
voice data, and under the condition that the
noise estimation value is greater than a target threshold value, image-assisted voice enhancement is adopted to improve the voice enhancement efficiency. Whether more
information needs to be brought through image data assistance or not can be determined in a targeted mode according to the complexity of the voice environment, voice enhancement in the complex environment is effectively achieved, and therefore the method adapts to different environment complexities, meets different voice enhancement requirements, adopts image data to carry out auxiliary enhancement on
voice data, and improves the voice enhancement efficiency. Compared with single-mode
speech enhancement, image signals insensitive to
noise are fused, the enhancement effect is better, the
environmental noise tolerance is improved, and the enhanced speech data have higher robustness, so that the calculation
power consumption and the calculation complexity of multi-mode fusion
speech enhancement are reduced while the
speech enhancement performance is guaranteed.