The invention discloses a multi-
modal speech enhancement method and device based on a
deep learning model, and relates to the technical field of
artificial intelligence, and the method comprises the steps: obtaining the
head posture data, binaural audio signals and visual context information of a user in a
virtual reality environment, coding the binaural audio signals into three-dimensional space acoustic features, and carrying out the coding of the three-dimensional space acoustic features; and extracting virtual sound source position features and lip motion features from the visual context information, inputting the features into an immersive fusion enhancement network, selectively enhancing or inhibiting acoustic features from different spatial directions, generating an enhanced audio
stream, and outputting the enhanced audio
stream through a binaural rendering engine. According to the method, the technical problems that in the prior art, due to the fact that multi-
modal prior information cannot be effectively fused, voice enhancement lacks spatial selectivity, and an interference sound source irrelevant to vision is difficult to restrain are solved, and
head posture dynamic attention and lip motion cross-
modal constraint are fused; the technical effects of natural binding of auditory attention and a visual focus and effective suppression of an interference sound source are achieved.