Multi-modal speech enhancement method and device based on deep learning model
By acquiring multimodal information in a virtual reality environment, using deep learning models for spatial audio coding and visual context coding, and combining head posture and lip movement features, an enhanced audio stream with spatial selectivity and consistency is generated. This solves the problems of spatial selectivity and interference source suppression in speech enhancement in virtual reality, and improves the effect of speech separation and enhancement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHUTIAN (HANGZHOU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies cannot effectively integrate multimodal prior information, resulting in speech enhancement lacking spatial selectivity and difficulty in suppressing visually irrelevant interfering sound sources, especially in virtual reality environments where speech separation and enhancement performance is limited.
By acquiring head pose data, binaural audio signals, and visual context information, an immersive fusion enhancement network is generated using the spatial audio coding and visual context coding modules of a deep learning model. Cross-modal consistency constraints are applied by combining head pose and lip movement features to generate an enhanced audio stream with spatial orientation awareness.
It achieves adaptive binding of the user's auditory attention direction and visual focus in the virtual reality environment, effectively suppresses interfering sound sources, generates an enhanced audio stream with spatial selectivity and environmental acoustic consistency, and improves the immersion and accuracy of speech enhancement.
Smart Images

Figure CN121963735A_ABST
Abstract
Description
Multimodal speech enhancement methods and devices based on deep learning models Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and specifically to a multimodal speech enhancement method and device based on a deep learning model. Background Technology
[0002] With the rapid development of virtual reality technology, immersive voice interaction has become a key requirement for improving user experience. In virtual reality environments, users often face problems such as multiple speakers speaking simultaneously, environmental noise interference, and difficulties in locating spatial sound sources. Traditional single-modal speech enhancement methods rely solely on audio signal processing, making it difficult to effectively utilize the rich spatial and visual information in virtual reality scenes, resulting in limited speech separation and enhancement performance in complex virtual scenes. Summary of the Invention
[0003] This application provides a multimodal speech enhancement method and device based on a deep learning model, which solves the technical problem that existing technologies cannot effectively integrate multimodal prior information, resulting in speech enhancement lacking spatial selectivity and difficulty in suppressing visually irrelevant interference sources.
[0004] The first aspect of this application provides a multimodal speech enhancement method based on a deep learning model. The method includes: acquiring head pose data of a target user, binaural audio signals, and visual context information in a virtual reality environment, wherein the visual context information includes at least spatial location information of sound sources in the virtual scene and facial animation data of a virtual speaker; encoding the binaural audio signals into acoustic features containing three-dimensional spatial orientation information using a spatial audio encoding module; extracting virtual sound source location features and virtual speaker lip movement features from the visual context information using a visual context encoding module; inputting the head pose data, the acoustic features, the virtual sound source location features, and the virtual speaker lip movement features into an immersive fusion enhancement network; the immersive fusion enhancement network dynamically adjusts attention weights based on the head pose data, selectively enhancing or suppressing acoustic features from different spatial directions, and combining the virtual speaker lip movement features with cross-modal consistency constraints to generate an enhanced audio stream with spatial orientation awareness; and outputting the enhanced audio stream through a binaural rendering engine to provide immersive speech interaction for the user.
[0005] A second aspect of this application provides an electronic device comprising: a processor coupled to a memory for storing a program that, when executed by the processor, implements the method described in any of the first aspects.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: The multimodal speech enhancement method and device based on deep learning models provided in this application relate to the field of artificial intelligence technology. By utilizing multimodal information such as head posture, sound source location, and speaker facial animation that can be accurately obtained from virtual reality scenes, a dynamic spatial attention mechanism guided by head posture and cross-modal consistency discrimination based on lip movement constraints are constructed. This enables adaptive binding of the user's auditory attention direction and visual focus, generating an enhanced audio stream with spatial orientation perception and environmental acoustic consistency. This solves the technical problem that existing technologies cannot effectively integrate multimodal prior information, resulting in speech enhancement lacking spatial selectivity and difficulty in suppressing visually irrelevant interference sources. It achieves the technical effect of integrating dynamic head posture attention and cross-modal constraints on lip movement, realizing natural binding of auditory attention and visual focus, and effective suppression of interference sources. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 is a schematic flowchart of a multimodal speech enhancement method based on a deep learning model provided in an embodiment of this application; Figure 2 is a schematic diagram of the structure of an electronic device provided in this application.
[0009] Explanation of reference numerals in the attached drawings: Electronic device 300, memory 301, processor 302, communication interface 303, bus architecture 304. Detailed Implementation
[0010] This application provides a multimodal speech enhancement method and device based on a deep learning model, which solves the technical problem that existing technologies cannot effectively integrate multimodal prior information, resulting in speech enhancement lacking spatial selectivity and difficulty in suppressing visually irrelevant interference sources.
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0012] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0013] Example 1, as shown in Figure 1, provides a multimodal speech enhancement method based on a deep learning model. The method includes: P10: In a virtual reality environment, acquiring the head posture data of the target user, binaural audio signals, and visual context information in the virtual reality environment, wherein the visual context information includes at least the spatial location information of the sound source in the virtual scene and the facial animation data of the virtual speaker.
[0014] Specifically, the first step is to collect various types of data within a virtual reality (VR) environment to ensure an accurate voice enhancement experience. This data includes the target user's head posture data, binaural audio signals, and visual contextual information within the VR environment.
[0015] In this process, the head-mounted display (HMD) or other inertial measurement unit (IMU) device worn by the user captures the user's head posture in real time. IMU devices typically include sensors such as accelerometers and gyroscopes, providing precise data on the user's head rotation and tilt. This data not only reflects the user's position in the virtual scene but also displays the user's gaze direction. For example, the inertial measurement unit outputs raw data of the user's three-axis angular velocity and three-axis linear acceleration in real time at a sampling frequency of 1000 Hz. Simultaneously, the head-mounted display's built-in camera array acquires image sequences of the surrounding environment at a frame rate of 90 Hz. Visual odometry is calculated based on simultaneous localization and mapping (SLT) algorithms to estimate the head-mounted display's pose change relative to the virtual world coordinate system. Then, an extended Kalman filter algorithm is used to fuse these two types of sensor data. Visual observation is used to correct the cumulative drift of the inertial data, and the inertial data is used to compensate for rapid motion blur in visual tracking. Finally, filtered and optimized six-DOF pose data is output. Using this information, the spatial orientation of the audio can be dynamically adjusted to synchronize with the user's head movements in real time. When the user turns their head or changes their perspective, the system will adjust the directionality of the audio accordingly, thereby enhancing immersion and realism and ensuring that the audio signal can reflect the correct orientation in three-dimensional space.
[0016] Binaural audio signals refer to audio data received through binaural headphones or headsets, possessing extremely high time and frequency resolution. In a virtual reality environment, binaural audio signals can provide spatial audio, meaning the direction and location of sound sources can be accurately conveyed to the user. For example, omnidirectional condenser microphone arrays with a spacing of 18 cm are symmetrically positioned on both earcups of a head-mounted display. The microphones have a frequency response range covering 20 Hz to 20000 Hz, and simultaneously acquire ambient sound waves at sampling rates of 16000 Hz or 48000 Hz and bit depths of 16 bits or 24 bits, outputting time-domain waveform sequences in pulse-code modulation format for both the left and right channels. Through this process, the user can not only hear audio signals from different directions but also perceive the spatial distance and direction of these sounds.
[0017] Visual contextual information primarily includes the location of sound sources in the virtual scene and facial animation data of virtual characters. Sound source location features are calculated by the virtual reality engine within the virtual scene, representing the position of each sound source in three-dimensional space. This position can be represented using a three-dimensional coordinate system, such as the position of virtual characters or the distribution of environmental sound sources. This location data is crucial for ensuring the accuracy of spatial audio, especially in dynamic interactive scenes where the location of sound sources may change over time and therefore must be continuously updated.
[0018] Meanwhile, the virtual speaker's facial animation data is generated using facial tracking technology or computer graphics algorithms. The virtual speaker's lip movements and facial expressions are synchronized with its spoken words. This data provides users with a more natural and vivid visual experience, especially during voice interaction, where users can see the virtual character's lip movements and expressions perfectly matching its voice. Facial animation data also provides important basis for subsequent cross-modal consistency constraints, ensuring consistency between audio and visual signals, thereby further enhancing immersion.
[0019] For example, visual context information is extracted in real time by querying the application programming interface of the virtual reality scene management system. Within each rendering frame cycle, a list of acoustic objects in the currently active scene is first obtained. Each sound source object is traversed, and its 3D transformation components are extracted. The 3D position coordinates of the sound source in the world coordinate system are deconstructed from the transformation matrix from the local coordinate system to the world coordinate system. Then, sound source objects tagged as speaker types are identified, and their bound virtual character models are accessed to obtain facial animation state data. This data covers basic action units describing lip movements, such as jaw opening and closing, lip pursing, and corners of the mouth turning up. The extracted facial animation data establishes a frame-level mapping relationship with the audio data through timestamps.
[0020] In summary, the head pose data, binaural audio signals, and visual context information collected through the above steps provide comprehensive data support for the realization of speech enhancement and immersive interaction. This data will be fed into subsequent deep learning models for spatial audio enhancement, cross-modal fusion, and other processing.
[0021] P20: The binaural audio signal is encoded into acoustic features containing three-dimensional spatial orientation information through the spatial audio encoding module.
[0022] Furthermore, step P20 in this embodiment of the application also includes: P21: performing a short-time Fourier transform on the binaural audio signal using a spatial audio coding module to obtain a complex spectrum; P22: extracting orientation-sensitive features associated with the head-related transfer function from the complex spectrum; P23: spatially aligning the orientation-sensitive features with the virtual sound source location features using a spatial attention layer to generate a spatial weight matrix; P24: weighting and adjusting the complex spectrum according to the spatial weight matrix to obtain acoustic features containing three-dimensional spatial orientation information.
[0023] It should be understood that the spatial audio coding module of this application is responsible for encoding binaural audio signals into acoustic features containing three-dimensional spatial orientation information, providing a foundation for subsequent speech enhancement and immersive audio processing.
[0024] First, a short-time Fourier transform is performed on the input binaural audio signal to convert the time-domain waveform signal into a complex frequency spectrum representation in the time-frequency domain. For example, the left and right channel audio signals are divided into frames with a preset frame length, such as 20 to 40 milliseconds. A 50% overlap rate is set between adjacent frames to ensure spectral continuity. A fast Fourier transform is performed on each frame, outputting the complex frequency values corresponding to each frame. The real part represents the amplitude of the cosine component of the signal, and the imaginary part represents the amplitude of the sine component. This yields a binaural complex spectrum, the dimension of which is the number of time frames multiplied by the number of frequency points multiplied by the number of channels.
[0025] Then, location-sensitive features associated with the head-related transfer function (HRTF) are extracted from the complex spectrum. HRTF is a mathematical model describing the transmission characteristics of sound from the sound source to the ear, especially for different orientations in three-dimensional space. By analyzing the relationship between binaural audio signals and HRTF, spatial location-related features can be extracted, especially the directionality of the sound source and the sense of spatial distance. For example, this extraction process first calculates the cross-power spectral density between the complex spectra of the left and right channels for each time frame and each frequency point based on binaural cue analysis, and then extracts binaural time difference features and binaural intensity difference features. Among them, the binaural time difference feature is obtained by calculating the phase angle of the cross-power spectrum, which represents the time delay of the sound wave reaching the ears; the binaural intensity difference feature is obtained by the logarithmic operation of the ratio of the amplitude spectra of the left and right channels, which represents the difference in energy attenuation of the sound wave reaching the ears. Next, the aforementioned binaural cues are pattern matched with a pre-stored head-related transfer function database to map each time-frequency unit of the spectrum to the corresponding spatial orientation category, generating an orientation-sensitive feature map. This feature map can more accurately capture the location of sound sources in the virtual scene, thereby achieving a more realistic spatial audio experience.
[0026] Subsequently, based on this, the spatial audio coding module uses a spatial attention layer to spatially align the orientation-sensitive features with the spatial location features of the virtual sound source. For example, the spatial attention layer receives two sets of input features: an orientation-sensitive feature map and virtual sound source location features from the visual context coding module. Next, the virtual sound source location features are projected into the same dimensional space as the orientation-sensitive feature map. A query vector is generated through a learnable linear transformation. The orientation-sensitive feature map is then transformed into a key vector and a value vector through another set of linear transformations. The dot product similarity between the query vector and the key vector is calculated. After scaling factor normalization and Softmax function activation, a spatial weight matrix representing the degree of spatial association between each time-frequency unit and each virtual sound source is obtained. The dimension of this matrix is adapted to the number of time frames and frequency points, and the matrix element values range from zero to one. A larger value indicates a higher probability that the acoustic energy of the corresponding time-frequency unit originates from that spatial orientation. This process ensures that the spatial localization of the audio features can accurately match the real location of the virtual sound source, thereby guaranteeing the orientation and localization of the sound source.
[0027] Finally, the complex spectrum is weighted and adjusted according to the spatial weight matrix to obtain acoustic features containing three-dimensional spatial orientation information. Specifically, the weighting method involves element-wise multiplication of the spatial weight matrix with the original complex spectrum, enhancing time-frequency components consistent with the spatial location of the virtual sound source and suppressing noise components deviating from the virtual sound source's spatial location. The weighted complex spectrum can be selectively subjected to inverse short-time Fourier transform to restore the enhanced time-domain signal, or it can be directly output to the downstream network module while maintaining its frequency domain form. The resulting acoustic features retain both the spectral content information of the original audio and the three-dimensional spatial orientation information of the sound source relative to the user's head, providing a feature foundation for subsequent speech enhancement.
[0028] P30: The virtual sound source location features and virtual speaker lip movement features are extracted from the visual context information through the visual context encoding module.
[0029] Furthermore, step P30 in this embodiment of the application also includes: P31: parsing the three-dimensional coordinates of the virtual sound source from the visual context information and converting them into spherical coordinate system parameters relative to the user's head to obtain the virtual sound source position features; P32: inputting the vertex coordinates of the lip region in the facial animation data into a three-dimensional convolutional neural network to output a lip vertex displacement sequence; P33: performing dynamic time warping on the lip vertex displacement sequence to generate virtual speaker lip movement features aligned with the complex spectrum time sequence.
[0030] Optionally, the core task of the visual context coding module is to extract features closely related to speech enhancement from the visual context information in the virtual reality environment, including virtual sound source location features and virtual speaker lip movement features.
[0031] First, the virtual sound source location features are extracted. The three-dimensional Cartesian coordinates of each virtual sound source in the world coordinate system are parsed from the input visual context information. These coordinates, in meters, represent the absolute spatial position of the sound source in the virtual scene. Then, the three-dimensional position coordinates and three-dimensional rotation vector of the user's head in the current frame are obtained. The relative position vector of each sound source relative to the user's head is calculated. This calculation is achieved by subtracting the sound source coordinates from the head coordinates in the world coordinate system and performing an inverse transformation using the head rotation matrix. The resulting relative position vectors are then transformed from the Cartesian coordinate system to the spherical coordinate system, and three parameters—azimuth, elevation, and distance—are extracted. The azimuth is defined as the angle between the projection of the relative position vector onto the horizontal plane and the direction directly in front, ranging from 0 to 360 degrees. The elevation is defined as the angle between the relative position vector and the horizontal plane, ranging from -90 to +90 degrees. The distance is defined as the magnitude of the relative position vector. Finally, the three parameters are concatenated into a three-dimensional vector, which is then mapped to a specified feature dimension through a learnable linear layer to obtain the virtual sound source location features. These features are represented by fixed-length real-valued vectors that characterize the spatial orientation of each sound source relative to the user's head. The purpose of this process is to ensure that the directionality of the sound aligns with the user's visual focus, thereby providing a more realistic and immersive auditory experience.
[0032] Next, the extraction of lip movement features of the virtual speaker is performed. The vertex coordinates of the lip region are extracted from the virtual speaker's facial animation data. Virtual speaker facial animations are typically generated by facial tracking systems or computer graphics algorithms, where lip movement is particularly crucial. This lip movement data is represented as multiple vertex positions in a three-dimensional coordinate system, reflecting the dynamic changes of the virtual character's lips, including opening and closing, vertical movement, and horizontal movement. To further process this lip movement data, these vertex coordinates are input into a three-dimensional convolutional neural network (3DCNN). The 3DCNN contains multiple spatiotemporal convolutional layers, each configured with a three-dimensional convolutional kernel to simultaneously capture the spatial structural features of the lip shape and the temporal dynamic features of vertex movement. The network ends with a global average pooling layer and a fully connected layer to output a fixed-length sequence of lip vertex displacements, which characterizes the degree of lip deformation relative to a neutral pose at each time point.
[0033] Subsequently, dynamic time warping is performed on the lip vertex displacement sequence to generate virtual speaker lip motion features aligned with the complex spectrum time sequence. Since the sampling frequency of visual context information is limited by the rendering frame rate of the virtual reality system, typically between 72 Hz and 120 Hz, while the temporal resolution of the complex spectrum is determined by the frame shift parameter of the short-time Fourier transform, typically 10 milliseconds corresponding to 100 Hz, a difference in sampling frequency exists. Therefore, a dynamic time warping algorithm is needed to establish a mapping relationship between visual and audio frames by calculating the optimal matching path between the lip vertex displacement sequence and the complex spectrum time axis. Linear interpolation or spline interpolation is then performed on the lip vertex displacement sequence for resampling to ensure its temporal resolution is consistent with the complex spectrum. Once the audio and visual signals are temporally aligned, the generated virtual speaker lip motion features can be used as part of subsequent processing, along with other audio features, for further speech enhancement and cross-modal consistency constraints.
[0034] These features not only help achieve more accurate spatial audio enhancement, but also provide strong support for the voice and visual synchronization of virtual characters, ensuring accurate matching of audio and visuals in virtual reality scenes.
[0035] P40: Input the head pose data, the acoustic features, the virtual sound source location features, and the virtual speaker lip movement features into the immersive fusion enhancement network.
[0036] Furthermore, the immersive fusion enhancement network includes: a spatial attention subnetwork for receiving the head pose data and the virtual sound source location features, and outputting orientation selection weights; a lip movement-speech alignment subnetwork for receiving the virtual speaker's lip movement features and the acoustic features, and outputting a cross-modal consistency score; and a fusion decision layer for combining the orientation selection weights and the cross-modal consistency score to generate a speech enhancement mask, and modulating the acoustic features based on the speech enhancement mask to output the enhanced audio stream.
[0037] Specifically, the immersive fusion enhancement network adopts a multi-branch parallel architecture, which includes three functional modules: a spatial attention subnetwork, a lip movement-speech alignment subnetwork, and a fusion decision layer. Each module works together to achieve spatial selective auditory attention based on head posture guidance and cross-modal speech enhancement based on lip movement constraints.
[0038] The spatial attention subnetwork receives head pose data and virtual sound source location features, and outputs azimuth selection weights. This subnetwork first performs spatial relationship calculations on the 3D rotation vector from the head pose data and the azimuth and elevation parameters from the virtual sound source location features. Using spherical distance or Euclidean distance formulas, it calculates the viewing angle offset of each virtual sound source relative to the user's current gaze direction. Then, it concatenates the viewing angle offset with the virtual sound source location features and inputs it into a multilayer perceptron. The multilayer perceptron contains two hidden layers, each configured with a batch normalization layer and an activation function. The output layer uses a Softmax activation function to generate azimuth selection weights in the form of a probability distribution. The dimension of this weight vector is equal to the number of virtual sound sources, and the value of each element ranges from zero to one, with a total sum of one. A larger element value indicates a higher probability that the sound source in the corresponding direction is actively focused on by the user or needs to be enhanced. Simultaneously, the spatial attention subnetwork adaptively adjusts the sharpness of the azimuth selection curve through learnable parameters, enabling flexible switching from wide beam to narrow beam.
[0039] The lip movement-speech alignment subnetwork receives the lip movement features and acoustic features of a virtual speaker and outputs a cross-modal consistency score. This subnetwork employs a dual-tower encoder structure. The first tower is a visual encoder, consisting of a cascaded one-dimensional temporal convolutional layer and bidirectional gated recurrent units, which input the virtual speaker's lip movement features and output a visual embedding vector sequence. The second tower is an auditory encoder, consisting of a cascaded two-dimensional convolutional layer and bidirectional gated recurrent units, which input the amplitude or power spectrum of the acoustic features and output an auditory embedding vector sequence. The output sequences of the two encoders are linearly projected onto the same dimensional space, and frame-by-frame cosine similarity or dot product similarity is calculated to obtain the initial alignment score. Next, a contrastive learning mechanism is introduced, using visual and auditory sample pairs from the same speaker as positive samples and sample pairs from different speakers or silent periods as negative samples. The encoder is trained using the InfoNCE loss function to extract modality-independent speech content representations. Finally, the cross-modal consistency score for each time frame is output. This score is a scalar in the range of zero to one, which represents the degree of matching between the current audio content and the visual content of lip movement. The higher the score, the more likely the audio of that frame is to come from a virtual character who is speaking.
[0040] The fusion decision layer combines orientation selection weights and cross-modal consistency scores to generate a speech enhancement mask. Based on this mask, acoustic features are modulated to output an enhanced audio stream. For example, the orientation selection weights are first extended to the time-frequency dimension to match the number of time frames and frequency points of the acoustic features, forming a spatial weight tensor. Simultaneously, the cross-modal consistency score is extended to the frequency dimension, forming a consistency weight tensor. Then, a learnable fusion strategy, such as weighted summation, gating, or attention mechanisms, is used to integrate the spatial weight tensor and the consistency weight tensor into a comprehensive weight tensor. This comprehensive weight tensor is then mapped using a sigmoid activation function to generate a speech enhancement mask. This mask has the same dimension as the complex spectrum of the acoustic features, with element values ranging from zero to one. The speech enhancement mask is then multiplied element-wise with the acoustic features to enhance speech components that are close to the target orientation and are audiovisually consistent, while suppressing noise components that are off-target or audiovisually inconsistent. Finally, the masked acoustic features are restored to time-domain waveforms through inverse short-time Fourier transform, or directly output as frequency-domain feature streams as enhanced audio streams with spatial awareness, cross-modal consistency, and immersion.
[0041] P50: The immersive fusion enhancement network dynamically adjusts attention weights based on the head pose data, selectively enhances or suppresses acoustic features from different spatial directions, and combines the virtual speaker's lip movement features to perform cross-modal consistency constraints, generating an enhanced audio stream with spatial orientation awareness.
[0042] Furthermore, selective enhancement or suppression of acoustic features from different spatial directions is performed. Step P50 in this embodiment further includes: P51: calculating the user's auditory focus direction in real time based on the head posture data; P52: performing spatial matching calculation between the virtual sound source location features and the auditory focus direction to generate a directional weight mask; P53: using the directional weight mask to weight the acoustic features, so that the speech components corresponding to the sound source spatially aligned with the auditory focus direction are enhanced, while the sound source components from other directions are suppressed; wherein, the calculation of the directional weight mask integrates a distance attenuation model and a spatial reverberation model to simulate the characteristics of a real three-dimensional sound field.
[0043] Optionally, the immersive fusion enhancement network first calculates the user's auditory focus direction in real time based on head posture data. Head posture data reflects the user's head movements, including rotation and tilt angles, and this data is used to determine the user's current auditory focus direction. The auditory focus direction can be understood as the user's preferred direction for perceiving sound, which changes as the user's head moves. For example, a three-dimensional rotation vector is extracted from the input head posture data, converted into a rotation matrix or quaternion representation, and then the mapping direction of the forward direction vector in the user's head coordinate system to the world coordinate system is calculated. This forward direction vector is the auditory focus direction, representing the user's current gaze orientation or auditory attention focus in the virtual environment. Furthermore, eye-tracking data can be used to fine-tune the auditory focus direction to reflect the user's actual visual-auditory joint attention area.
[0044] Next, the virtual sound source location features are spatially matched with the auditory focus direction to generate a directional weight mask. The virtual sound source location features represent the sound source's position in three-dimensional space, while the auditory focus direction reflects the user's current point of focus. By matching these two, the positional relationship of each sound source relative to the user's auditory focus direction can be calculated, generating a weight mask that reflects the relative importance of sound sources in different spatial directions. For example, for each virtual sound source, its azimuth and elevation angles relative to the user's head are calculated. Simultaneously, the spatial angle between the sound source direction and the auditory focus direction is calculated. This angle is obtained through the spherical cosine theorem or vector dot product operations, with a value ranging from 0 to 180 degrees. Then, the spatial angle is input into a preset directional gain function, which uses a Gaussian function, a cosine function, or a learned nonlinear mapping, such that the gain is maximized when the angle is 0 degrees and monotonically decreases as the angle increases. Simultaneously, a distance attenuation model is fused. This model simulates the effect of the distance between the sound source and the user on sound intensity, calculated based on the distance parameter in the virtual sound source location features. The sound wave propagation energy attenuation coefficient follows the inverse square law or the logarithmic distance law, with the attenuation coefficient decreasing as the sound source moves further away. Furthermore, a spatial reverberation model is integrated. This model simulates the reflection and scattering of sound in space, calculating the proportion of early reflection and late reverberation energy based on the geometric and material properties of the virtual scene, and applying differentiated reverberation characteristics to sound sources at different distances and orientations. Finally, by combining the spatial angle gain, distance attenuation coefficient, and reverberation characteristics, a directional weight mask is generated. This mask is represented as a time-frequency matrix, with matrix elements ranging from zero to one. A larger value indicates a higher probability that the corresponding time-frequency unit originates from a sound source near the direction of the auditory focus. For example, a higher weight mask value indicates that sound sources located in the direction of the user's auditory focus should be enhanced, while a lower weight mask value indicates that the audio components of sound sources from other directions need to be suppressed.
[0045] Subsequently, directional weighted masks are used to weight the acoustic features, enhancing the speech components corresponding to sound sources spatially aligned with the auditory focus direction, while suppressing sound source components from other directions. Specifically, the directional weighted mask is multiplied element-wise with the complex spectrum or amplitude spectrum of the acoustic features. Time-frequency units with mask values close to one are preserved and amplified, while those with mask values close to zero are attenuated and filtered out. Finally, in the weighted acoustic features, the energy of the target speech signal from the auditory focus direction is enhanced, while interference noise and competing speaker signal energy from the sides and rear are suppressed, allowing the user to focus their attention on the sound source they are currently interested in.
[0046] Meanwhile, the immersive fusion enhancement network incorporates virtual speaker lip movement features for cross-modal consistency constraints. Weighted acoustic features and virtual speaker lip movement features are input into the lip-speech alignment subnet to calculate the cross-modal consistency score of the current enhancement result. If the consistency score is below a preset threshold, it indicates that the enhanced audio may be out of sync with lip movements or originate from a non-target speaker. This triggers a consistency constraint mechanism, adjusting the generation strategy of the directional weight mask or introducing additional temporal alignment loss to suppress visually silent but acoustically active interference sources and enhance visually active and acoustically matched reliable speech components. Through iterative optimization or end-to-end training, the enhanced audio stream output by the network is made temporally synchronized with lip movements while originating from the visually focused speaker, thus achieving audiovisually consistent spatially selective speech enhancement.
[0047] Furthermore, by combining the lip movement features of the virtual speaker with cross-modal consistency constraints to generate an enhanced audio stream with spatial orientation awareness, step P50 of this embodiment further includes: P54: identifying the visual identity information of multiple virtual speakers in the virtual reality environment; P55: extracting the voiceprint features of each virtual speaker from the mixed binaural audio signals; P56: performing cross-modal association matching between the visual identity information and the voiceprint features to generate an independent identity vector for each virtual speaker; P57: based on the identity vector, separating the speech stream of each target speaker from the mixed audio through a speech separation network; P58: for the separated speech stream of a single target speaker, combining its corresponding virtual speaker lip movement features, performing a speech enhancement operation to generate an enhanced audio stream with spatial orientation awareness.
[0048] In one possible embodiment of this application, the process of combining virtual speaker lip movement features with cross-modal consistency constraints to generate an enhanced audio stream with spatial orientation awareness can be further refined.
[0049] First, the visual identity information of multiple virtual speakers in the virtual reality environment is identified. This identification process is achieved by querying the metadata tags of the virtual characters or calling the identity query interface of the scene management system to obtain the unique identifier and role attribute information of each virtual speaker, such as role name and user binding relationship. Simultaneously, facial embedding features are extracted from the facial animation data of the virtual speakers. These features are generated through a pre-trained face recognition network or role classification network to characterize the visual appearance characteristics of the virtual speakers. The identifier and facial embedding features are then concatenated to form a visual identity feature vector.
[0050] Then, the voiceprint features of each virtual speaker are extracted from the mixed binaural audio signals. Binaural audio signals refer to audio data collected through binaural microphones or headphones, which typically contain the voices of multiple speakers. When multiple speakers speak simultaneously, the audio signals mix together, forming complex sound source information. Voiceprint recognition technology can extract the unique voiceprint features of each speaker from the mixed audio. For example, the directionally weighted acoustic features are input into a voiceprint coding network, which consists of a time-delay neural network and a statistical pooling layer, and outputs a fixed-dimensional speaker embedding vector. For overlapping speech segments of multiple speakers, a clustering algorithm or a permutation-invariant training strategy is used to assign time-frequency units to different speaker clusters, and the voiceprint features of each speaker are estimated.
[0051] Subsequently, visual identity information and voiceprint features are cross-modal correlated and matched to generate an independent identity vector for each virtual speaker. The goal of this step is to associate the visual identity of each virtual speaker with its voiceprint features in the audio signal, thereby generating an independent identity vector for each virtual speaker. For example, a similarity matrix between the visual identity feature vector and the voiceprint feature vector is calculated, and the optimal correspondence between visual identity and voiceprint features is established using the Hungarian algorithm or a greedy matching strategy to eliminate identity ambiguity between modalities. For successfully matched virtual speakers, their visual identity feature vector and voiceprint feature vector are fused using methods such as vector concatenation, bilinear pooling, or attention weighting to generate a unified identity vector. This vector, with a fixed-length real number, represents the consistency of a specific virtual speaker's identity in both the audiovisual and visual modalities. Isolated modal features that do not match are marked as unknown identity or environmental noise and are not included in subsequent separation processing.
[0052] Next, based on the identity vectors, a speech separation network is used to separate the speech stream of each target speaker from the mixed audio. The speech separation network is a deep learning model specifically designed to separate the speech of a single speaker from mixed audio containing multiple speakers. This network employs an encoder-splitter-decoder architecture. The encoder encodes the mixed audio into a high-dimensional latent representation. The splitter receives the identity vectors of each target speaker as conditional input, injects the identity information into the latent representation space through conditional normalization, and generates a separation mask corresponding to each speaker's identity. The decoder restores the weighted latent representation to a time-domain waveform or spectrogram, outputting an independent speech stream for each target speaker. The separation process must preserve the spatial orientation information of each speech stream, ensuring that the separated speech retains its three-dimensional spatial location attributes within the original virtual scene.
[0053] Finally, for the separated individual target speaker's speech stream, combined with the corresponding virtual speaker's lip movement features, speech enhancement is performed to generate an enhanced audio stream with spatial orientation awareness. For example, the single-speaker speech stream is input into a lip movement-speech alignment subnet, and the cross-modal consistency score between the speech stream and the corresponding virtual speaker's lip movement features is calculated. If the consistency score meets a preset threshold, it indicates that the separation result is correct and the audio and video are synchronized, and the speech stream is directly output as the enhancement result. If the consistency score is lower than the threshold, it indicates that there may be a separation error or audio-video asynchrony, and an enhancement correction mechanism is triggered, including adjusting the confidence threshold of the separation mask, introducing lip movement-guided time alignment constraints, or fusing spatial orientation information for secondary separation decisions. The single-speaker speech stream, after consistency verification and enhancement processing, is superimposed with its original spatial orientation information and output by a binaural rendering engine as immersive audio with clear spatial positioning, achieving accurate separation, enhancement, and spatialized presentation of target speech in multi-speaker scenarios.
[0054] Furthermore, the cross-modal association matching is implemented through a learnable memory network. In this embodiment, step P56 further includes: P56-1: storing the registration entry of each virtual speaker in the memory network, wherein the registration entry includes a visual identity feature template and a voiceprint feature template; P56-2: calculating the similarity score between the visual identity feature and voiceprint feature of the current frame and the registration entry in real time; P56-3: determining the identity vector of the target speaker through an optimal matching algorithm; P56-4: creating a new registration entry in the memory network when a new speaker is detected.
[0055] Optionally, cross-modal association matching is implemented through a learnable memory network. Specifically, this memory network can create a registration entry for each known virtual speaker during the initialization phase or before system operation. The registration entry includes a visual identity feature template and a voiceprint feature template. The visual identity feature template is obtained by extracting high-dimensional embedding vectors of the virtual speaker's facial animation data through a pre-trained face recognition network or role classification network, containing the virtual speaker's appearance information in the virtual environment, such as facial features, head posture, and lip movements. The voiceprint feature template is obtained by extracting clean speech samples of the virtual speaker in a quiet environment through a voiceprint coding network, containing the virtual speaker's vocal characteristics, such as timbre and intonation. Both types of feature templates are stored in the memory network's registry in key-value pairs. Each registration entry is appended with a unique identifier, a creation timestamp, and a confidence weight. The confidence weight is initialized to a uniform value or pre-assigned based on the template quality.
[0056] During real-time operation, the memory network performs similarity calculations between the current features and registered entries for each processing frame. Whenever a new audio or video signal is input, the visual identity features and voiceprint features of the current frame are extracted. These features are then compared with the features of all registered entries stored in the memory network to calculate a similarity score. For example, the visual identity features extracted from the current frame are compared with the visual identity feature templates of each registered entry in the memory network using cosine similarity or the inverse of Euclidean distance to obtain a visual similarity score vector. Next, the voiceprint features extracted from the current frame are compared with the voiceprint feature templates of each registered entry to obtain an acoustic similarity score vector. The visual and acoustic similarity score vectors are then weighted and fused. The fusion weights are determined by learnable parameters or an adaptive mechanism based on modal reliability to generate a comprehensive similarity score, representing the matching probability between the current speaker and a known identity.
[0057] Next, based on the comprehensive similarity score matrix, the target speaker's identity vector is determined using an optimal matching algorithm. For example, the Hungarian algorithm can be used to find the globally optimal speaker-identity assignment scheme, maximizing the overall similarity score while satisfying the one-to-one matching constraint; alternatively, a greedy matching algorithm can be used to determine matching relationships sequentially in descending order of similarity score until all current speakers have completed identity assignment or reached the lower limit of the similarity threshold. For a successfully matched speaker, its current visual identity features and voiceprint features are updated using a moving average with the template of the corresponding registered entry, enabling online learning and evolution of the identity template. Finally, the identity identifier of the matched registered entry is concatenated with the fused audiovisual feature vector to generate the target speaker's identity vector, which is used for subsequent speech separation and enhancement processing.
[0058] When a new speaker is detected, a new registration entry is created in the memory network. New speaker detection is based on outliers in the comprehensive similarity score matrix. If the maximum similarity between the current speaker and all existing registration entries is below a preset threshold, and the speaker maintains a stable presence across multiple consecutive frames, then the speaker is identified as a newly entered virtual speaker. A globally unique identifier is assigned to the new speaker. Using the visual and vocal features of the current frame as initial templates, a new registration entry is created and incorporated into the memory network. The confidence weight of the new registration entry is set to a low initial value. As the speaker remains active in the scene and accumulates observation data across multiple frames, its template quality and confidence weight are gradually improved, enabling dynamic expansion and adaptive updating of the memory network.
[0059] This memory network optimizes similarity metric parameters and template update strategies through end-to-end training, enabling cross-modal association matching to maintain robustness and real-time performance in multi-user scenarios with dynamically changing numbers of virtual speakers.
[0060] Furthermore, based on the identity vector, the speech stream of each target speaker is separated from the mixed audio by a speech separation network. Step P57 of this embodiment further includes: P57-1: using the identity vector as a conditional input to modulate the channel attention mechanism of the speech separation network; P57-2: generating a time-frequency mask corresponding to the target speaker based on the channel attention mechanism; P57-3: applying the time-frequency mask to filter the binaural audio signal to obtain the separated speech stream.
[0061] Specifically, the process of separating the speech stream of each target speaker from mixed audio can be further refined using a speech separation network.
[0062] First, the speech separation network adopts an encoder-splitter-decoder architecture, where the identity vector serves as the channel attention mechanism of the conditional input modulation splitter. The encoder first downsamples the binaural audio signal, then maps the temporal waveform or complex spectrum to a high-dimensional latent feature space through one-dimensional convolutional layers and multiple residual convolutional blocks, outputting a latent feature tensor. The dimension of this tensor is the number of channels multiplied by the number of time frames. The channel dimension represents the abstract feature components of the audio signal, and the time frame dimension represents the temporal evolution of the signal.
[0063] Next, the identity vector is input into the conditional embedding layer, which is composed of a multilayer perceptron. This layer maps the fixed-dimensional identity vector to a conditional vector matching the number of latent feature channels. Then, the conditional vector is expanded to the same time dimension as the latent features and subjected to broadcast addition or feature concatenation with the latent features, achieving initial fusion of identity information and acoustic features. Following this, the channel attention mechanism receives the fused feature representation and calculates the importance weights of each feature channel. Specifically, global average pooling is performed on the fused features to obtain channel statistical descriptions. After dimensionality reduction and mapping with an activation function, a channel weight vector is generated. This vector has the same dimension as the number of feature channels, and its element values range from zero to one. The conditional embedding of the identity vector is modulated element-wise with the channel weight vector, making the channel attention distribution adaptive to the target speaker's identity characteristics. This enhances feature channels related to the target speaker's acoustic characteristics and suppresses feature channels related to other speakers or noise.
[0064] Based on the modulated channel attention mechanism, a time-frequency mask corresponding to the target speaker is generated. The time-frequency mask is a binary mask that helps separate the speech components of a specific target speaker from a mixed audio signal by weighting the audio signal in both the time and frequency domains. The channel attention mechanism generates a time-frequency mask associated with the target speaker by weighting different frequency components of the input audio signal. Specifically, each element of the mask represents the weight of the target speaker's speech component at a specific time and frequency point. For audio components corresponding to the target speaker, the mask value is close to 1; for other speakers or noise components, the mask value is close to 0. This mask guides the subsequent speech stream separation process, enhancing the target speaker's speech while suppressing other irrelevant audio components.
[0065] Next, a time-frequency mask is applied to filter the binaural audio signal, resulting in a separated speech stream. An element-wise complex multiplication operation is then performed between the time-frequency mask and the short-time Fourier transform complex spectrum of the binaural audio signal. This operation enhances the target speaker's speech components while suppressing the sounds of other speakers or environmental noise, thus separating the target speaker's speech. The filtered signal is the target speaker's speech stream, containing the sound information related to the target speaker extracted from the mixed audio. For multiple target speakers, the above conditional modulation, mask generation, and filtering / reconstruction processes are executed in parallel to generate independent speech streams corresponding one-to-one with each identity vector. Each speech stream retains its spatial orientation attributes in the original virtual scene, providing a foundation for subsequent spatial rendering and immersive presentation.
[0066] Furthermore, the immersive fusion enhancement network applies acoustic effects matching the environment to the enhanced audio stream using acoustic material properties. In this embodiment, step P50 further includes: P59: generating a reflection coefficient and sound absorption coefficient matrix based on the material type of the virtual scene object; P510: calculating the environmental impulse response using a differentiable acoustic simulator based on the reflection coefficient and sound absorption coefficient matrix; P511: performing a convolution operation between the impulse response and the enhanced audio stream to simulate the acoustic effects of a real environment.
[0067] Optionally, the immersive fusion augmented network can also utilize acoustic material properties to apply acoustic effects processing to the augmented audio stream that matches the virtual environment, thereby enhancing the user's immersion and interactive experience.
[0068] First, reflection and absorption coefficient matrices are generated based on the material types of objects in the virtual scene. Each object in the virtual scene, such as walls, floors, and ceilings, has different material properties, which determine the reflection and absorption of sound on that object. For example, hard surfaces, such as glass or metal, produce strong reflections, while soft surfaces, such as fabric or carpet, absorb more sound. Based on the material type of each scene object, the reflection coefficient describing the intensity of sound reflection and the absorption coefficient describing the intensity of sound absorption are calculated. For example, the geometric model and material labels in the virtual reality scene are parsed to identify the surface material categories of each virtual object, including preset types such as concrete, wood, fabric, metal, and glass. Then, the acoustic characteristic parameters of each material type in a wide frequency range are queried from the acoustic physics database, including the reflection coefficient, which characterizes the proportion of sound wave energy reflected by the surface after incident; and the absorption coefficient, which characterizes the proportion of sound wave energy absorbed and converted by the surface. Finally, for each virtual object in the scene, a reflection coefficient matrix and a sound absorption coefficient matrix are constructed based on the orientation, area, and material label of its geometric surface. The matrix dimension is the number of object surfaces multiplied by the number of frequency bands, and the matrix elements represent the energy interaction characteristics of each surface in each frequency band. In addition, for complex materials or composite structures, a multi-layer acoustic model can be used to calculate the equivalent coefficients.
[0069] Next, based on the reflection and absorption coefficient matrices, the environmental impulse response is calculated using a differentiable acoustic simulator. The environmental impulse response is an important tool for describing the propagation characteristics of sound in a virtual environment; it represents the change of a sound signal after acoustic processes such as reflection, refraction, and absorption in the environment. Using a differentiable acoustic simulator, it is possible to simulate how each sound source in the virtual environment interacts with different surfaces in the scene and calculate the corresponding impulse response. Specifically, the differentiable acoustic simulator employs either the mirror source method based on geometric acoustics or the finite difference time-domain method based on wave acoustics. The mirror source method tracks the propagation path of sound waves from the sound source through reflections at various surfaces to the receiving point, calculating the arrival time, direction, and energy attenuation of each order of reflected sound. It is suitable for mid-to-high frequency bands and early reflection simulations. The finite difference time-domain method simulates wave propagation and diffraction effects by discretely solving the sound wave equations, making it suitable for low frequency bands and complex geometric scenarios. The simulator uses the reflection coefficient matrix and absorption coefficient matrix as boundary condition inputs, and the virtual sound source position and the user's head position as excitation and receiving points. It iteratively calculates the propagation and attenuation process of the sound energy pulse in the scene, outputting an environmental impulse response sequence. This sequence characterizes the acoustic transmission characteristics of the entire path from the sound source to the receiving point, including three stages: direct sound, early reflection, and late reverberation. It can provide detailed information about sound propagation, such as sound delay, reflection, and reverberation.
[0070] Subsequently, the impulse response is convolved with the enhanced audio stream to simulate realistic environmental acoustics. For each target speaker's enhanced audio stream, the environmental impulse response corresponding to the speaker's virtual location and the user's head position is selected. A fast convolution algorithm or frequency domain multiplication is used to perform linear convolution, where the output signal equals the convolution integral of the input signal and the impulse response. In the discrete implementation, this is a point-by-point multiplication and summation. Next, the convolution operation integrates the signal characteristics of the enhanced audio stream with the environmental acoustic characteristics of the scene, resulting in output audio that matches the spatial sense, immersion, and distance of the virtual scene's geometry and material distribution. This process adds environmental acoustic effects to the enhanced audio stream, enabling the audio signal to not only contain spatial location information but also reflect the acoustic characteristics of the virtual environment. For example, if the audio signal passes through a hard surface with a high reflectivity, the convolution operation will make the reflected sound more pronounced; if the audio signal interacts with an object with a high absorption coefficient, the high-frequency components of the audio signal will attenuate, thus enhancing the realism of the environment. Through convolution operations, the enhanced audio stream is given acoustic effects that match the virtual environment, enabling users to perceive a natural interaction between sound and scene.
[0071] P60: The enhanced audio stream is output through the binaural rendering engine to provide users with immersive voice interaction.
[0072] It should be understood that the enhanced audio stream generated through the above steps will be output through a binaural rendering engine to provide users with an immersive voice interaction experience. First, the enhanced audio stream already includes audio features that have undergone spatial audio coding, speech separation, cross-modal consistency constraints, and environmental acoustic effect simulation. These processes ensure that the audio signal has a high degree of realism and immersion in terms of directionality, temporal synchronization, and acoustic matching with the virtual environment. Next, these enhanced audio features will be processed and output through the binaural rendering engine.
[0073] The task of a binaural rendering engine is to simulate spatial sound perception in the real world, based on the user's headphone or headset settings and combining audio from the virtual environment with the user's spatial location information. This engine uses a head-related transfer function (HRTF) model to simulate how sound reaches the user's ears, taking into account the spatial orientation, distance, and acoustic characteristics of the sound source. Based on the user's head posture data and the spatial location of the sound source in the virtual scene, the binaural rendering engine can accurately locate the source of sound in the virtual environment, enabling the user to perceive sounds from different directions and enhancing the immersive experience of the virtual environment.
[0074] Specifically, the binaural rendering engine adjusts the directionality of the audio signal based on spatial orientation information in the enhanced audio stream and the user's current head posture, such as through sensors in the head-mounted display or headphones. It simulates the effect of sound traveling to both ears from different directions by calculating the time difference and loudness difference between them. Furthermore, the binaural rendering engine also simulates acoustic reflections and reverberation in the environment, making the final output audio signal more consistent with the real-world auditory experience.
[0075] Through processing by the binaural rendering engine, the final generated audio signal will be transmitted to the user via headphones or other audio output devices, providing a highly immersive audio experience. Users can accurately perceive the direction, location, and movement of multiple sound sources in the virtual environment through the audio signal, while simultaneously interacting with other elements in the virtual world. For example, in a virtual meeting, users can clearly hear the voices of each participant and accurately determine their location; in virtual games or immersive experiences, users can perceive changes in ambient sound, enhancing the realism of the interaction.
[0076] In summary, the embodiments of this application have at least the following technical effects: This application dynamically adjusts attention weights through head posture data to achieve a natural binding between the user's auditory focus direction and visual focus direction, thereby improving the spatial selectivity of speech enhancement; it integrates virtual sound source location features and acoustic features to construct an explicit three-dimensional spatial orientation perception capability, enabling the enhanced speech to possess locatable spatial attributes; it combines virtual speaker lip movement features for cross-modal consistency constraints, utilizes visual information to assist speech separation and enhancement, and suppresses visually irrelevant interference sound sources; it achieves continuous tracking and association of multiple speaker identities through a learnable memory network, supporting multi-user interaction scenarios such as dynamically changing virtual meetings; it generates environment-matched impulse responses using acoustic material properties, making the speech reverberation effect consistent with the geometric appearance of the virtual scene, enhancing physical realism; and it provides users with an immersive voice interaction experience with externalized perception, distance, and a sense of envelopment through personalized binaural rendering output.
[0077] It achieves the technical effect of integrating dynamic attention to head posture and cross-modal constraints of lip movement, realizing the natural binding of auditory attention and visual focus, and effectively suppressing interfering sound sources.
[0078] Example 2: Exemplary electronic device.
[0079] The electronic device of an embodiment of this application will now be described with reference to FIG2.
[0080] Based on the same inventive concept as the multimodal speech enhancement method based on deep learning models in the foregoing embodiments, this application also provides an electronic device, including: a processor coupled to a memory, the memory being used to store a program, which, when executed by the processor, implements the steps of the method described in Embodiment 1.
[0081] The electronic device 300 includes a processor 302, a communication interface 303, and a memory 301. Optionally, the electronic device 300 may also include a bus architecture 304. The communication interface 303, processor 302, and memory 301 can be interconnected via the bus architecture 304. The bus architecture 304 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus architecture 304 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in Figure 2, but this does not indicate that there is only one bus or one type of bus.
[0082] Processor 302 may be a CPU, microprocessor, ASIC, or one or more integrated circuits used to control the execution of programs according to the present application.
[0083] Communication interface 303 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), wired access network, etc.
[0084] Memory 301 can be ROM or other types of static storage devices capable of storing static information and instructions, RAM or other types of dynamic storage devices capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory can exist independently and be connected to the processor via bus architecture 304. Memory can also be integrated with the processor.
[0085] The memory 301 stores computer execution instructions for implementing the scheme of this application, and the processor 302 controls the execution. The processor 302 executes the computer execution instructions stored in the memory 301, thereby implementing the multimodal speech enhancement method based on a deep learning model provided in the above embodiments of this application.
[0086] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0087] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0088] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A multimodal speech enhancement method based on a deep learning model, characterized in that, The method includes: acquiring head posture data, binaural audio signals, and visual context information of a target user in a virtual reality environment, wherein the visual context information includes at least spatial location information of sound sources in the virtual scene and facial animation data of a virtual speaker; encoding the binaural audio signals into acoustic features containing three-dimensional spatial orientation information using a spatial audio encoding module; extracting virtual sound source location features and virtual speaker lip movement features from the visual context information using a visual context encoding module; inputting the head posture data, the acoustic features, the virtual sound source location features, and the virtual speaker lip movement features into an immersive fusion enhancement network; the immersive fusion enhancement network dynamically adjusts attention weights based on the head posture data, selectively enhances or suppresses acoustic features from different spatial directions, and combines the virtual speaker lip movement features for cross-modal consistency constraints to generate an enhanced audio stream with spatial orientation awareness; and outputting the enhanced audio stream through a binaural rendering engine to provide immersive voice interaction for the user.
2. The multimodal speech enhancement method based on a deep learning model as described in claim 1, characterized in that, The spatial audio coding module encodes the binaural audio signal into acoustic features containing three-dimensional spatial orientation information, including: performing a short-time Fourier transform on the binaural audio signal to obtain a complex spectrum; extracting orientation-sensitive features associated with the head-related transfer function from the complex spectrum; spatially aligning the orientation-sensitive features with virtual sound source location features through a spatial attention layer to generate a spatial weight matrix; and weighting the complex spectrum according to the spatial weight matrix to obtain acoustic features containing three-dimensional spatial orientation information.
3. The multimodal speech enhancement method based on a deep learning model as described in claim 2, characterized in that, The visual context encoding module extracts virtual sound source location features and virtual speaker lip movement features from the visual context information, including: parsing the three-dimensional coordinates of the virtual sound source from the visual context information and converting them into spherical coordinate system parameters relative to the user's head to obtain virtual sound source location features; inputting the vertex coordinates of the lip region in the facial animation data into a three-dimensional convolutional neural network to output a lip vertex displacement sequence; and performing dynamic time warping on the lip vertex displacement sequence to generate virtual speaker lip movement features aligned with the complex spectral time sequence.
4. The multimodal speech enhancement method based on a deep learning model as described in claim 1, characterized in that, The immersive fusion enhancement network applies acoustic effects matching the environment to the enhanced audio stream using acoustic material properties, including: generating reflection coefficient and sound absorption coefficient matrices based on the material type of virtual scene objects; calculating the environmental impulse response using a differentiable acoustic simulator based on the reflection coefficient and sound absorption coefficient matrices; and performing a convolution operation between the impulse response and the enhanced audio stream to simulate the acoustic effects of a real environment.
5. The multimodal speech enhancement method based on a deep learning model as described in claim 1, characterized in that, The immersive fusion enhancement network includes: a spatial attention subnetwork for receiving the head pose data and the virtual sound source location features, and outputting orientation selection weights; a lip movement-speech alignment subnetwork for receiving the virtual speaker's lip movement features and the acoustic features, and outputting a cross-modal consistency score; and a fusion decision layer for combining the orientation selection weights and the cross-modal consistency score to generate a speech enhancement mask, and modulating the acoustic features based on the speech enhancement mask to output the enhanced audio stream.
6. The multimodal speech enhancement method based on a deep learning model as described in claim 1, characterized in that, The immersive fusion enhancement network dynamically adjusts attention weights based on the head posture data, selectively enhancing or suppressing acoustic features from different spatial directions. This includes: calculating the user's auditory focus direction in real time based on the head posture data; performing spatial matching calculations between the virtual sound source location features and the auditory focus direction to generate a directional weight mask; and weighting the acoustic features using the directional weight mask so that speech components corresponding to sound sources spatially aligned with the auditory focus direction are enhanced, while sound source components from other directions are suppressed. The calculation of the directional weight mask integrates a distance attenuation model and a spatial reverberation model to simulate the characteristics of a real three-dimensional sound field.
7. The multimodal speech enhancement method based on a deep learning model as described in claim 6, characterized in that, Combining the lip movement features of the virtual speaker with cross-modal consistency constraints, an enhanced audio stream with spatial orientation awareness is generated, including: identifying the visual identity information of multiple virtual speakers in the virtual reality environment; extracting the voiceprint features of each virtual speaker from the mixed binaural audio signals; performing cross-modal association matching between the visual identity information and the voiceprint features to generate an independent identity vector for each virtual speaker; based on the identity vector, separating the speech stream of each target speaker from the mixed audio through a speech separation network; and for the separated speech stream of a single target speaker, combining its corresponding virtual speaker lip movement features, performing a speech enhancement operation to generate an enhanced audio stream with spatial orientation awareness.
8. The multimodal speech enhancement method based on a deep learning model as described in claim 7, characterized in that, The cross-modal association matching is achieved through a learnable memory network, including: storing a registration entry for each virtual speaker in the memory network, wherein the registration entry includes a visual identity feature template and a voiceprint feature template; calculating the similarity score between the visual identity feature and voiceprint feature of the current frame and the registration entry in real time; determining the identity vector of the target speaker through an optimal matching algorithm; and creating a new registration entry in the memory network when a new speaker is detected.
9. The multimodal speech enhancement method based on a deep learning model as described in claim 7, characterized in that, Based on the identity vector, the speech stream of each target speaker is separated from the mixed audio by a speech separation network, including: using the identity vector as a conditional input to modulate the channel attention mechanism of the speech separation network; generating a time-frequency mask corresponding to the target speaker based on the channel attention mechanism; and applying the time-frequency mask to filter the binaural audio signal to obtain the separated speech stream.
10. An electronic device, characterized in that, include: A processor coupled to a memory for storing a program that, when executed by the processor, implements the steps of the method as claimed in any one of claims 1 to 9.
Citation Information
Cited By
A financial service video post-event quality inspection method, device, equipment and medium
CN122290225A