Audio processing method and apparatus, and XR device

By constructing a relationship library between voiceprint features and face information, using the voice separation model and the audio and video data of XR devices, the problem that XR devices are difficult to recognize target voice in noisy environments is solved, and the target audio signal is accurately identified without increasing hardware costs, reducing power consumption and computing power consumption.

CN120388580APending Publication Date: 2025-07-29HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 19 Cites 0 Cited by

Patent Information

Application Number
CN202510875034.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

It is difficult for existing XR devices to accurately identify the voice of the target speaking object in noisy environments, especially in multi-person scenarios. The existing hardware improvement solutions are costly and have limited effects, and cannot effectively suppress noise in the gaze direction and disturb vocals.

Method used

By constructing an association relationship library between voiceprint features and face information, using the voice separation model and audio and video data obtained by the XR device, identify the target audio and video signals of the target object, use the camera and microphone necessary for XR device to obtain the audio and video data, and build an association relationship library between voiceprint features and face information, directly call the association library to obtain the target voiceprint features, and identify the target audio signal from the audio to be processed.

Benefits of technology

Without increasing hardware costs, accurately identify the voice of the target speaking object, suppress background noise and interfere with human voice, save computing power, and reduce power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388580A_ABST
    Figure CN120388580A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method and device and XR equipment, and relates to the technical field of voice processing, and the method is applied to the XR equipment, and comprises the steps: obtaining a to-be-processed audio and the face information of a target object; obtaining a target voiceprint feature of the target object based on a pre-constructed association relationship library of voiceprint features and face information and the face information of the target object; wherein the association relationship library of the voiceprint features and the face information is constructed based on a voice separation model and audio and video data acquired by the XR equipment; and based on the target voiceprint feature of the target object, identifying a target audio signal of the target object from the to-be-processed audio. According to the method, the audio signal of the target object can be accurately identified in a noisy environment, especially a multi-person scene, on the premise that the hardware cost of XR equipment is not increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and particularly to an audio processing method, apparatus and XR device. Background Art

[0002] In scenarios such as on-site meetings and exhibitions, users can facilitate communication between people of different languages by using the listening and translation function of XR (Extended Reality) devices. However, the quality of translation is often significantly affected by various interferences such as background noise, reverberation, and interfering speakers. Currently, typical speech enhancement schemes mainly focus on removing background noise and reverberation, and cannot filter out interfering human voices. Therefore, how to accurately identify and extract the speech of the target speaker and suppress other interfering sounds by XR devices in a noisy environment (especially with interfering human voices) has great practical value.

[0003] In the prior art, the hardware of XR devices is generally improved. For example, the microphone is changed to a microphone array, the target position is determined based on the user's gaze, and beamforming is performed on the audio using the microphone array; or, an image sensor is added to capture the human eye, and eye movement tracking is performed to calculate the fixation point, and then the sound source in the direction of the fixation point is detected by the sensor. In addition, the above methods can only suppress the interference in non-gaze directions and cannot effectively suppress the noise and other interfering speech sounds in the gaze direction. Therefore, how to accurately identify the speech of the target speaker in a noisy environment (especially in a multi-person scenario) without increasing the hardware cost is an urgent problem to be solved currently. Summary of the Invention

[0004] The present invention provides an audio processing method, apparatus and XR device to solve the defects that in the process of audio processing based on XR devices in the prior art, the hardware cost of XR devices is relatively high, and the accuracy of the audio recognition result of the target object is relatively poor.

[0005] The present invention provides an audio processing method applied to an XR device, including: [[ID=2I]]Obtaining the audio to be processed and the face information of the target object; Based on the pre-constructed association relationship library between the voiceprint feature and the face information, and the face information of the target object, obtaining the target voiceprint feature of the target object; wherein, the association relationship library between the voiceprint feature and the face information is constructed based on the voice separation model and the audio-visual data obtained by the XR device; Based on the target voiceprint feature of the target object, identifying the target audio signal of the target object from the audio to be processed.

[0006] According to the audio processing method provided by the present invention, the construction steps of the association relationship library between the voiceprint feature and the face information include: Obtain a scene image, perform face detection on the scene image to obtain at least one piece of face information; Obtain audio-visual data including the at least one piece of face information, and use the speech separation model to process the audio-visual data to determine an audio signal corresponding to the at least one piece of face information; Based on the audio signal corresponding to the at least one piece of face information, obtain a voiceprint feature corresponding to the at least one piece of face information; Based on the voiceprint feature corresponding to the at least one piece of face information, construct an association relationship library between the voiceprint feature and the face information.

[0007] According to an audio processing method provided by the present invention, the speech separation model includes an audio encoder, a visual encoder, an audio-visual fusion module, a mask estimation network, and an audio decoder. The audio-visual data includes audio data and video data. Using the speech separation model to process the audio-visual data to determine an audio signal corresponding to at least one piece of face information includes: Input the audio data into the audio encoder to obtain audio features output by the audio encoder; Input the video data into the visual encoder to obtain facial visual features output by the visual encoder; Input the audio features and the facial visual features into the audio-visual fusion module to obtain audio-visual fusion features output by the audio-visual fusion module; Input the audio-visual fusion features into the mask estimation network to obtain a time-frequency mask output by the mask estimation network; Input the time-frequency mask into the audio decoder to obtain an audio signal corresponding to at least one piece of face information output by the audio decoder.

[0008] According to an audio processing method provided by the present invention, the audio encoder includes a U-shaped network U-Net encoder, a time-frequency transformer transformer network, and an audio fusion module. The visual encoder includes a face attribute sub-network, a lip movement sub-network, and a visual fusion module. Inputting the audio data into the audio encoder to obtain audio features output by the audio encoder includes: Input the audio data into the U-Net encoder to obtain short-term audio features output by the U-Net encoder; Input the audio data into the time-frequency transformer transformer network to obtain long-distance dependency features output by the time-frequency transformer transformer network; Input the short-term audio features and the long-distance dependence features into the audio fusion module to obtain the audio features output by the audio fusion module; The step of inputting the video data into the visual encoder to obtain the facial visual features output by the visual encoder includes: Input the video data into the face attribute sub-network to obtain the face appearance features output by the face attribute sub-network; Input the video data into the lip movement sub-network to obtain the lip movement sequence features output by the lip movement sub-network; Input the face appearance features and the lip movement sequence features into the visual fusion module to obtain the facial visual features output by the visual fusion module.

[0009] According to an audio processing method provided by the present invention, the step of inputting the audio features and the facial visual features into the audio-visual fusion module to obtain the audio-visual fusion features output by the audio-visual fusion module includes: Input the audio features and the facial visual features into the audio-visual fusion module, and perform dynamic fusion on the audio features and the facial visual features through the audio-visual fusion module to obtain the current audio-visual fusion features; Based on the current audio-visual fusion features, weighted fusion of historical audio-visual fusion features is performed from a multi-level feature memory bank to obtain audio-visual fusion features; wherein, the historical audio-visual fusion features include one or more of short-term memory audio-visual fusion features, long-term memory audio-visual fusion features, and global memory audio-visual fusion features.

[0010] According to an audio processing method provided by the present invention, the step of acquiring the audio-visual data including the at least one face information and using the voice separation model to process the audio-visual data to determine the audio signals corresponding to the at least one face information includes: Acquire real-time audio-visual data including the at least one face information, where the real-time audio-visual data includes real-time audio data and real-time video data; Perform block processing on the real-time audio data through a time sliding window, and perform frame processing on the real-time video data; Input the block-processed real-time audio data and the frame-processed real-time video data into the voice separation model to obtain the block audio signals corresponding to the at least one face information output by the voice separation model; Stitch the block audio signals through a smoothing window to obtain the audio signals corresponding to the at least one face information.

[0011] An audio processing method provided by the present invention, for identifying a target audio signal of a target object from the audio to be processed based on the target voiceprint feature of the target object, includes: Identifying, from the audio to be processed, audio signals of at least one object and the voiceprint feature of each audio signal; Based on the voiceprint feature of each audio signal and the target voiceprint feature of the target object, obtaining the target audio signal of the target object from the audio signals of the at least one object; After identifying, from the audio to be processed, audio signals of at least one object and the voiceprint feature of each audio signal, it further includes: If it is detected that at least one voiceprint feature in the voiceprint features of each audio signal is not in the association relationship library, updating the association relationship library based on the detection result.

[0012] An audio processing method provided by the present invention, after identifying the target audio signal of the target object from the audio to be processed based on the target voiceprint feature of the target object, it further includes: Performing intensity change processing on the target audio signal of the target object in the audio to be processed; or, Performing type conversion processing on the target audio signal of the target object.

[0013] The present invention further provides an audio processing apparatus, including the following modules: A first acquisition module, configured to acquire an audio to be processed and the face information of a target object; A second acquisition module, configured to obtain the target voiceprint feature of the target object based on a pre-constructed association relationship library between voiceprint features and face information, and the face information of the target object; wherein, the association relationship library between voiceprint features and face information is constructed based on audio-visual data acquired by a voice separation model and an XR device; An audio recognition module, configured to identify the target audio signal of the target object from the audio to be processed based on the target voiceprint feature of the target object.

[0014] The present invention further provides an XR device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the audio processing method described in any one of the above is implemented.

[0015] The audio processing method, device, and XR device provided by the present invention construct a correlation relationship library between voiceprint features and face information through a voice separation model and the audio-visual data obtained by the XR device. When the audio to be processed and the face information of the target object are obtained, according to the face information of the target object, the target voiceprint feature corresponding to the target object is obtained from the pre-constructed correlation relationship library; furthermore, based on the target voiceprint feature, the target audio signal of the target object is identified from the audio to be processed. The present invention does not require improvement of the hardware of the XR device. It only needs to use the camera and microphone necessary for the XR device to obtain audio-visual data, and then uses the voice separation model to determine the voices of each speaker in the audio-visual data, and then constructs a correlation relationship library between voiceprint features and face information. In subsequent audio processing scenarios, by directly calling the correlation relationship library to obtain the target voiceprint feature of the target object, the target audio signal of the target object can be accurately identified from the audio to be processed based on the target voiceprint feature. Through the above method, the audio signal of the target object can be completely identified from various interference sources such as background noise, reverberation, and interfering human voices. At the same time, the present invention pre-constructs a correlation relationship library. In the audio processing scenario, only the audio data needs to be processed. Compared with processing the audio-visual data using the voice separation model each time, the present invention can also save computing power and reduce power consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 is the system architecture diagram of the audio processing system provided by the present invention; Figure 2 is one of the flow diagrams of the audio processing method provided by the present invention; Figure 3 is another flow diagram of the audio processing method provided by the present invention; Figure 4 is the structural diagram of the audio processing device provided by the present invention; Figure 5 is the structural diagram of the XR device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] In scenarios such as on-site meetings and exhibitions, users can facilitate communication between people of different languages by using the listening and translation function of XR (Extended Reality) devices. However, the quality of translation is often significantly affected by various interferences such as background noise, reverberation, and interfering speakers. Currently, typical speech enhancement solutions mainly focus on removing background noise and reverberation and cannot filter out interfering human voices. Therefore, it has great practical value for XR devices to accurately identify and extract the speech of the target speaker and suppress other interfering sounds in a noisy environment (especially interfering human voices).

[0020] Related technologies use beamforming of a microphone array to be directed at the target speaker. Beamforming includes combining the audio of each microphone in the array in a specific manner based on the position of the target speaker. Based on the user's gaze to determine the position of the target speaker and perform beamforming on the audio accordingly, AR applications can use this eye-tracking beamforming (i.e., gaze-point beamforming) to enhance the sound from the gaze direction and suppress the sound from other directions.

[0021] There is also related technology that provides a head-mounted unit for assisting users (such as hearing-impaired users). The head-mounted unit includes a tracking sensor, a sensor, and a user interface. Among them, the tracking sensor is used to monitor the user wearing the head-mounted unit to determine the gaze direction that the user is looking at; the sensor is used to detect the sound source located in the identified gaze direction; the user interface provides information to the user to help the user identify the sound from the sound source.

[0022] The main disadvantages of the above-mentioned related technologies are as follows: (1) Limited performance. Due to the influence of physical size, price, etc., the microphone array equipped in XR devices has poor beamforming effect in the low-frequency band, resulting in insufficient noise suppression ability; at the same time, the directional beam cannot suppress the noise in the gaze direction and the speech of other interfering people.

[0023] (2) Increased cost and power consumption. High-performance beamforming requires strict matching of the gain, phase, and position accuracy of microphone elements, resulting in an increase in hardware cost; the gaze-point function requires more than two image sensors to capture the human eye and requires sufficient computing power to perform eye movement tracking to calculate the gaze point, resulting in a further increase in hardware cost and R & D cost, and at the same time increasing power consumption.

[0024] Therefore, how to accurately identify the voice of the target speaker in a noisy environment (especially in a multi-person scenario) without increasing the hardware cost is an urgent problem to be solved at present.

[0025] Based on the above problems, the present invention proposes an audio processing method, apparatus and XR device, which will be described below in conjunction with Figures 1 - 5 for description.

[0026] Figure 1 is the system architecture diagram of the audio processing system provided by the present invention. As Figure 1 shown, the audio processing system includes an Extended Reality (XR) device 01 and a server 02.

[0027] Among them, the XR device 01 includes but is not limited to: VR (Virtual Reality) device, AR (Augmented Reality) device and MR (Mixed Reality) device, which is used to execute the audio processing method of the present invention. The server 02 can be a server, which is used to train the voice separation model.

[0028] After the server 02 finishes training the voice separation model, it deploys the trained voice separation model to the XR device 01. The XR device 01 can build a correlation library of voiceprint features and face information based on the voice separation model. In addition, the user can wear the XR device 01 and perform specific interaction operations, such as gesture interaction, head movement interaction or eye movement interaction, etc. After the XR device 01 detects the user's interaction operation, it obtains the face information of the target object through the camera based on the interaction operation. At the same time, it collects the audio to be processed through the microphone, and then, based on the pre-built correlation library of voiceprint features and face information, and the face information of the target object, it obtains the target voiceprint feature of the target object, and then, based on the target voiceprint feature of the target object, it identifies the target audio signal of the target object from the audio to be processed.

[0029] In some other embodiments of the present invention, the XR device 01 may not deploy the voice separation model, and the server 02 builds the correlation library of voiceprint features and face information. Specifically, the XR device 01 obtains audio-visual data and sends it to the server 02. The server 02 builds the correlation library of voiceprint features and face information based on the received audio-visual data and the voice separation model. Then, the server 02 sends the correlation library of voiceprint features and face information to the XR device 01 for the XR device 01 to use.

[0030] Figure 2 is one of the flow diagrams of the audio processing method provided by the present invention. AsFigure 2 As shown in Figure 2 , the audio processing method includes: Step S110: Obtain the audio to be processed and the face information of the target object.

[0031] In this embodiment, the audio processing method is applied to XR devices.

[0032] An XR device refers to a wearable or portable device that realizes human-computer interaction by integrating virtual and real environments through hardware and software technologies. XR devices include, but are not limited to: VR (Virtual Reality) devices, AR (Augmented Reality) devices, and MR (Mixed Reality) devices. Among them, a VR device uses computer technology to simulate and generate a three-dimensional virtual space, allowing users to immerse themselves in it and interact with it to obtain an immersive experience; an AR device fuses virtual information with the real world through technology and superimposes it on the real scene in real time to enhance the sensory experience; an MR device mixes the real world and the virtual world together to generate a new visual environment, which contains both physical entities and virtual information, and users can interact with these physical entities and virtual information in real time.

[0033] The audio processing method is applicable to scenarios that require precise association between faces and voices, such as conference translation, AR / VR social interactions, and security monitoring.

[0034] Here, the audio to be processed is the audio recorded by the XR device, which includes the voice of the target object and may also include the voices of other objects (non-target objects) and background noise (such as environmental noise).

[0035] The target object is the target speaking object that the XR device wearer (hereinafter referred to as the user) wants to obtain. The face information can be a facial image.

[0036] As an implementation manner, the user can select the target object through interactive operations (such as gestures, head rotation gaze, and eye rotation gaze), and determine the face area in the camera through the spatial coordinate system mapping matrix to obtain the face information of the target object.

[0037] As an implementation manner, the facial images of one or more objects in the currently captured scene image can be displayed in the display area of the XR device for the user to select the target object.

[0038] Step S120: Based on the pre-constructed association relationship library between voiceprint features and face information, and the face information of the target object, obtain the target voiceprint feature of the target object; wherein, the association relationship library between voiceprint features and face information is constructed based on the voice separation model and the audio-visual data obtained by the XR device.

[0039] Here, the association relationship library between voiceprint features and face information is constructed based on a voice separation model and audio-visual data. The audio-visual data is obtained through a camera (such as an RGB (Red-Green-Blue) camera) and a microphone that are essential on XR devices.

[0040] As an implementation, the scene image where the user is located can be obtained through the camera on the XR device, face detection is performed on the scene image, and the audio-visual data corresponding to the detected face information is obtained, where at least one piece of detected face information is included. Then, the audio-visual data is processed using the voice separation model to determine the audio signal corresponding to the face information. Next, based on the audio signal corresponding to the face information, the voiceprint features corresponding to the face information are obtained. Finally, the face information (such as a facial image) is associated and saved with the voiceprint features corresponding to the face information to construct the association relationship library between voiceprint features and face information.

[0041] By pre-constructing the association relationship library between voiceprint features and face information, in subsequent audio processing scenarios, only the audio data needs to be processed, that is, the target voiceprint features corresponding to the target object can be directly obtained by calling this association relationship library, and then the target audio signal of the target object can be identified from the audio to be processed based on the target voiceprint features. Compared with processing the audio-visual data using the voice separation model each time to determine the target audio signal of the target object, this embodiment can save computing power and reduce power consumption. This is of significant significance for consumer-grade XR devices.

[0042] Step S130, based on the target voiceprint features of the target object, identify the target audio signal of the target object from the audio to be processed.

[0043] Specifically, the audio to be processed can be separated first to obtain the separated audio signal; then, the voiceprint features in the separated audio signal are extracted, and the extracted voiceprint features are compared with the target voiceprint features, and the target audio signal of the target object is determined according to the similarity comparison result. Further, the similarity comparison can be performed based on cosine similarity or probabilistic linear discriminant analysis.

[0044] The audio processing method provided by the present invention is applied to an XR device. By using a voice separation model and the audio-visual data obtained by the XR device, an association relationship library between voiceprint features and face information is constructed. When the audio to be processed and the face information of the target object are obtained, according to the face information of the target object, the target voiceprint feature corresponding to the target object is obtained from the pre-constructed association relationship library; and then, based on the target voiceprint feature, the target audio signal of the target object is identified from the audio to be processed. The present invention does not require any improvement to the hardware of the XR device. It only needs to use the camera and microphone that are essential for the XR device to obtain audio-visual data, and then uses the voice separation model to determine the voices of each speaker in the audio-visual data, and then constructs an association relationship library between voiceprint features and face information. In subsequent audio processing scenarios, by directly calling this association relationship library to obtain the target voiceprint feature of the target object, the target audio signal of the target object can be accurately identified from the audio to be processed based on this target voiceprint feature. In this way, the audio signal of the target object can be completely identified from various interference sources such as background noise, reverberation, and interfering voices. At the same time, by pre-constructing the association relationship library, in the audio processing scenario, only the audio data needs to be processed. Compared with using the voice separation model to process the audio-visual data each time, the present invention can also save computing power and reduce power consumption.

[0045] Figure 3 is the second schematic flowchart of the audio processing method provided by the present invention. As Figure 3 shown, the construction steps of the association relationship library between voiceprint features and face information include: step S10, step S11, step S12, and step S13.

[0046] In step S10, a scene image is obtained, and face detection is performed on the scene image to obtain at least one piece of face information.

[0047] The scene image is obtained through the camera that is essential for the XR device. The scene image can be an image taken based on the current gaze direction of the user, or multiple images in multiple gaze directions taken based on the rotation of the user.

[0048] During face detection, computer vision algorithms can be used for detection.

[0049] In step S11, audio-visual data including the at least one piece of face information is obtained, and the audio-visual data is processed by using the voice separation model to determine the audio signal corresponding to the at least one piece of face information.

[0050] After detecting at least one piece of face information, audio-visual data including this face information is obtained, and then, the audio-visual data is processed by using the voice separation model to determine the audio signal corresponding to this face information.

[0051] Among them, the voice separation model is trained based on sample audio-visual data, sample faces, and the annotation results of the audio signals of the sample faces.

[0052] Step S12: Based on the audio signals corresponding to the at least one face information, obtain the voiceprint features corresponding to the at least one face information.

[0053] Voiceprint features include but are not limited to: MFCC (Mel Frequency Cepstrum Coefficient), LPCC (Linear Predictive Cepstral Coefficients), spectral envelope, i-vector, x-vector. The extraction methods of various voiceprint features can refer to the prior art and will not be elaborated here.

[0054] Step S13: Based on the voiceprint features corresponding to the at least one face information, construct a correlation relationship library between the voiceprint features and the face information.

[0055] Associate and save the face information (such as facial images) with their corresponding voiceprint features to construct a correlation relationship library between the voiceprint features and the face information.

[0056] In this embodiment, by constructing a correlation relationship library between the voiceprint features and the face information, in subsequent audio processing scenarios, only the audio data needs to be processed. That is, directly call this correlation relationship library to obtain the target voiceprint features corresponding to the target object, and then based on this target voiceprint feature, the target audio signal of the target object can be identified from the audio data to be processed. Compared with using the voice separation model to process the audio-visual data every time to determine the target audio signal of the target object, this embodiment can save computing power and reduce power consumption. This is of significant significance for consumer-grade XR devices.

[0057] In one embodiment, the voice separation model includes an audio encoder, a visual encoder, an audio-visual fusion module, a mask estimation network, and an audio decoder. The audio-visual data includes audio data and video data. The step of "using the voice separation model to process the audio-visual data to determine the audio signals corresponding to at least one face information" includes steps S111, S112, S113, S114, and S115.

[0058] Step S111: Input the audio data into the audio encoder to obtain the audio features output by the audio encoder.

[0059] Here, the audio features include but are not limited to: short-term audio features, long-distance dependency features, and audio features obtained by fusing short-term audio features and long-distance dependency features.

[0060] As an implementation, when the audio feature is a short-term audio feature, the audio encoder can be a U-Net (U-shaped network) encoder to extract the short-term audio feature through this U-Net encoder.

[0061] As an implementation, when the audio feature is a long-distance dependency feature, the audio encoder can be a time-frequency transformer network to extract the long-distance dependency feature through this time-frequency transformer network.

[0062] As an implementation, when the audio feature is an audio feature obtained by fusing short-term audio features and long-distance dependency features, the audio encoder includes a U-Net encoder, a time-frequency transformer network, and an audio fusion module. The short-term audio feature can be extracted through the U-Net encoder, and at the same time, the long-distance dependency feature can be extracted through the time-frequency transformer network. Then, the short-term audio feature and the long-distance dependency feature are fused by the audio fusion module.

[0063] Step S112: Input the video data into the visual encoder to obtain the facial visual features output by the visual encoder.

[0064] Here, the facial visual features include but are not limited to: human face appearance features, lip movement sequence features, and facial visual features obtained by fusing human face appearance features and lip movement sequence features.

[0065] As an implementation, when the facial visual feature is based on the human face appearance feature, the visual encoder can be a ResNet-18 network structure (a deep residual network) to extract the human face appearance feature through this network structure.

[0066] As an implementation, when the facial visual feature is a lip movement sequence feature, the visual encoder can be a network structure composed of 3D-Conv (3D-Convolutional Neural Networks, three-dimensional convolutional neural network), ShuffleNet V2 (an efficient convolutional neural network), and TCN (Temporal Convolutional Network, time-domain convolutional network) to extract the lip movement sequence feature through this network structure.

[0067] As an implementation manner, when the facial visual feature is a facial visual feature obtained by fusing the facial appearance feature and the lip movement sequence feature, the visual encoder may include a face attribute sub-network, a lip movement sub-network, and a visual fusion module. The face appearance feature can be extracted through the face attribute sub-network, and the lip movement sequence feature can be extracted through the lip movement sub-network. Then, the face appearance feature and the lip movement sequence feature are fused by the visual fusion module.

[0068] Step S113: Input the audio feature and the facial visual feature into the audio-visual fusion module to obtain the audio-visual fusion feature output by the audio-visual fusion module.

[0069] As an implementation manner, the audio-visual fusion module dynamically evaluates the reliability of the audio feature and the facial visual feature through cross-modal dynamic gating fusion technology, calculates the cross-modal fusion weight based on the reliability, and then dynamically fuses the audio feature, the facial visual feature, and the cross-modal fusion weight to obtain the audio-visual fusion feature. In the case of non-streaming processing, this implementation manner can be adopted.

[0070] As an implementation manner, input the audio feature and the facial visual feature into the audio-visual fusion module, dynamically fuse the audio feature and the facial visual feature through the audio-visual fusion module to obtain the current audio-visual fusion feature; then, weighted-fuse the historical audio-visual fusion feature from the multi-level feature memory bank based on the current audio-visual fusion feature to obtain the audio-visual fusion feature; where the historical audio-visual fusion feature includes one or more of the short-term memory audio-visual fusion feature, the long-term memory audio-visual fusion feature, and the global memory audio-visual fusion feature. In the case of streaming processing, this implementation manner can be adopted, and the specific process can refer to the following embodiments.

[0071] Step S114: Input the audio-visual fusion feature into the mask estimation network to obtain the time-frequency mask output by the mask estimation network.

[0072] The mask estimation network (Mask Estimator) is used to automatically estimate and generate the time-frequency mask of the target object.

[0073] As an implementation manner, in the case of non-streaming processing, the time-frequency mask is determined according to the following method: ; where is the time-frequency mask, is the audio feature, is the audio-visual fusion feature fused by the cross-modal dynamic gating fusion technology, [[ID=,31]]is the frequency domain mask function.

[0074] As an implementation manner, in the case of streaming processing, the time-frequency mask is determined as follows: ; Wherein, is the time-frequency mask, is the audio feature, is the audio-visual fusion feature obtained by weighted fusion processing, is the frequency-domain masking function.

[0075] Step S115, input the time-frequency mask into the audio decoder to obtain the audio signal corresponding to at least one piece of face information output by the audio decoder.

[0076] The audio decoder is used to restore the time-frequency mask to a time-domain waveform and then output the corresponding audio signal. Among them, the process of restoring the time-frequency mask to a time-domain waveform is as follows: s ; Wherein, s is the time-domain waveform, is the time-frequency mask, is the short-time Fourier transform of the audio to be processed, is the Hadamard product, is the inverse short-time Fourier transform function. That is, after multiplying the time-frequency mask by the STFT of the audio to be processed, the time-domain signal is reconstructed through the inverse STFT.

[0077] In this embodiment, by processing the audio-visual data through the above voice separation model, the audio signal corresponding to the face information can be accurately determined.

[0078] In one embodiment, the audio encoder includes a U-shaped network U-Net encoder, a time-frequency transformer transformer network, and an audio fusion module. The above step S111 includes: step S1111, step S1112, and step S1113.

[0079] It should be noted that the execution order of step S1111 and step S1112 is not in sequence.

[0080] Step S1111, input the audio data into the U-Net encoder to obtain the short-term audio feature output by the U-Net encoder.

[0081] U-Net is a convolutional neural network architecture mainly used for separation. U-Net adopts an encoder-decoder structure. Here, the audio waveform of the input audio data can obtain short-term audio features through 1D convolution and residual blocks by the U-Net encoder. The short-term audio features mainly include time-domain features and frequency-domain features, which focus on the changing characteristics of audio data in time and frequency.

[0082] Step S1112: Input the audio data into the time-frequency transformer network to obtain the long-range dependence features output by the time-frequency transformer network.

[0083] The time-frequency transformer network is used to transform the audio waveform of the audio data into a time-frequency spectrum and extract long-range dependence features through the transformer. The long-range dependence features can reflect the complex associations across time or frequency in the audio data.

[0084] Step S1113: Input the short-term audio features and the long-range dependence features into the audio fusion module to obtain the audio features output by the audio fusion module.

[0085] The audio fusion module is used to perform feature fusion on the short-term audio features and the long-range dependence features. Specifically, it can perform linear projection and splicing processing on the short-term audio features and the long-range dependence features in sequence to obtain audio features. Specifically as follows: ; Among them, is the audio feature, is the short-term audio feature, is the long-range dependence feature, is the linear projection function of the short-term audio feature, is the linear projection function of the long-range dependence feature, and Concat is the connection function.

[0086] In this embodiment, by extracting and fusing the short-term audio features and the long-range dependence features of the audio data, both local feature details and global feature information are taken into account, which can help improve the accuracy of the recognition results of the speech separation model.

[0087] In one embodiment, the visual encoder includes a face attribute sub-network, a lip movement sub-network, and a visual fusion module. The above step S112 includes: step S1121, step S1122, and step S1123.

[0088] Step S1121: Input the video data into the face attribute sub-network to obtain the face appearance features output by the face attribute sub-network.

[0089] As an implementation, the face attribute sub-network can be a ResNet-18 network structure (a deep residual network) for extracting face appearance features.

[0090] Step S1122: Input the video data into the lip movement sub-network to obtain the lip movement sequence features output by the lip movement sub-network.

[0091] As an implementation, the lip movement sub-network can be a network structure composed of 3D-Conv (3D-Convolutional Neural Networks), ShuffleNet V2 (an efficient convolutional neural network), and TCN (Temporal Convolutional Network) for extracting lip movement sequence features.

[0092] Step S1123: Input the face appearance features and the lip movement sequence features into the visual fusion module to obtain the facial visual features output by the visual fusion module.

[0093] As an implementation, the visual fusion module can be a transformer network or a small attention network for evaluating the visual frame quality and outputting facial visual features. Specifically, it performs weighted fusion on the face appearance features and the lip movement sequence features to obtain the facial visual features. Specifically as follows: ; Where is the facial visual feature, is the face appearance feature, is the weight coefficient of the face appearance feature, is the lip movement sequence feature, is the weight coefficient of the lip movement sequence feature.

[0094] In this embodiment, the face appearance features and the lip movement sequence features of the video data are extracted and fused to obtain the facial visual features. By fusing multiple facial features related to speech separation for speech separation, it can help improve the accuracy of the recognition result of the speech separation model.

[0095] In one embodiment, the above step S113 includes: step S1131 and step S1132.

[0096] Step S1131, input the audio feature and the facial visual feature into the audio-visual fusion module, and dynamically fuse the audio feature and the facial visual feature through the audio-visual fusion module to obtain the current audio-visual fusion feature.

[0097] The audio-visual fusion module may include a Dynamic Cross-modal Gating (DCG) module, which is used to determine the cross-modal fusion weight based on the reliability of the audio feature and the reliability of the facial visual feature obtained by dynamic evaluation. Specifically as follows: ; Among them, is the cross-modal fusion weight, is the reliability of the audio feature, is the reliability of the facial visual feature, σ is the Sigmoid function, 、 and b are learnable parameters. Among them, is used to measure the quality of the audio signal at the current moment or the credibility of the output of the audio branch, and generally takes values between [0,1]. Connect a small MLP (Multi-Layer Perceptron) at the end of the U-Net / transformer audio branch, input its feature vector, and after two fully connected layers and the Sigmoid activation function, directly predict a scalar, which is . is used to measure the credibility of the output of the visual branch (such as face / lip movement, etc.) at the current moment, and is also between [0,1]. Similarly, a small MLP can be added at the end of the visual feature branch to generate .

[0098] Next, based on the cross-modal fusion weight, dynamically fuse the audio feature and the facial visual feature to obtain the current audio-visual fusion feature. Specifically as follows: ; Among them, is the current audio-visual fusion feature, is the audio feature, is the facial visual feature. represents concatenating the audio feature and the facial visual feature, and then using several layers of MLP to learn the fused result. represents making a linear projection of the audio feature to keep the vector dimension unchanged or mapping it to the target dimension.

[0099] Step S1132: Based on the current audio-visual fusion feature, weighted fusion of historical audio-visual fusion features is performed from a multi-level feature memory bank to obtain an audio-visual fusion feature; wherein, the historical audio-visual fusion features include one or more of short-term memory audio-visual fusion features, long-term memory audio-visual fusion features, and global memory audio-visual fusion features.

[0100] The multi-level feature memory bank includes one or more of short-term memory audio-visual fusion features, long-term memory audio-visual fusion features, and global memory audio-visual fusion features. Therefore, the weighted fusion historical audio fusion features can also be one or more of them.

[0101] Among them, the short-term memory audio-visual fusion feature can be the audio-visual fusion feature generated in real time in the recent several windows, the long-term memory audio-visual fusion feature can be the audio-visual fusion feature of the target object accumulated in the current session, and the global memory audio-visual fusion feature can be the audio-visual fusion feature of the target object across multiple sessions.

[0102] It can be understood that in the case of streaming processing, the historical audio-visual fusion features can include one or more of short-term memory audio-visual fusion features, long-term memory audio-visual fusion features, and global memory audio-visual fusion features. In the case of non-streaming processing, the historical audio-visual fusion features include global memory audio-visual fusion features.

[0103] The audio-visual fusion module may also include a Dynamic Window Fusion (DWF) module for performing cross-window weighted fusion processing on historical audio-visual fusion features and current audio-visual fusion features. Specifically, historical window features can be dynamically integrated based on the self-attention mechanism, high-quality information in multiple historical windows can be adaptively selected, and cross-window fusion can be achieved through a transformer. The specific weighted fusion process is as follows: ; Among them, is the audio-visual fusion feature, is the current audio-visual fusion feature, is the short-term memory audio-visual fusion feature, is the long-term memory audio-visual fusion feature, is the global memory audio-visual fusion feature.

[0104] In this embodiment, by performing cross-window weighted fusion processing on the current audio-visual fusion feature and the historical audio-visual fusion feature, high-quality audio-visual fusion features can be fused, thereby further improving the accuracy of the recognition result of the speech separation model.

[0105] In one embodiment, the voice separation model is trained based on sample mixed audio, sample face, and the audio signal annotation result of the sample face.

[0106] As an implementation manner, the sample mixed audio is audio that mixes multiple human voices and noises.

[0107] In one embodiment, the voice separation model is trained using a preset loss function, and the preset loss function is determined based on voice separation loss, speaker identification loss, and visual quality assessment loss. Among them, the voice separation loss is determined based on the target component and the error component, the speaker identification loss is determined based on the prediction probability and the true label, and the visual quality assessment loss is determined based on the true visual reliability and the predicted visual reliability.

[0108] In a specific embodiment, the preset loss function is determined based on the following formula: ; Among them, is the total loss, is the voice separation loss, is the speaker identification loss, is the visual quality assessment loss, and α and β are two hyperparameters used to balance the weights of each sub-loss in the overall optimization process. They need to be selected through small-scale search (such as grid search or Bayesian optimization) before training so that the final model achieves an optimal balance among voice separation quality, speaker discrimination ability, and visual fusion rationality.

[0109] Among them, the voice separation loss is determined based on the following formula: ; Among them, is the estimated signal, and s is the reference signal (i.e., the clean signal); is the target component, which refers to the part that coincides as much as possible with the reference signal in the direction of the reference signal and measures how much of the estimated signal is truly aligned with the reference signal; is the error component, which refers to the part of the estimated signal that is not aligned with the reference signal and includes residual noise, distortion, or content other than the source signal. is determined based on the following formula: .

[0110] The speaker identification loss is determined based on the following formula: ; Among them, C is the total number of speaker categories, is an indicator variable for the true label on the c-th class, usually using one-hot encoding. If the sample belongs to the speaker category c, = 1, otherwise = 0; is the predicted probability that the voice separation model assigns to the sample belonging to the c-th class, usually normalized by Softmax.

[0111] The visual quality assessment loss is determined based on the following formula: ; where is the true visual reliability, is the predicted visual reliability.

[0112] The above preset loss function combines the voice separation loss , the speaker recognition loss and the visual quality assessment loss , which can better urge the model to optimize the processing results, thereby further improving the processing effect of the voice separation model.

[0113] In one embodiment, the training process of the voice separation model is as follows: 1) Phase 1: Pre-training phase.

[0114] Train the audio branch (i.e., the audio encoder) and the visual branch (i.e., the visual encoder) separately. Among them, for the audio branch, train the U-Net encoder and the time-frequency transformer network separately, and the voice separation loss can be used as the target during training; for the visual branch, the visual fusion module can be trained separately using the face lip movement database, and the classification accuracy can be used as the target during training.

[0115] 2) Phase 2: Cross-modal Fusion Pre-training.

[0116] Fix the audio encoding and visual encoders obtained from the training in Phase 1, train the cross-modal dynamic gating fusion module, and optimize the parameters of the cross-modal dynamic gating fusion module. The voice separation loss can be used as the main target during training.

[0117] 3) Phase 3: Multi-level Memory Training.

[0118] Fix the network parameters of the cross-modal dynamic gating fusion module obtained in the second stage of training, train the short-term, long-term, and global memory update mechanisms, and the supervision objectives during training include the speaker recognition loss 。

[0119] 4) Stage 4: Dynamic Cross-Window Fusion Optimization.

[0120] Use long-sequence speech data to optimize the dynamic cross-window fusion module and further enhance the cross-window feature integration ability.

[0121] 5) Overall Fine-tuning Stage.

[0122] Finally, unfreeze all modules and perform end-to-end fine-tuning with the comprehensive loss. Among them, the comprehensive loss includes the speech separation loss 、the speaker recognition loss and the visual quality assessment loss , specifically, perform fine-tuning through a preset loss function.

[0123] Based on any of the above embodiments, step S11 may further include: step S116, step S117, step S118, and step S119.

[0124] Step S116, obtain real-time audio-visual data including the at least one face information, where the real-time audio-visual data includes real-time audio data and real-time video data.

[0125] Considering the latency issue, in this embodiment, when using the speech separation model to determine the audio signal corresponding to the face information, a streaming processing method is adopted, that is, during the process of obtaining the audio-visual data, the real-time obtained audio-visual data (referred to as real-time audio-visual data) is processed, rather than processing the obtained audio-visual data after obtaining all the audio-visual data. Exemplarily, assuming that the recording duration of the audio-visual data is set to 3s, in the prior art, generally the recorded 3s audio-visual data is input into the speech separation model for processing, while in this embodiment, the streaming processing method is adopted, and during the recording process of the audio-visual data, every 400ms of the recorded real-time audio-visual data is input into the speech separation model for processing each time. Through the streaming processing, it can be processed immediately when the audio-visual data is initially generated, thus realizing the fast processing of the audio-visual data and greatly reducing the latency.

[0126] Step S117, perform block processing on the real-time audio data through a time-sliding window and perform frame processing on the real-time video data.

[0127] During streaming processing, the input of the voice separation model is generally processed in chunks (blocks) with a time window as the unit. Therefore, after obtaining real-time audio data, the real-time audio data can be chunked, and correspondingly, the real-time video data obtained synchronously can be framed.

[0128] Furthermore, considering the latency issue and user experience, the audio length of chunking processing can be determined based on the key metric of streaming processing efficiency - end-to-end latency.

[0129] Specifically, the end-to-end latency consists of the following items: ; Among them, is the end-to-end latency, is the time of each chunk of audio, is the forward calculation time of the model. Considering the user experience issue, should be controlled within <500ms, should be controlled within <100ms, correspondingly, can be set to 400ms.

[0130] Exemplarily, the common audio sampling rate is 16kHz; the length of each chunk of audio (chunk): 6400, that is, the audio length of each chunk of audio processed each time is 400ms; the sliding step (hop) is regarded as 3200, and it slides backward 3200 sampling points each time (i.e., 200ms); the overlap ratio is 50%, that is, the current chunk of audio overlaps with the previous chunk of audio by half. Correspondingly, the framed processing data of the real-time video data is as follows: Assuming the video frame rate is 25FPS, corresponding to one frame every 40ms, then each 400ms chunk of audio contains about 10 frames of images. When processing the nth chunk of audio, the real-time video data in the corresponding time period is synchronously collected, framed to obtain the video frame sequence , and sent into the model at the same time.

[0131] Furthermore, for the first frame, we can use the "half-padding" method to occupy the position of the second half of the samples that have not arrived with "useless" or zero signals first. In this way, when the real 200ms real-time audio data arrives, a 400ms chunk of audio can be immediately filled and input into the voice separation model, thereby further reducing the end-to-end latency to <300ms.

[0132] Furthermore, a sliding window (overlap) can also be used to further reduce the latency. Specifically, the audio buffer is updated continuously with the sliding step. That is, in the nth processing cycle, read from the input buffer from the time point to The audio data forms the chunked audio for the current processing. . Among them, is the sliding step size, and is the length (number of samples) of the audio chunk processed each time.

[0133] Step S118: Input the real-time audio data after chunk processing and the real-time video data after frame processing into the voice separation model to obtain at least one chunked audio signal corresponding to the face information output by the voice separation model.

[0134] Then, input the real-time audio data after chunk processing and the real-time video data after frame processing into the voice separation model, and use the voice separation model to determine the chunked audio signal corresponding to the face information.

[0135] Step S119: Stitch the chunked audio signals through a smoothing window to obtain the audio signal corresponding to at least one face information.

[0136] The time-domain output signal generated for each chunked audio is denoted as , and is processed using a smoothing window (such as a Hanning window): ; Among them, is the smoothed signal after applying the window function to the in-block signal , is the value of the smoothing window function at time point t.

[0137] Add and stitch with the overlapping section of the previous chunk to ensure seamless connection. Specifically as follows: ; Among them, output represents the processed result, n represents the nth processing cycle, is the smoothed signal after applying the window function to the in-block signal , is the sliding step size.

[0138] In this embodiment, through streaming processing, fast processing of audio and video data can be achieved, greatly reducing latency and improving the user experience.

[0139] Based on any of the above embodiments, step S130 may include: step S131 and step S132.

[0140] Step S131: Identify the audio signals of at least one object and the voiceprint features of each audio signal from the audio to be processed.

[0141] Separate the audio to be processed to obtain the separated audio signal. It should be understood that the separated audio signal includes the audio signals of at least one object. In the scenario of multiple people speaking, the separated audio signal includes the audio signals of multiple objects; in the scenario of a single person speaking, the separated audio signal includes the audio signal of one object.

[0142] Then, extract the voiceprint features in each recognized audio signal to obtain the voiceprint features of each audio signal. Among them, the voiceprint features include but are not limited to: MFCC, LPCC, spectral envelope, i-vector, x-vector. The extraction methods of various voiceprint features can refer to the prior art and will not be elaborated here.

[0143] Step S132, based on the voiceprint features of each audio signal and the target voiceprint feature of the target object, obtain the target audio signal of the target object from the audio signals of the at least one object.

[0144] Compare the target voiceprint feature of the target object with the voiceprint features of each recognized audio signal respectively. Specifically, the similarity comparison can be performed based on cosine similarity or probabilistic linear discriminant analysis.

[0145] Then, according to the similarity comparison result, obtain the target audio signal of the target object from the recognized audio signals.

[0146] Exemplarily, if cosine similarity is used as the similarity comparison index, it can be set that when the cosine similarity is greater than or equal to the preset threshold, it is determined that the compared voiceprint features are of the same object.

[0147] In this embodiment, by recognizing the audio signals and their voiceprint features in the audio to be processed, and then comparing the target voiceprint feature with the recognized voiceprint features, the target audio signal of the target object can be determined according to the comparison result. Through the above method, the target audio signal corresponding to the target object can be accurately determined from the audio to be processed.

[0148] Furthermore, based on the above embodiment, after step S131, the audio processing method further includes: Step S141, if it is detected that at least one of the voiceprint features in the voiceprint features of each audio signal is not in the association relationship library, update the association relationship library based on the detection result.

[0149] After identifying the audio signals of at least one object from the audio to be processed, it is possible to further detect whether the voiceprint features of the identified audio signals are all in the association relationship library to determine whether there is a new speaker joining the scene. The specific detection method can also be through similarity comparison. For example, if it is detected that the cosine similarity between a certain voiceprint feature and all the voiceprint features stored in the association relationship library is less than the preset threshold, it is determined that the voiceprint feature is not in the association relationship library.

[0150] If it is detected that at least one of the identified voiceprint features is not in the association relationship library, the association relationship library is updated. Specifically, obtain the face information corresponding to the voiceprint feature that is not in the association relationship library. The obtaining method can be: obtain the current scene image, detect the faces in the current scene image, compare the detected face information with the face information in the association relationship library, and determine the face information of the new speaker based on the detection result.

[0151] Then, add the face information and the voiceprint feature to the association relationship library to expand and update the association relationship library.

[0152] Through the above method, the dynamic update of the association relationship library can be realized to support the joining of new speakers, which helps to support the switching of new speakers in complex scenes and the rapid recognition of their audio signals.

[0153] Based on any of the above embodiments, after step S130, the audio processing method may further include: Step S142, perform intensity change processing on the target audio signal of the target object in the audio to be processed.

[0154] In this embodiment, considering that the background noise and / or the speech of non-target objects have a large interference, after identifying the target audio signal of the target object, it is possible to further perform intensity change processing on the target audio signal of the target object in the audio to be processed. Among them, the intensity change processing methods can include: 1) only retain the target audio signal of the target object, and eliminate the audio signals of other objects and the audio signal of the background noise; 2) emphasize the target audio signal of the target object, and weakly retain the audio signals of other objects and the audio signal of the background noise.

[0155] Through the intensity change processing, it is possible to reduce or even eliminate the volume of non-target objects and background noise in the audio output based on the intensity change processed target audio signal, so as to highlight the volume of the target object, facilitate the user to better hear the voice of the target object, reduce the interference of other interfering sounds, and thus improve the user experience.

[0156] In a specific embodiment, the audio to be processed includes audio signals of at least two objects. By setting weight values for the audio signals of the at least two objects, intensity variation processing is performed on the target audio signal of the target object.

[0157] Further, considering that in a multi-person scenario, the audio to be processed may include audio signals of multiple objects, intensity variation processing can be performed on the target audio signal of the target object by setting weight values for the audio signals of at least two objects.

[0158] It can be understood that if the audio to be processed includes only the audio signal of one object, that is, the target audio signal of the target object, no processing is required. If the audio to be processed includes only the audio signal of one object but also includes background noise, the background noise can be removed or weakly retained. Similarly, this can also be achieved by setting weight values. Specifically, different weight values can be set for the target audio signal of the target object and the audio signal of the background noise.

[0159] In a specific embodiment, step S142 may include: When setting the weight value of the audio signal of the non-target object to 0, separating the target audio signal of the target object from the audio signal to be processed corresponding to the audio to be processed; or, When setting the weight value of the audio signal of the target object to be greater than the weight value of the audio signal of the non-target object, enhancing the target audio signal of the target object in the audio signal to be processed corresponding to the audio to be processed; wherein, the non-target object is other objects except the target object among the at least two objects.

[0160] It should be noted that in this embodiment, the influence of background noise is not considered, and only the scenario of multiple people speaking is considered.

[0161] As an implementation manner, the weight value of the audio signal of the non-target object can be set to 0 to eliminate the interfering sounds of other objects. The weight value of the audio signal of the target object can be set to 1. In this case, the target audio signal of the target object can be separated from the audio signal to be processed corresponding to the audio to be processed.

[0162] As an implementation, the weight value of the audio signal of the target object can be set to be greater than the weight value of the audio signal of the non-target object to emphasize the voice of the target object and weakly retain the voice of the non-target object. Exemplarily, the weight value of the audio signal of the target object can be set to 1, and the weight value of the audio signal of the non-target object is in the range greater than 0 and less than 1. In this case, in the to-be-processed audio signal corresponding to the to-be-processed audio, the target audio signal of the target object is enhanced. Specifically, each sampling point of the target audio signal and the non-target audio signal is weighted and summed to obtain an output audio signal, where the non-target audio signal is the audio signal of other objects except the target object in the to-be-processed audio.

[0163] Further, after the intensity change processing, the target audio signal after the intensity change processing can be converted into an output audio and output, so that the user can obtain the audio of the target object.

[0164] Based on any of the above embodiments, after step S130, the audio processing method may further include: Step S143, performing type conversion processing on the target audio signal of the target object.

[0165] Considering the specific application scenario and user requirements, after the target audio signal of the target object is recognized, the audio signal of the target object can be further subjected to type conversion processing. Among them, the type conversion processing includes but is not limited to: translation, audio-to-text conversion, and audio-to-sign language conversion, etc.

[0166] In this embodiment, through the type conversion processing, the audio signal can be converted into different types, and then the speech content of the target object can be displayed in different forms, thereby improving the user experience.

[0167] In a specific embodiment, step S143 may include any one of the following steps S1431, step S1432, and step S1433.

[0168] Step S1431, converting the target audio signal of the target object into an audio signal in another language.

[0169] As an implementation, the target audio signal can be converted into an audio signal in another language to translate the speech content of the target object.

[0170] After obtaining the target audio signal of the target object, further detect whether the language of the audio signal is a preset language, where the preset language is the preferred language pre-set by the user. For example, Chinese. The detection method of the language can be: by analyzing the acoustic features of the target audio, and determining whether it is a preset voice type according to the analysis result. Or, use a deep learning model to detect the language of the target audio signal, obtain the detection result, and then determine whether it is a preset language according to the detection result. Among them, the deep learning model includes but is not limited to: Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and transformer model.

[0171] When it is detected that the language of the target audio signal is not the preset language, it can be translated and converted into an audio signal in the preset language, and then the audio in the preset language is output. Through the conversion of the language, it is convenient for the user to quickly and effectively obtain the speech content of the target object.

[0172] Step S1432, convert the target audio signal of the target object into text.

[0173] As an implementation manner, the target audio signal of the target object can be converted into text, that is, the speech content of the target object is displayed in the form of text.

[0174] Specifically, a speech recognition model can be used to perform speech recognition on the target audio signal of the target object to obtain the recognition result, and then it is displayed in the form of text.

[0175] Step S1433, convert the target audio signal of the target object into sign language.

[0176] As an implementation manner, the target audio signal of the target object can be converted into sign language, that is, the speech content of the target object is displayed in the form of sign language. In this case, it is convenient for the deaf to quickly and effectively obtain the speech content of the target object.

[0177] Specifically, a speech recognition model can be used to perform speech recognition on the target audio signal of the target object to obtain the recognition result, and then the recognition result is converted into sign language and displayed.

[0178] Based on any of the above embodiments, step S110 may include: In response to the user's interaction operation, obtain the face information of the target object, and collect the audio to be processed through the microphone; Among them, the interaction operation includes any one of a gesture interaction operation, a head movement interaction operation, and an eye movement interaction operation.

[0179] In this embodiment, the process of obtaining the audio to be processed and the face information of the target object may be as follows: The user performs specific interaction operations, such as gesture interaction, head movement interaction, or eye movement interaction. The specific manner of gesture interaction can be to specify the direction of the target object with the hand. The specific manner of head movement interaction can be to turn the head to look at the direction of the target object. The specific manner of eye movement interaction can be to turn the eyes to look at the direction of the target object.

[0180] Correspondingly, after the XR device detects the user's interaction operation, it can obtain the face information of the target object based on the user's interaction operation, and at the same time collect the audio to be processed through the microphone.

[0181] In the above manner, based on the user's interaction operation, the face information of the target object can be accurately and quickly obtained, and at the same time the audio to be processed is collected, which is convenient for subsequent further audio processing.

[0182] Based on any of the above embodiments, before step S110, the audio processing method may further include: step S01 and step S02.

[0183] Step S01, perform a calibration process on the XR device.

[0184] In this embodiment, when recording audio and video data through the XR device, it is necessary to first perform a calibration process on the XR device. Through the calibration process, the real-world space and the camera imaging space can be aligned to ensure the geometric consistency of the subsequent collected data.

[0185] Taking an AR glasses as an example, its calibration process aims to accurately align the digital virtual image with the real scene, so that the virtual object presents the correct position, scale, and orientation in the physical world. The specific calibration methods include but are not limited to the following methods: (1) User perspective calibration: Calibration is performed through the wearer's perspective, that is, the wearer actively aligns the virtual and real reference objects to calculate the calibration parameters. For example, the commonly used SPAAM (Single Point Active Alignment Method) requires the wearer to repeatedly coincide the virtual aiming mark with the feature points in the real world. Correspondingly, the AR glasses will obtain corresponding multiple sets of virtual and real coordinates, so as to calculate the conversion relationship between the real world and the virtual display coordinates, and calibrate the AR glasses based on this conversion relationship. This calibration based on the wearer's manual alignment ensures the correct perspective of the system starting from the user's eyes, but requires the user to participate in multiple alignment operations.

[0186] (2) Optical-Image Alignment Calibration: For optically see-through AR glasses, the alignment relationship between the optical view and the digital image needs to be calibrated. The core is to establish a mapping between the tracking sensor coordinate system and the glasses display coordinate system: by measuring the corresponding positions of a series of spatial points in the real-world tracking coordinates (such as camera coordinates) and the glasses screen pixel coordinates, the projection matrix (usually a 3×4 matrix G) is solved, so that the virtual image can accurately coincide with the real object in the field of view. After completing the optical-image alignment, the system can convert any real-world three-dimensional point into a two-dimensional projection on the screen according to this mapping, and realize the superposition of virtual information on the corresponding real-world position (for example, stably superimposing text or 3D models on the surface of real objects).

[0187] (3) Spatial Coordinate System Mapping Calibration: The corresponding relationship between various spatial coordinate systems is established through mathematical transformation. This usually includes the mapping between the camera coordinate system, the glasses device coordinate system, the world coordinate system, and even the user's eye coordinate system. For example, calibrating the position relationship between the tracking camera and the eyeball / display can determine the conversion from the world coordinates to the user's field of view. Essentially, this is to solve the external parameter (rotation and translation) transformation matrix between different sensor / device coordinate systems, so that all virtual and real objects are in the same unified spatial reference. Through the coordinate system mapping calibration, the AR glasses system can correctly convert the virtual content from its virtual coordinates to the real-world scene coordinates and keep them aligned under the cooperation of multiple devices / sensors.

[0188] Step S02, based on the calibrated XR device, record the audio-visual data.

[0189] Then, based on the calibrated XR device, record the audio-visual data for constructing the correlation relationship library of voiceprint features and face information.

[0190] In this embodiment, by calibrating the XR device, the geometric consistency between the real-world space and the camera imaging space of the XR device can be ensured. Combining face detection and synchronous acquisition of audio-visual data, the spatio-temporal alignment of vision and audio-visual data is achieved. That is, the fusion accuracy of multi-modal data is optimized, providing reliable spatial constraints for subsequent voiceprint matching.

[0191] Next, the audio processing device provided by the present invention will be described. The audio processing device described below can be mutually referred to the audio processing method described above.

[0192] Figure 4 is a schematic structural diagram of the audio processing device provided by the present invention, as Figure 4 shown, the device includes a first acquisition module 410, a second acquisition module 420, and an audio recognition module 430; where: The first acquisition module 410 is configured to acquire the audio to be processed and the face information of the target object; The second acquisition module 420 is configured to acquire the target voiceprint feature of the target object based on the pre-constructed association relationship library between the voiceprint feature and the face information, and the face information of the target object; wherein, the association relationship library between the voiceprint feature and the face information is constructed based on the audio-visual data acquired by the voice separation model and the XR device; The audio recognition module 430 is configured to recognize the target audio signal of the target object from the audio to be processed based on the target voiceprint feature of the target object.

[0193] The audio processing device provided by the present invention constructs an association relationship library between the voiceprint feature and the face information through the audio-visual data acquired by the voice separation model and the XR device. When the audio to be processed and the face information of the target object are acquired, according to the face information of the target object, the target voiceprint feature corresponding to the target object is acquired from the pre-constructed association relationship library; and then based on the target voiceprint feature, the target audio signal of the target object is recognized from the audio to be processed. The present invention does not need to improve the hardware of the XR device, and only needs to use the camera and microphone necessary for the XR device to acquire the audio-visual data, and then uses the voice separation model to determine the voices of each speaker in the audio-visual data, and then constructs an association relationship library between the voiceprint feature and the face information. In subsequent audio processing scenarios, by directly calling the association relationship library to acquire the target voiceprint feature of the target object, the target audio signal of the target object can be accurately recognized from the audio to be processed based on the target voiceprint feature. Through the above method, it can completely identify the audio signal of the target object from various interference sources such as background noise, reverberation, and interfering human voices. At the same time, the present invention constructs an association relationship library in advance. In the audio processing scenario, only the audio data needs to be processed. Compared with processing the audio-visual data using the voice separation model each time, the present invention can also save computing power and reduce power consumption.

[0194] It should be noted here that the above audio processing device provided by the embodiments of the present invention can implement all the method steps implemented by the above audio processing method embodiments, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described in this embodiment.

[0195] Figure 5 Illustrates a schematic diagram of the physical structure of an XR device, such as Figure 5As shown in the figure, the XR device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute an audio processing method, and this method includes: Obtain the audio to be processed and the face information of the target object; Based on the pre-constructed association relationship library between voiceprint features and face information, and the face information of the target object, obtain the target voiceprint feature of the target object; wherein, the association relationship library between voiceprint features and face information is constructed based on the voice separation model and the audio-visual data obtained by the XR device; Based on the target voiceprint feature of the target object, identify the target audio signal of the target object from the audio to be processed.

[0196] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0197] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the audio processing method provided by the above-mentioned various methods. This method includes: Obtain the audio to be processed and the face information of the target object; Based on the pre-constructed association relationship library between voiceprint features and face information, and the face information of the target object, obtain the target voiceprint feature of the target object; wherein, the association relationship library between voiceprint features and face information is constructed based on the voice separation model and the audio-visual data obtained by the XR device; Identify the target audio signal of the target object from the audio to be processed based on the target voiceprint feature of the target object.

[0198] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the audio processing method provided by the above-mentioned various methods, and the method includes: Obtain the audio to be processed and the face information of the target object; Based on the pre-constructed association relationship library between voiceprint features and face information, and the face information of the target object, obtain the target voiceprint feature of the target object; wherein, the association relationship library between voiceprint features and face information is constructed based on the audio-visual data obtained by the voice separation model and the XR device; Identify the target audio signal of the target object from the audio to be processed based on the target voiceprint feature of the target object.

[0199] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0200] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An audio processing method, characterized in that, Applied to an extended reality (XR) device, including: Obtain the audio to be processed and the face information of the target object; Based on the pre-constructed association relationship library between voiceprint features and face information, and the face information of the target object, obtain the target voiceprint feature of the target object; wherein, the association relationship library between voiceprint features and face information is constructed based on a voice separation model and the audio-visual data obtained by the XR device; Based on the target voiceprint feature of the target object, identify the target audio signal of the target object from the audio to be processed.

2. The audio processing method according to claim 1, wherein The construction steps of the association relationship library between voiceprint features and face information include: Obtain a scene image, perform face detection on the scene image, and obtain at least one piece of face information; Obtain audio-visual data including the at least one piece of face information, and use the voice separation model to process the audio-visual data to determine the audio signal corresponding to the at least one piece of face information; Based on the audio signal corresponding to the at least one piece of face information, obtain the voiceprint feature corresponding to the at least one piece of face information; Based on the voiceprint feature corresponding to the at least one piece of face information, construct the association relationship library between voiceprint features and face information.

3. The audio processing method according to claim 2, wherein The voice separation model includes an audio encoder, a visual encoder, an audio-visual fusion module, a mask estimation network, and an audio decoder. The audio-visual data includes audio data and video data. Using the voice separation model to process the audio-visual data to determine the audio signal corresponding to the at least one piece of face information includes: Input the audio data into the audio encoder to obtain the audio feature output by the audio encoder; Input the video data into the visual encoder to obtain the facial visual feature output by the visual encoder; Input the audio feature and the facial visual feature into the audio-visual fusion module to obtain the audio-visual fusion feature output by the audio-visual fusion module; Input the audio-visual fusion feature into the mask estimation network to obtain the time-frequency mask output by the mask estimation network; Input the time-frequency mask into the audio decoder to obtain the audio signal corresponding to the at least one piece of face information output by the audio decoder.

4. The audio processing method according to claim 3, wherein The audio encoder includes a U-Net encoder, a time-frequency transformer network, and an audio fusion module. The visual encoder includes a face attribute sub-network, a lip movement sub-network, and a visual fusion module. Inputting the audio data into the audio encoder to obtain the audio feature output by the audio encoder includes: Input the audio data into the U-Net encoder to obtain the short-term audio feature output by the U-Net encoder; Input the audio data into the time-frequency transformer network to obtain the long-distance dependence feature output by the time-frequency transformer network; Input the short-term audio feature and the long-distance dependence feature into the audio fusion module to obtain the audio feature output by the audio fusion module; Inputting the video data into the visual encoder to obtain the facial visual features output by the visual encoder includes: Inputting the video data into the face attribute sub-network to obtain the face appearance features output by the face attribute sub-network; Inputting the video data into the lip movement sub-network to obtain the lip movement sequence features output by the lip movement sub-network; Inputting the face appearance features and the lip movement sequence features into the visual fusion module to obtain the facial visual features output by the visual fusion module.

5. The audio processing method according to claim 3, wherein Inputting the audio features and the facial visual features into the audio-visual fusion module to obtain the audio-visual fusion features output by the audio-visual fusion module includes: Inputting the audio features and the facial visual features into the audio-visual fusion module, and dynamically fusing the audio features and the facial visual features through the audio-visual fusion module to obtain the current audio-visual fusion features; Weightedly fusing the historical audio-visual fusion features from the multi-level feature memory bank based on the current audio-visual fusion features to obtain the audio-visual fusion features; wherein, the historical audio-visual fusion features include one or more of short-term memory audio-visual fusion features, long-term memory audio-visual fusion features, and global memory audio-visual fusion features.

6. The audio processing method according to any one of claims 2 to 5, characterized in that Obtaining the audio-visual data including the at least one face information, and processing the audio-visual data by using the voice separation model to determine the audio signals corresponding to the at least one face information includes: Obtaining the real-time audio-visual data including the at least one face information, where the real-time audio-visual data includes real-time audio data and real-time video data; Performing block processing on the real-time audio data through a time sliding window, and performing frame processing on the real-time video data; Inputting the block-processed real-time audio data and the frame-processed real-time video data into the voice separation model to obtain the block audio signals corresponding to the at least one face information output by the voice separation model; Stitching the block audio signals through a smoothing window to obtain the audio signals corresponding to the at least one face information.

7. The audio processing method according to any one of claims 1 to 5, characterized in that, Based on the target voiceprint feature of the target object, identifying the target audio signal of the target object from the audio to be processed includes: Identifying the audio signals of at least one object and the voiceprint features of each audio signal from the audio to be processed; Based on the voiceprint features of each audio signal and the target voiceprint feature of the target object, obtaining the target audio signal of the target object from the audio signals of the at least one object; After identifying the audio signals of at least one object and the voiceprint features of each audio signal from the audio to be processed, further includes: If it is detected that at least one of the voiceprint features in the voiceprint features of each audio signal is not in the association relationship library, updating the association relationship library based on the detection result.

8. The audio processing method according to any one of claims 1 to 5, characterized in that, After identifying the target audio signal of the target object from the audio to be processed based on the target voiceprint feature of the target object, further includes: Perform intensity change processing on the target audio signal of the target object in the audio to be processed; or, Perform type conversion processing on the target audio signal of the target object.

9. An audio processing device, characterized in that, It includes: A first acquisition module for acquiring the audio to be processed and the face information of the target object; A second acquisition module for acquiring the target voiceprint feature of the target object based on the pre-constructed association relationship library between the voiceprint feature and the face information, and the face information of the target object; wherein, the association relationship library between the voiceprint feature and the face information is constructed based on the audio-visual data acquired by the voice separation model and the XR device; An audio recognition module for recognizing the target audio signal of the target object from the audio to be processed based on the target voiceprint feature of the target object.

10. An XR device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the audio processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Audio processing method and device, terminal and storage medium

    CN112820300A

  • Image fusion method, video image processing equipment and computer readable storage medium

    CN114511481A

  • Target detection network training method and device, processing equipment and storage medium

    CN114596519A

  • Multi-modal sentiment analysis method based on multi-task learning and stacked cross-modal fusion

    CN114694076A

  • Voice operation method and device of equipment and electronic equipment

    CN115050375A