Audio-visual speech separation method and apparatus, electronic device, and storage medium

By using a multimodal separation network to perform multiple feature fusions on video and audio features, the problem of noise interference in speech separation is solved, achieving higher speech separation accuracy and noise suppression effect.

CN116863537BActive Publication Date: 2026-02-03TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310816352.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-04
Publication Date
2026-02-03
Estimated Expiration
2043-07-04

AI Technical Summary

Technical Problem

Existing technologies suffer from severe noise interference during speech separation, leading to a decline in the performance of the separation system, especially in noisy environments where it is difficult to accurately extract the target speech.

Method used

By acquiring image frame sequences and mixed audio from video information, a multimodal separation network is used for feature fusion, including an auditory subnetwork, a visual subnetwork, a top module, a middle module, and a bottom module. Multiple feature fusions are performed to generate a sound mask for the target object, thereby improving the accuracy of speech separation.

Benefits of technology

It enhances intramodal contextual information, improves audio-visual separation performance, obtains more accurate audio separation results, and reduces the impact of noise interference on speech separation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116863537B_ABST
    Figure CN116863537B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio-visual speech separation method and device, electronic equipment and storage medium, video information including target object sound and at least one reference object sound is acquired, and an image frame sequence composed of target object lip image frames in the video information and mixed audio are extracted, and the video feature and the audio feature corresponding to the target object are obtained by encoding respectively. The video feature and the audio feature are input into a trained multi-modal separation network, and the sound mask corresponding to the target object is obtained after multiple feature fusions. The multi-modal separation network includes a top module, a middle module and a bottom module for three times of feature fusion. The target audio recording the target object sound is determined according to the sound mask and the audio feature. The present disclosure fuses the information of two levels of vision and hearing multiple times through three modules, enhances the intra-modal context information, improves the audio-visual separation performance, and obtains an accurate audio separation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and storage medium for separating audiovisual speech. Background Technology

[0002] The real world environment contains various forms of information and complex interactions between them. The human brain can effectively integrate different modalities to enhance perception of the surrounding environment. In our daily lives, audio and visual signals are among the main means of information transmission, providing rich clues for people to obtain valuable information in noisy environments. This innate ability is known as the "cocktail party effect," or "speech separation" in computer science. Speech separation has broad research implications, such as helping individuals with hearing impairments, enhancing the auditory experience of wearable devices, and improving transcription recognition accuracy in video conferencing. Researchers have previously attempted to address the cocktail party effect using separate audio streams. However, when speech is interfered with by noise, the performance of the separation system can degrade significantly. Summary of the Invention

[0003] In view of this, this disclosure proposes an audiovisual speech separation method, apparatus, electronic device, and storage medium, aiming to reduce noise interference in the speech separation process and improve the accuracy of speech separation results.

[0004] According to a first aspect of this disclosure, an audiovisual speech separation method is provided, the method comprising:

[0005] Acquire video information including the sound of the target object and the sound of at least one reference object;

[0006] Extract the image frame sequence and mixed audio from the video information, wherein the image frame sequence is a sequence composed of image frames of the lip area of ​​the target object;

[0007] The mixed audio and the image frame sequence are encoded separately to obtain audio features and video features corresponding to the target object;

[0008] The video features and audio features are input into a multimodal separation network trained on the network. After multiple feature fusions, the sound mask corresponding to the target object is obtained. The multimodal separation network includes a top module, a middle module, and a bottom module. The top module is used to perform preliminary fusion of the video features and audio features. The middle module is used to perform further fusion of the video features and audio features at each time scale. The bottom module is used to perform fine-grained feature fusion of the video features and auditory features.

[0009] The target audio for recording the sound of the target object is determined based on the sound mask and the audio features.

[0010] In one possible implementation, the multimodal separation network includes an auditory subnetwork and a visual subnetwork. The process of inputting the video features and the audio features into the trained multimodal separation network, and obtaining the sound mask corresponding to the target object through multiple feature fusions, includes:

[0011] The video features and the audio features are fused an iteratively through the auditory subnetwork, the visual subnetwork, the top module, the middle module, and the bottom module to obtain intermediate audio features and intermediate video features.

[0012] The intermediate audio features are processed iteratively through the auditory sub-network for a second preset number of audio processing steps to obtain the target audio features as the sound mask corresponding to the target object.

[0013] In one possible implementation, the iterative feature fusion of the video features and the audio features through the auditory subnetwork, visual subnetwork, top module, middle module, and bottom module to obtain intermediate audio features and intermediate video features includes:

[0014] The video features and audio features are input into the auditory subnetwork and the visual subnetwork, respectively, to obtain first visual features and first auditory features at different time scales;

[0015] The top module fuses first visual features and first auditory features based on different time scales to obtain modulated second visual features and second auditory features;

[0016] The second visual feature and the second auditory feature are fused by the middle module to obtain the third auditory feature and the third visual feature after re-modulation.

[0017] The bottom module performs feature fusion on the third auditory feature and the third visual feature to obtain the fused fourth auditory feature and fourth visual feature, thus completing one iteration;

[0018] In response to the current iteration number being the first preset number, the iteration process ends and the fourth auditory feature and the fourth visual feature are determined to be the intermediate audio feature and the intermediate video feature, respectively.

[0019] If the current iteration number is not the first preset number, the fourth auditory feature is updated to an audio feature, the fourth visual feature is updated to a video feature, and the iteration is restarted.

[0020] In one possible implementation, the step of inputting the video features and audio features into the auditory subnetwork and the visual subnetwork, respectively, to obtain first visual features and first auditory features at different time scales includes:

[0021] The visual subnetwork downsamples the video features based on different time scales to obtain the first visual feature corresponding to each time scale.

[0022] The auditory subnetwork downsamples the audio features based on different time scales to obtain the first auditory feature corresponding to each time scale.

[0023] In one possible implementation, the top module includes an average normalization layer and a multilayer perceptron. The process of fusing first visual features and first auditory features based on different time scales through the top module to obtain modulated second visual features and second auditory features includes:

[0024] The first visual feature and the first auditory feature at each time scale are processed by the average normalization layer to obtain fused visual features and fused auditory features;

[0025] The fused auditory features are multiplied by the fused visual features processed by the sigma function, and then input into the multilayer perceptron for feature fusion to obtain the second auditory features.

[0026] The fused auditory features, processed by the sigma function, are multiplied by the fused visual features and then input into the multilayer perceptron for feature fusion to obtain the second visual features.

[0027] In one possible implementation, the feature fusion of the second visual feature and the second auditory feature through the central module to obtain the remodulated third auditory feature and third visual feature includes:

[0028] Through the self-attention module and dual-output module in the visual sub-network, a reference visual feature and a third visual feature for each time scale are generated based on the second visual feature and the first visual feature for each time scale.

[0029] The self-attention module in the audio sub-network generates reference auditory features for each time scale based on the second auditory feature and the first auditory feature for each time scale.

[0030] The central module performs feature fusion on the reference visual features and reference auditory features for each time scale to obtain the remodulated third auditory feature.

[0031] In one possible implementation, generating reference visual features and third visual features for each time scale based on the second visual features and the first visual features for each time scale via the self-attention module and dual-output module in the visual sub-network includes:

[0032] The second visual feature is upsampled based on different time scales to obtain the first intermediate visual feature corresponding to each time scale.

[0033] The intermediate visual features and the first intermediate visual features corresponding to each time scale are input into the self-attention module to obtain the second intermediate visual features corresponding to the time scale.

[0034] The second intermediate visual feature corresponding to each time scale is input into the dual output module, and the second intermediate visual feature corresponding to each time scale is output as a reference visual feature, while the third visual feature is output.

[0035] The process of determining the third visual feature includes:

[0036] Upsample the second intermediate visual feature corresponding to the smallest time scale;

[0037] The current upsampled visual features are repeatedly input into the attention module along with the second intermediate visual features corresponding to adjacent and larger time scales in an iterative manner, and the corresponding candidate visual features are then output.

[0038] In response to the fact that the time scale corresponding to the candidate visual feature is not the largest time scale, the candidate visual feature is upsampled;

[0039] In response to the candidate visual feature being the largest time scale, the candidate visual feature is determined to be the third visual feature.

[0040] In one possible implementation, generating reference auditory features for each time scale based on the second auditory feature and the first auditory feature for each time scale via a self-attention module in the audio sub-network includes:

[0041] The second auditory feature is upsampled based on different time scales to obtain the intermediate auditory feature corresponding to each time scale;

[0042] The intermediate auditory features and intermediate auditory features corresponding to each time scale are input into the self-attention module to obtain the reference auditory features corresponding to the time scale.

[0043] In one possible implementation, the central module performs feature fusion on the reference visual features and reference auditory features for each time scale to obtain the remodulated third auditory feature, including:

[0044] The central module performs a dot product of the reference visual features and reference auditory features for each time scale to obtain the feature product for each time scale.

[0045] Upsample the feature product corresponding to the smallest time scale;

[0046] The feature product after upsampling is input into the attention module multiple times in an iterative manner, along with the feature product corresponding to the adjacent and larger time scales. The updated feature product of the adjacent and larger time scales is then output.

[0047] If the time scale corresponding to the updated feature product is not the largest time scale, the feature product is upsampled.

[0048] In response to the updated feature product corresponding to the largest time scale, the feature product is determined to be the third auditory feature and output.

[0049] In one possible implementation, the step of fusing the third auditory feature and the third visual feature through the bottom module to obtain the fused fourth auditory feature and fourth visual feature includes:

[0050] The third auditory feature and the third visual feature are input into the bottom module, and the third auditory feature and the third visual feature are processed by the convolutional layer of the bottom module to obtain convolutional auditory features and convolutional visual features.

[0051] The fourth auditory feature is obtained by multiplying the convolutional auditory feature and the convolutional visual feature processed by the sigma function;

[0052] The fourth visual feature is obtained by multiplying the convolutional auditory feature and the convolutional visual feature after processing with the sigma function.

[0053] In one possible implementation, extracting the image frame sequence and mixed audio from the video information includes:

[0054] The video information is sampled for both image and audio using preset sampling parameters to obtain a sampled image sequence and mixed audio.

[0055] Identify and extract the lip region of the target object in each of the sampled images in the sampled image sequence to obtain an image frame sequence.

[0056] In one possible implementation, determining the target audio for recording the sound of the target object based on the sound mask and the audio features includes:

[0057] The dot product of the sound mask and the audio feature is calculated and input into the audio decoder, which then transcodes the sound of the target object to obtain the target audio.

[0058] According to a second aspect of this disclosure, an audiovisual speech separation device is provided, the device comprising:

[0059] The information acquisition module is used to acquire video information including the sound of the target object and the sound of at least one reference object;

[0060] The content extraction module is used to extract the image frame sequence and mixed audio from the video information, wherein the image frame sequence is a sequence composed of image frames of the lip area of ​​the target object;

[0061] The encoding module is used to encode the mixed audio and the image frame sequence respectively to obtain audio features and video features corresponding to the target object;

[0062] The feature fusion module is used to input the video features and the audio features into the trained multimodal separation network, and obtain the sound mask corresponding to the target object after multiple feature fusions. The multimodal separation network includes a top module, a middle module, and a bottom module. The top module is used to perform preliminary fusion of the video features and the audio features. The middle module is used to perform further fusion of the video features and the audio features at each time scale. The bottom module is used to perform fine-grained feature fusion of the video features and the auditory features.

[0063] An audio determination module is used to determine the target audio for recording the sound of the target object based on the sound mask and the audio features.

[0064] In one possible implementation, the multimodal separation network includes an auditory subnetwork and a visual subnetwork, and the feature fusion module is further configured to:

[0065] The video features and the audio features are fused an iteratively through the auditory subnetwork, the visual subnetwork, the top module, the middle module, and the bottom module to obtain intermediate audio features and intermediate video features.

[0066] The intermediate audio features are processed iteratively through the auditory sub-network for a second preset number of audio processing steps to obtain the target audio features as the sound mask corresponding to the target object.

[0067] In one possible implementation, the feature fusion module is further configured to:

[0068] The video features and audio features are input into the auditory subnetwork and the visual subnetwork, respectively, to obtain first visual features and first auditory features at different time scales;

[0069] The top module fuses first visual features and first auditory features based on different time scales to obtain modulated second visual features and second auditory features;

[0070] The second visual feature and the second auditory feature are fused by the middle module to obtain the third auditory feature and the third visual feature after re-modulation.

[0071] The bottom module performs feature fusion on the third auditory feature and the third visual feature to obtain the fused fourth auditory feature and fourth visual feature, thus completing one iteration;

[0072] In response to the current iteration number being the first preset number, the iteration process ends and the fourth auditory feature and the fourth visual feature are determined to be the intermediate audio feature and the intermediate video feature, respectively.

[0073] If the current iteration number is not the first preset number, the fourth auditory feature is updated to an audio feature, the fourth visual feature is updated to a video feature, and the iteration is restarted.

[0074] In one possible implementation, the feature fusion module is further configured to:

[0075] The visual subnetwork downsamples the video features based on different time scales to obtain the first visual feature corresponding to each time scale.

[0076] The auditory subnetwork downsamples the audio features based on different time scales to obtain the first auditory feature corresponding to each time scale.

[0077] In one possible implementation, the top module includes an average normalization layer and a multilayer perceptron, and the feature fusion module is further configured to:

[0078] The first visual feature and the first auditory feature at each time scale are processed by the average normalization layer to obtain fused visual features and fused auditory features;

[0079] The fused auditory features are multiplied by the fused visual features processed by the sigma function, and then input into the multilayer perceptron for feature fusion to obtain the second auditory features.

[0080] The fused auditory features, processed by the sigma function, are multiplied by the fused visual features and then input into the multilayer perceptron for feature fusion to obtain the second visual features.

[0081] In one possible implementation, the feature fusion module is further configured to:

[0082] Through the self-attention module and dual-output module in the visual sub-network, a reference visual feature and a third visual feature for each time scale are generated based on the second visual feature and the first visual feature for each time scale.

[0083] The self-attention module in the audio sub-network generates reference auditory features for each time scale based on the second auditory feature and the first auditory feature for each time scale.

[0084] The central module performs feature fusion on the reference visual features and reference auditory features for each time scale to obtain the remodulated third auditory feature.

[0085] In one possible implementation, the feature fusion module is further configured to:

[0086] The second visual feature is upsampled based on different time scales to obtain the first intermediate visual feature corresponding to each time scale.

[0087] The intermediate visual features and the first intermediate visual features corresponding to each time scale are input into the self-attention module to obtain the second intermediate visual features corresponding to the time scale.

[0088] The second intermediate visual feature corresponding to each time scale is input into the dual output module, and the second intermediate visual feature corresponding to each time scale is output as a reference visual feature, while the third visual feature is output.

[0089] The process of determining the third visual feature includes:

[0090] Upsample the second intermediate visual feature corresponding to the smallest time scale;

[0091] The current upsampled visual features are repeatedly input into the attention module along with the second intermediate visual features corresponding to adjacent and larger time scales in an iterative manner, and the corresponding candidate visual features are then output.

[0092] In response to the fact that the time scale corresponding to the candidate visual feature is not the largest time scale, the candidate visual feature is upsampled;

[0093] In response to the candidate visual feature being the largest time scale, the candidate visual feature is determined to be the third visual feature.

[0094] In one possible implementation, the feature fusion module is further configured to:

[0095] The second auditory feature is upsampled based on different time scales to obtain the intermediate auditory feature corresponding to each time scale;

[0096] The intermediate auditory features and intermediate auditory features corresponding to each time scale are input into the self-attention module to obtain the reference auditory features corresponding to the time scale.

[0097] In one possible implementation, the feature fusion module is further configured to:

[0098] The central module performs a dot product of the reference visual features and reference auditory features for each time scale to obtain the feature product for each time scale.

[0099] Upsample the feature product corresponding to the smallest time scale;

[0100] The feature product after upsampling is input into the attention module multiple times in an iterative manner, along with the feature product corresponding to the adjacent and larger time scales. The updated feature product of the adjacent and larger time scales is then output.

[0101] If the time scale corresponding to the updated feature product is not the largest time scale, the feature product is upsampled.

[0102] In response to the updated feature product corresponding to the largest time scale, the feature product is determined to be the third auditory feature and output.

[0103] In one possible implementation, the feature fusion module is further configured to:

[0104] The third auditory feature and the third visual feature are input into the bottom module, and the third auditory feature and the third visual feature are processed by the convolutional layer of the bottom module to obtain convolutional auditory features and convolutional visual features.

[0105] The fourth auditory feature is obtained by multiplying the convolutional auditory feature and the convolutional visual feature processed by the sigma function;

[0106] The fourth visual feature is obtained by multiplying the convolutional auditory feature and the convolutional visual feature after processing with the sigma function.

[0107] In one possible implementation, the content extraction module is further configured to:

[0108] The video information is sampled for both image and audio using preset sampling parameters to obtain a sampled image sequence and mixed audio.

[0109] Identify and extract the lip region of the target object in each of the sampled images in the sampled image sequence to obtain an image frame sequence.

[0110] In one possible implementation, the audio determination module is further configured to:

[0111] The dot product of the sound mask and the audio feature is calculated and input into the audio decoder, which then transcodes the sound of the target object to obtain the target audio.

[0112] According to a third aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-described method when executing instructions stored in the memory.

[0113] According to a fourth aspect of this disclosure, a non-volatile computer-readable storage medium is provided that stores computer program instructions thereon, wherein the computer program instructions, when executed by a processor, implement the above-described method.

[0114] According to a fifth aspect of this disclosure, a computer program product is provided, including computer-readable code or a non-volatile computer-readable storage medium carrying the computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0115] In this embodiment, video information including the sound of a target object and the sound of at least one reference object is acquired. An image frame sequence consisting of lip image frames of the target object and mixed audio are extracted from the video information and encoded to obtain video features and audio features corresponding to the target object. The video features and audio features are input into a trained multimodal separation network. After multiple feature fusions, a sound mask corresponding to the target object is obtained. The multimodal separation network includes a top module, a middle module, and a bottom module for performing three feature fusions. The target audio for recording the target object's sound is determined based on the sound mask and audio features. This disclosure enhances intramodal contextual information by fusing visual and auditory information multiple times through three modules, improving audiovisual separation performance and obtaining accurate audio separation results.

[0116] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0117] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0118] Figure 1 A flowchart of an audiovisual speech separation method according to an embodiment of the present disclosure is shown;

[0119] Figure 2 A schematic diagram of a module structure according to an embodiment of the present disclosure is shown;

[0120] Figure 3 A schematic diagram of a multimodal separation network structure according to an embodiment of the present disclosure is shown;

[0121] Figure 4 A schematic diagram illustrating an audiovisual speech separation process according to an embodiment of the present disclosure is shown;

[0122] Figure 5 A schematic diagram of an audiovisual speech separation device according to an embodiment of the present disclosure is shown;

[0123] Figure 6 A schematic diagram of an electronic device according to an embodiment of the present disclosure is shown;

[0124] Figure 7 A schematic diagram of another electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0125] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0126] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0127] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0128] The audiovisual-speech separation method of this disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be any fixed or mobile terminal, such as a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, or wearable device. The server can be a single server or a server cluster consisting of multiple servers. Any electronic device can implement the audiovisual-speech separation method of this disclosure by having its processor call computer-readable instructions stored in its memory.

[0129] Figure 1 A flowchart illustrating an audiovisual speech separation method according to an embodiment of this disclosure is shown. Figure 1 As shown, an audiovisual speech separation method according to an embodiment of this disclosure may include the following steps S10-S50.

[0130] Step S10: Obtain video information including the sound of the target object and the sound of at least one reference object.

[0131] In one possible implementation, embodiments of this disclosure acquire video information requiring audio-visual-speech separation via an electronic device. This video information includes image frames and audio components. The audio component contains the voice of the target object to be separated and the voice of reference objects other than the target object. For example, in a video containing a dialogue between three individuals (A, B, and C), if it is necessary to extract the audio of A's speech, A is identified as the target object, and B and C are identified as reference objects. The audio-visual-speech separation process of this disclosure is used to extract the voice of the target object from the audio component of the video information. The electronic device can directly acquire video information by capturing the dialogue between at least two objects through an audio-visual acquisition device, or it can receive video information acquired and transmitted by other devices.

[0132] Step S20: Extract the image frame sequence and mixed audio from the video information.

[0133] In one possible implementation, after acquiring video information, the electronic device extracts the required image frame portions from the video information to obtain an image frame sequence, and simultaneously extracts the audio portion from it to obtain mixed audio. The image frame sequence can be a sequence composed of lip image frames of the target object for which audiovisual separation is required. The mixed audio is the original audio that combines the target object's sound and the sound of at least one reference object, and may also include background noise and other noise in addition to the target object's sound and the reference object's sound.

[0134] Optionally, the electronic device can acquire image frame sequences and mixed audio through sampling. Specifically, it can perform image and audio sampling on the video information using preset sampling parameters to obtain a sampled image sequence and mixed audio. The sampled images in the sampled image sequence are single, complete frames from the video information. After obtaining the sampled image sequence, the electronic device can identify and extract the lip region of the target object in each sampled image to obtain the image frame sequence. The preset sampling parameters can include image sampling parameters for image sampling and audio sampling parameters for audio sampling; for example, the image sampling parameter can be set to 25 FPS and the audio sampling parameter to 16 kHz.

[0135] Step S30: Encode the mixed audio and the image frame sequence respectively to obtain audio features and video features corresponding to the target object.

[0136] In one possible implementation, after extracting the mixed audio and image frame sequences from video information, the electronic device performs audio encoding and video encoding on the mixed audio and image frame sequences respectively to obtain corresponding audio features and video features. The audio and video encoding processes can be implemented using pre-set audio encoders and video encoders; that is, the mixed audio and image frame sequences are input into the audio encoder and video encoder respectively, and the corresponding audio features and video features are output respectively. The video encoder can be a pre-trained lip-reading model, and the audio encoder can be a one-dimensional convolutional layer; that is, the lip-reading model extracts video features from the image frame sequence, and the one-dimensional convolutional layer extracts audio features from the mixed audio. The video features are obtained by extracting the content information corresponding to the lips of the target object in the image frame, and the audio features are obtained by extracting the conversation content information from the mixed audio.

[0137] Step S40: Input the video features and the audio features into the trained multimodal separation network, and obtain the sound mask corresponding to the target object after multiple feature fusions.

[0138] In one possible implementation, after obtaining video and audio features through encoding, these features can be input into a trained multimodal separation network. Multiple feature fusions are then performed to obtain a sound mask representing the sound portion of the target object in the audio features. The multimodal separation network can include an auditory sub-network and a visual sub-network, respectively processing audio and video features. It can also include a top module, a middle module, and a bottom module. The top module performs initial fusion of video and audio features, the middle module performs further fusion at each time scale, and the bottom module performs fine-grained feature fusion. In other words, the multimodal separation sub-network can perform multiple feature fusions of audio and video features from different dimensions through different modules.

[0139] Optionally, in this embodiment of the present disclosure, after inputting video features and audio features into the trained multimodal separation network, the two types of features can be iteratively fused multiple times through a visual subnetwork and an auditory subnetwork. Then, the fused audio features are processed multiple times iteratively through the auditory subnetwork alone to output a sound mask. Specifically, the video and audio features are iteratively fused a first preset number of times through the auditory subnetwork, visual subnetwork, top module, middle module, and bottom module to obtain intermediate audio features and intermediate video features. Then, the intermediate audio features are iteratively processed a second preset number of times through the auditory subnetwork to obtain the target audio features as the sound mask corresponding to the target object. The audio processing can be used to filter and enhance the audio features multiple times to further improve the clarity of the separated signal.

[0140] Furthermore, the feature fusion process for video and audio features can include inputting video and audio features into the auditory and visual subnetworks respectively to obtain first visual and first auditory features at different time scales. Then, the top module fuses the first visual and first auditory features based on the different time scales to obtain modulated second visual and second auditory features. The middle module fuses the second visual and second auditory features to obtain remodulated third auditory and third visual features. The bottom module fuses the third auditory and third visual features to obtain fused fourth auditory and fourth visual features, completing one iteration. If the current iteration number is a first preset number, the iteration process ends, and the fourth auditory and fourth visual features are determined to be intermediate audio and video features, respectively. If the current iteration number is not the first preset number, the fourth auditory feature is updated to an audio feature, and the fourth visual feature is updated to a video feature, and the iteration restarts. In other words, in the feature fusion process of video and audio features, each iteration requires three feature fusion processes through the top, middle, and bottom modules. Before the first preset number of iterations is reached, the fourth auditory feature obtained after fusing the bottom module features is updated to an audio feature, and the fourth visual feature is updated to a video feature, thus entering the next iteration process. When the first preset number of iterations is reached, the fourth auditory feature and the fourth visual feature obtained after fusing the bottom module features are used as intermediate audio features and intermediate video features, respectively, to further process the intermediate audio features to obtain the sound mask corresponding to the target object.

[0141] In one possible implementation, the auditory and visual subnetworks can obtain the first auditory and first visual features corresponding to each time scale through downsampling. Specifically, the visual subnetwork downsamples video features based on different time scales to obtain the first visual feature corresponding to each time scale. Similarly, the auditory subnetwork downsamples audio features based on different time scales to obtain the first auditory feature corresponding to each time scale. The visual and auditory features corresponding to different time scales can be determined by downsampling a preset time scale with different numbers of iterations. Each downsampled time scale then has corresponding first visual and first auditory features. For example, when the audio and video features are E... S and E VIn this case, the first auditory feature S0 and the first auditory feature V0 are obtained by downsampling the audio features and video features once based on a preset visual scale through the auditory subnetwork and the visual subnetwork, respectively. The first auditory feature S1 and the first auditory feature V1 are obtained by downsampling the audio features and video features twice based on a preset time scale, respectively. The first auditory feature S2 and the first auditory feature V2 are obtained by downsampling the audio features and video features three times based on a preset time scale.

[0142] Optionally, after obtaining the first auditory and first visual features corresponding to each time scale during the current iteration through the auditory and visual sub-networks, the top module first performs feature fusion based on the first auditory and first visual features at each time scale to obtain modulated second visual and second auditory features. The top module may include an average normalization layer and a multilayer perceptron. The average normalization layer is used to fuse the first visual and first auditory features at different scales, and the multilayer perceptron is used to fuse the visual and auditory features. That is, the average normalization layer can process the first visual and first auditory features at each time scale separately to obtain fused visual and fused auditory features. Then, the fused auditory features are multiplied by the fused visual features processed by the sigma function and input into the multilayer perceptron for feature fusion to obtain the second auditory feature. Finally, the fused auditory features processed by the sigma function are multiplied by the fused visual features and input into the multilayer perceptron for feature fusion to obtain the second visual feature.

[0143] Figure 2 A schematic diagram of a module structure according to an embodiment of the present disclosure is shown. Figure 2 As shown, the top module includes a multilayer perceptron and an average normalization layer. The average normalization layer can process the first auditory feature S corresponding to each time scale. i Obtaining fused auditory features And processing the first visual feature V corresponding to each time scale. i Obtain fused visual features The auditory and temporal features are further fused using a sigma function σ. The fused auditory features are then multiplied by the fused visual features processed with the sigma function, and input into a multilayer perceptron for feature fusion to obtain the second auditory feature. The fused auditory features processed with the sigma function are then multiplied by the fused visual features, and input into a multilayer perceptron for feature fusion to obtain the second visual feature. The top module performs the first audiovisual feature fusion based on the above method. This fusion process integrates visual and auditory features from multiple time scales, captures the relevance of the context, and considers the important parts of each time scale in both auditory and visual perception.

[0144] Furthermore, after fusing the first auditory and first visual features through the top module to obtain the second visual and second auditory features, the middle module further fuses these features to obtain the re-modulated third auditory and third visual features. This feature fusion process can begin by using the self-attention module and dual-output module in the visual sub-network to generate reference visual and third visual features for each time scale based on the second visual features and the first visual features at each time scale. Then, the self-attention module in the audio sub-network generates reference auditory features for each time scale based on the second auditory features and the first auditory features at each time scale. Finally, the middle module fuses the reference visual and reference auditory features at each time scale to obtain the re-modulated third auditory feature.

[0145] Optionally, the process of determining the reference visual feature and the third visual feature may include upsampling the second visual feature based on different time scales to obtain a first intermediate visual feature corresponding to each time scale. The intermediate visual feature and the first intermediate visual feature corresponding to each time scale are input into a self-attention module to obtain a second intermediate visual feature corresponding to each time scale. The second intermediate visual feature corresponding to each time scale is input into a dual-output module, which outputs the second intermediate visual feature corresponding to each time scale as the reference visual feature, and simultaneously outputs the third visual feature.

[0146] like Figure 2 As shown in the dual-output module structure, the process of determining the third visual feature by the dual-output module may include: upsampling the second intermediate visual feature corresponding to the smallest time scale, iteratively inputting the currently upsampled visual feature together with the second intermediate visual feature corresponding to the adjacent and larger time scale into the attention module, and outputting the corresponding candidate visual feature. If the time scale corresponding to the candidate visual feature is not the largest time scale, the candidate visual feature is upsampled. If the time scale corresponding to the candidate visual feature is the largest time scale, the candidate visual feature is determined as the third visual feature. Optionally, the time scale size is determined based on the number of downsampling operations performed on the audio and video features; the more downsampling operations, the smaller the time scale, and the fewer downsampling operations, the larger the time scale.

[0147] In one possible implementation, the reference auditory feature for each time scale can be determined based on the second auditory feature and the first auditory feature for each time scale. Specifically, the second auditory feature can be upsampled at different time scales to obtain the intermediate auditory feature corresponding to each time scale. The intermediate auditory feature and the intermediate auditory feature corresponding to each time scale are then input into the self-attention module to obtain the reference auditory feature for that time scale. The process of upsampling the second auditory feature at different time scales can be achieved by directly upsampling the second auditory feature at each time scale to obtain the corresponding intermediate auditory feature. Alternatively, the second auditory feature can be upsampled multiple times at the same time scale, obtaining the intermediate auditory feature for each time scale.

[0148] like Figure 2 As shown in the middle module structure, after obtaining the reference visual and auditory features for each time scale, the middle module performs a dot product on the reference visual and auditory features for each time scale to obtain the feature product for each time scale. The feature product corresponding to the smallest time scale is then upsampled. Iteratively, the currently upsampled feature product is input together with the feature products corresponding to adjacent and larger time scales into the attention module, and the updated feature product for adjacent and larger time scales is output. If the time scale corresponding to the updated feature product is not the largest time scale, the feature product is upsampled. If the time scale corresponding to the updated feature product is the largest time scale, the feature product is determined as the third auditory feature and output. The middle module performs a second audiovisual feature fusion train in the above manner. This process captures relevant information between visual and auditory features, preserves more intramodal details, and introduces intermodal correlations.

[0149] Furthermore, after a second audiovisual feature fusion via the middle module, a third auditory feature and a third visual feature are obtained. Then, based on the third auditory and third visual features, a third audiovisual feature fusion is performed via the bottom module, resulting in a fused fourth auditory feature and a fourth visual feature. For example... Figure 2As shown in the bottom module structure, the third audiovisual feature fusion process involves inputting the third auditory and third visual features into the bottom module. The convolutional layers of the bottom module then process these features separately, resulting in convolutional auditory and visual features. A dot product is then performed between the convolutional auditory features and the convolutional visual features processed by the sigma function to obtain the fourth auditory feature. Similarly, a dot product is performed between the convolutional auditory features processed by the sigma function and the convolutional visual features to obtain the fourth visual feature. This third audiovisual fusion process in the bottom module obtains the finest-grained auditory and visual features through multi-layer upsampling and balances the information flow between visual and auditory features through attention, avoiding mutual interference between different modalities at deeper levels.

[0150] Figure 3 A schematic diagram of a multimodal separation network structure according to an embodiment of the present disclosure is shown. Figure 3 As shown, the electronic device acquires audio feature E S and video features E V Then, the data is input into a multimodal separation network. The auditory and visual sub-networks within the multimodal separation network acquire first visual and first auditory features at different time scales from the audio and video features, respectively. Further, the top, middle, and bottom modules perform first, second, and third audiovisual feature fusions, respectively, yielding second and second auditory, third and third auditory, and fourth auditory and fourth visual features after each fusion. Before reaching a first preset number of iterations, the fourth auditory feature obtained from the bottom module's feature fusion is updated to an audio feature, and the fourth visual feature is updated to a video feature, before proceeding to the next iteration. Upon reaching the first preset number of iterations, the fourth auditory and fourth visual features obtained from the bottom module's feature fusion are used as intermediate audio and video features, respectively, to further process the intermediate audio features and obtain the sound mask corresponding to the target object.

[0151] In one possible implementation, after completing the third audiovisual feature fusion through the bottom module, an iteration process is completed, and the current iteration number is determined. Before the iteration number reaches a first preset number, the fourth auditory feature obtained after feature fusion in the bottom module is updated to an audio feature, and the fourth visual feature is updated to a video feature, thus entering the next iteration process. When the iteration number reaches the first preset number, the fourth auditory feature and the fourth visual feature obtained after feature fusion in the bottom module are used as intermediate audio features and intermediate video features, respectively, to further process the intermediate audio features to obtain the sound mask corresponding to the target object.

[0152] Step S50: Determine the target audio for recording the sound of the target object based on the sound mask and the audio features.

[0153] In one possible implementation, after the electronic device performs audio-visual separation through a multimodal separation network to obtain a sound mask representing the sound portion of the target object in the mixed audio, it determines the target audio for recording the target object's sound based on the sound mask and audio features. The method for determining the target audio for recording the target object's sound can be to calculate the dot product of the sound mask and audio features and input it into an audio decoder, which then transcodes the audio to obtain the target audio for recording the target object's sound. Optionally, the audio decoder can be a one-dimensional transposed convolutional layer used to convert the speech features of the target object back to the time domain, generating the target audio recording the content of the target object's speech.

[0154] Figure 4 A schematic diagram illustrating an audiovisual speech separation process according to an embodiment of the present disclosure is shown. Figure 4 As shown, in this embodiment of the present disclosure, when it is necessary to separate the target object's speaking audio from the mixed speech portion of video information, the image frame sequence and mixed audio are first extracted from the video information. These are then encoded by a video encoder and an audio encoder respectively to obtain video features and audio features. These are input into a multimodal separation network for audiovisual speech separation. Features are learned from the video features to separate the target object's features from the audio features, resulting in a sound mask. The dot product of the nonlinear sound mask and the audio features is calculated and input into an audio decoder. The audio decoder then transcodes the audio to obtain the target audio recording the target object's voice.

[0155] Based on the aforementioned technical features, this embodiment of the disclosure utilizes cross-attention to perform audiovisual feature fusion at multiple levels through three modules—a top module, a middle module, and a bottom module—in a multimodal separation network. The top module extracts context from multiple visual and auditory scales and fuses multimodal information using cross-attention, enhancing intramodal contextual information to improve audiovisual separation performance. The middle module uses lateral connections and cross-attention to simultaneously aggregate visual and auditory features from the same layer, preserving fine-grained details of multimodal features while maintaining good computational efficiency. The bottom module fuses fine-grained visual and auditory features through cross-attention, providing global-level guidance for the reconstruction of fine-grained auditory features. This audiovisual separation method learns explicit auditory perception between vision and hearing through multi-stage fusion, producing clearer target audio.

[0156] Figure 5 A schematic diagram of an audiovisual speech separation device according to an embodiment of the present disclosure is shown. Figure 5 As shown, the audiovisual speech separation device in this embodiment of the present disclosure may include:

[0157] Information acquisition module 50 is used to acquire video information including the sound of the target object and the sound of at least one reference object;

[0158] Content extraction module 51 is used to extract image frame sequences and mixed audio from the video information, wherein the image frame sequence is a sequence composed of lip image frames of the target object;

[0159] Encoding module 52 is used to encode the mixed audio and the image frame sequence respectively to obtain audio features and video features corresponding to the target object;

[0160] The feature fusion module 53 is used to input the video features and the audio features into the trained multimodal separation network, and obtain the sound mask corresponding to the target object after multiple feature fusions. The multimodal separation network includes a top module, a middle module, and a bottom module. The top module is used to perform preliminary fusion of the video features and the audio features. The middle module is used to perform further fusion of the video features and the audio features at each time scale. The bottom module is used to perform fine-grained feature fusion of the video features and the auditory features.

[0161] The audio determination module 54 is used to determine the target audio for recording the sound of the target object based on the sound mask and the audio features.

[0162] In one possible implementation, the multimodal separation network includes an auditory subnetwork and a visual subnetwork, and the feature fusion module 53 is further configured to:

[0163] The video features and the audio features are fused an iteratively through the auditory subnetwork, the visual subnetwork, the top module, the middle module, and the bottom module to obtain intermediate audio features and intermediate video features.

[0164] The intermediate audio features are processed iteratively through the auditory sub-network for a second preset number of audio processing steps to obtain the target audio features as the sound mask corresponding to the target object.

[0165] In one possible implementation, the feature fusion module 53 is further configured to:

[0166] The video features and audio features are input into the auditory subnetwork and the visual subnetwork, respectively, to obtain first visual features and first auditory features at different time scales;

[0167] The top module fuses first visual features and first auditory features based on different time scales to obtain modulated second visual features and second auditory features;

[0168] The second visual feature and the second auditory feature are fused by the middle module to obtain the third auditory feature and the third visual feature after re-modulation.

[0169] The bottom module performs feature fusion on the third auditory feature and the third visual feature to obtain the fused fourth auditory feature and fourth visual feature, thus completing one iteration;

[0170] In response to the current iteration number being the first preset number, the iteration process ends and the fourth auditory feature and the fourth visual feature are determined to be the intermediate audio feature and the intermediate video feature, respectively.

[0171] If the current iteration number is not the first preset number, the fourth auditory feature is updated to an audio feature, the fourth visual feature is updated to a video feature, and the iteration is restarted.

[0172] In one possible implementation, the feature fusion module 53 is further configured to:

[0173] The visual subnetwork downsamples the video features based on different time scales to obtain the first visual feature corresponding to each time scale.

[0174] The auditory subnetwork downsamples the audio features based on different time scales to obtain the first auditory feature corresponding to each time scale.

[0175] In one possible implementation, the top module includes an average normalization layer and a multilayer perceptron, and the feature fusion module 53 is further configured to:

[0176] The first visual feature and the first auditory feature at each time scale are processed by the average normalization layer to obtain fused visual features and fused auditory features;

[0177] The fused auditory features are multiplied by the fused visual features processed by the sigma function, and then input into the multilayer perceptron for feature fusion to obtain the second auditory features.

[0178] The fused auditory features, processed by the sigma function, are multiplied by the fused visual features and then input into the multilayer perceptron for feature fusion to obtain the second visual features.

[0179] In one possible implementation, the feature fusion module 53 is further configured to:

[0180] Through the self-attention module and dual-output module in the visual sub-network, a reference visual feature and a third visual feature for each time scale are generated based on the second visual feature and the first visual feature for each time scale.

[0181] The self-attention module in the audio sub-network generates reference auditory features for each time scale based on the second auditory feature and the first auditory feature for each time scale.

[0182] The central module performs feature fusion on the reference visual features and reference auditory features for each time scale to obtain the remodulated third auditory feature.

[0183] In one possible implementation, the feature fusion module 53 is further configured to:

[0184] The second visual feature is upsampled based on different time scales to obtain the first intermediate visual feature corresponding to each time scale.

[0185] The intermediate visual features and the first intermediate visual features corresponding to each time scale are input into the self-attention module to obtain the second intermediate visual features corresponding to the time scale.

[0186] The second intermediate visual feature corresponding to each time scale is input into the dual output module, and the second intermediate visual feature corresponding to each time scale is output as a reference visual feature, while the third visual feature is output.

[0187] The process of determining the third visual feature includes:

[0188] Upsample the second intermediate visual feature corresponding to the smallest time scale;

[0189] The current upsampled visual features are repeatedly input into the attention module along with the second intermediate visual features corresponding to adjacent and larger time scales in an iterative manner, and the corresponding candidate visual features are then output.

[0190] In response to the fact that the time scale corresponding to the candidate visual feature is not the largest time scale, the candidate visual feature is upsampled;

[0191] In response to the candidate visual feature being the largest time scale, the candidate visual feature is determined to be the third visual feature.

[0192] In one possible implementation, the feature fusion module 53 is further configured to:

[0193] The second auditory feature is upsampled based on different time scales to obtain the intermediate auditory feature corresponding to each time scale;

[0194] The intermediate auditory features and intermediate auditory features corresponding to each time scale are input into the self-attention module to obtain the reference auditory features corresponding to the time scale.

[0195] In one possible implementation, the feature fusion module 53 is further configured to:

[0196] The central module performs a dot product of the reference visual features and reference auditory features for each time scale to obtain the feature product for each time scale.

[0197] Upsample the feature product corresponding to the smallest time scale;

[0198] The feature product after upsampling is input into the attention module multiple times in an iterative manner, along with the feature product corresponding to the adjacent and larger time scales. The updated feature product of the adjacent and larger time scales is then output.

[0199] If the time scale corresponding to the updated feature product is not the largest time scale, the feature product is upsampled.

[0200] In response to the updated feature product corresponding to the largest time scale, the feature product is determined to be the third auditory feature and output.

[0201] In one possible implementation, the feature fusion module 53 is further configured to:

[0202] The third auditory feature and the third visual feature are input into the bottom module, and the third auditory feature and the third visual feature are processed by the convolutional layer of the bottom module to obtain convolutional auditory features and convolutional visual features.

[0203] The fourth auditory feature is obtained by multiplying the convolutional auditory feature and the convolutional visual feature processed by the sigma function;

[0204] The fourth visual feature is obtained by multiplying the convolutional auditory feature and the convolutional visual feature after processing with the sigma function.

[0205] In one possible implementation, the content extraction module 50 is further configured to:

[0206] The video information is sampled for both image and audio using preset sampling parameters to obtain a sampled image sequence and mixed audio.

[0207] Identify and extract the lip region of the target object in each of the sampled images in the sampled image sequence to obtain an image frame sequence.

[0208] In one possible implementation, the audio determination module 51 is further configured to:

[0209] The dot product of the sound mask and the audio feature is calculated and input into the audio decoder, which then transcodes the sound of the target object to obtain the target audio.

[0210] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0211] This disclosure also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the above-described method. The computer-readable storage medium can be volatile or non-volatile.

[0212] This disclosure also proposes an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0213] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0214] Figure 6 A schematic diagram of an electronic device 800 according to an embodiment of the present disclosure is shown. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0215] Reference Figure 6 The electronic device 800 may include one or more of the following components: processing component 802, memory 804, power supply component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.

[0216] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0217] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of this data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0218] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.

[0219] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0220] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0221] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0222] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0223] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0224] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0225] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 804 including computer program instructions that can be executed by a processor 820 of an electronic device 800 to perform the above-described method.

[0226] Figure 7 A schematic diagram of another electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. (Refer to...) Figure 7 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0227] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output (I / O) interface 1958. Electronic device 1900 can operate on an operating system stored in memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0228] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.

[0229] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0230] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0231] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0232] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0233] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0234] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0235] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0236] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0237] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for separating audiovisual speech, characterized in that, The method includes: Acquire video information including the sound of the target object and the sound of at least one reference object; Extract the image frame sequence and mixed audio from the video information, wherein the image frame sequence is a sequence composed of image frames of the lip area of ​​the target object; The mixed audio and the image frame sequence are encoded separately to obtain audio features and video features corresponding to the target object; The video features and audio features are input into a multimodal separation network trained on the network. After multiple feature fusions, the sound mask corresponding to the target object is obtained. The multimodal separation network includes a top module, a middle module, and a bottom module. The top module is used to perform preliminary fusion of the video features and audio features. The middle module is used to perform further fusion of the video features and audio features at each time scale. The bottom module is used to perform fine-grained feature fusion of the video features and audio features. The target audio for recording the sound of the target object is determined based on the sound mask and the audio features.

2. The method according to claim 1, characterized in that, The multimodal separation network includes an auditory subnetwork and a visual subnetwork. The multimodal separation network, trained by inputting the video features and the audio features, obtains the sound mask corresponding to the target object through multiple feature fusions, including: The video features and the audio features are fused an iteratively through the auditory subnetwork, the visual subnetwork, the top module, the middle module, and the bottom module to obtain intermediate audio features and intermediate video features. The intermediate audio features are processed iteratively through the auditory sub-network for a second preset number of audio processing steps to obtain the target audio features as the sound mask corresponding to the target object.

3. The method according to claim 2, characterized in that, The feature fusion process, which iteratively integrates the video features and the audio features through the auditory subnetwork, visual subnetwork, top module, middle module, and bottom module for a first preset number of iterations to obtain intermediate audio features and intermediate video features, includes: The video features and audio features are input into the auditory subnetwork and the visual subnetwork, respectively, to obtain first visual features and first auditory features at different time scales; The top module fuses first visual features and first auditory features based on different time scales to obtain modulated second visual features and second auditory features; The second visual feature and the second auditory feature are fused by the middle module to obtain the third auditory feature and the third visual feature after re-modulation. The bottom module performs feature fusion on the third auditory feature and the third visual feature to obtain the fused fourth auditory feature and fourth visual feature, thus completing one iteration; In response to the current iteration number being the first preset number, the iteration process ends and the fourth auditory feature and the fourth visual feature are determined to be the intermediate audio feature and the intermediate video feature, respectively. If the current iteration number is not the first preset number, the fourth auditory feature is updated to an audio feature, the fourth visual feature is updated to a video feature, and the iteration is restarted.

4. The method according to claim 3, characterized in that, The step of inputting the video features and audio features into the auditory subnetwork and the visual subnetwork respectively to obtain first visual features and first auditory features at different time scales includes: The visual subnetwork downsamples the video features based on different time scales to obtain the first visual feature corresponding to each time scale. The auditory subnetwork downsamples the audio features based on different time scales to obtain the first auditory feature corresponding to each time scale.

5. The method according to claim 3 or 4, characterized in that, The top module includes an average normalization layer and a multilayer perceptron. The process of fusing first visual features and first auditory features based on different time scales through the top module to obtain modulated second visual features and second auditory features includes: The first visual feature and the first auditory feature at each time scale are processed by the average normalization layer to obtain fused visual features and fused auditory features; The fused auditory features are multiplied by the fused visual features processed by the sigma function, and then input into the multilayer perceptron for feature fusion to obtain the second auditory features. The fused auditory features, processed by the sigma function, are multiplied by the fused visual features and then input into the multilayer perceptron for feature fusion to obtain the second visual features.

6. The method according to claim 3, characterized in that, The step of fusing the second visual feature and the second auditory feature through the middle module to obtain the remodulated third auditory feature and third visual feature includes: Through the self-attention module and dual-output module in the visual sub-network, a reference visual feature and a third visual feature for each time scale are generated based on the second visual feature and the first visual feature for each time scale. The self-attention module in the auditory sub-network generates reference auditory features for each time scale based on the second auditory feature and the first auditory feature for each time scale. The central module performs feature fusion on the reference visual features and reference auditory features for each time scale to obtain the remodulated third auditory feature.

7. The method according to claim 6, characterized in that, The step of generating reference visual features and third visual features for each time scale based on the second visual features and the first visual features for each time scale through the self-attention module and dual-output module in the visual sub-network includes: The second visual feature is upsampled based on different time scales to obtain the first intermediate visual feature corresponding to each time scale. The intermediate visual features and the first intermediate visual features corresponding to each time scale are input into the self-attention module to obtain the second intermediate visual features corresponding to the time scale. The second intermediate visual feature corresponding to each time scale is input into the dual output module, and the second intermediate visual feature corresponding to each time scale is output as a reference visual feature, while the third visual feature is output. The process of determining the third visual feature includes: Upsample the second intermediate visual feature corresponding to the smallest time scale; The current upsampled visual features are repeatedly input into the attention module along with the second intermediate visual features corresponding to adjacent and larger time scales in an iterative manner, and the corresponding candidate visual features are then output. In response to the fact that the time scale corresponding to the candidate visual feature is not the largest time scale, the candidate visual feature is upsampled; In response to the candidate visual feature being the largest time scale, the candidate visual feature is determined to be the third visual feature.

8. The method according to claim 6, characterized in that, The step of generating reference auditory features for each time scale using the self-attention module in the auditory sub-network, based on the second auditory feature and the first auditory feature for each time scale, includes: The second auditory feature is upsampled based on different time scales to obtain the intermediate auditory feature corresponding to each time scale; The intermediate auditory features and intermediate auditory features corresponding to each time scale are input into the self-attention module to obtain the reference auditory features corresponding to the time scale.

9. The method according to claim 6, characterized in that, The process involves fusing the reference visual and auditory features at each time scale using the central module to obtain a remodulated third auditory feature, including: The central module performs a dot product of the reference visual features and reference auditory features for each time scale to obtain the feature product for each time scale. Upsample the feature product corresponding to the smallest time scale; The feature product after upsampling is input into the attention module multiple times in an iterative manner, along with the feature product corresponding to the adjacent and larger time scales. The updated feature product of the adjacent and larger time scales is then output. If the time scale corresponding to the updated feature product is not the largest time scale, the feature product is upsampled. In response to the updated feature product corresponding to the largest time scale, the feature product is determined to be the third auditory feature and output.

10. The method according to claim 3, characterized in that, The step of fusing the third auditory feature and the third visual feature through the bottom module to obtain the fused fourth auditory feature and fourth visual feature includes: The third auditory feature and the third visual feature are input into the bottom module, and the third auditory feature and the third visual feature are processed by the convolutional layer of the bottom module to obtain convolutional auditory features and convolutional visual features. The fourth auditory feature is obtained by multiplying the convolutional auditory feature and the convolutional visual feature processed by the sigma function; The fourth visual feature is obtained by multiplying the convolutional auditory feature and the convolutional visual feature after processing with the sigma function.

11. The method according to claim 1, characterized in that, The step of extracting the image frame sequence and mixed audio from the video information includes: The video information is sampled for both image and audio using preset sampling parameters to obtain a sampled image sequence and mixed audio. Identify and extract the lip region of the target object in each of the sampled images in the sampled image sequence to obtain an image frame sequence.

12. The method according to claim 1, characterized in that, Determining the target audio for recording the sound of the target object based on the sound mask and the audio features includes: The dot product of the sound mask and the audio feature is calculated and input into the audio decoder, which then transcodes the sound of the target object to obtain the target audio.

13. An audiovisual speech separation device, characterized in that, The device includes: The information acquisition module is used to acquire video information including the sound of the target object and the sound of at least one reference object; The content extraction module is used to extract the image frame sequence and mixed audio from the video information, wherein the image frame sequence is a sequence composed of image frames of the lip area of ​​the target object; The encoding module is used to encode the mixed audio and the image frame sequence respectively to obtain audio features and video features corresponding to the target object; The feature fusion module is used to input the video features and the audio features into the trained multimodal separation network, and obtain the sound mask corresponding to the target object after multiple feature fusions. The multimodal separation network includes a top module, a middle module, and a bottom module. The top module is used to perform preliminary fusion of the video features and the audio features. The middle module is used to perform further fusion of the video features and the audio features at each time scale. The bottom module is used to perform fine-grained feature fusion of the video features and the audio features. An audio determination module is used to determine the target audio for recording the sound of the target object based on the sound mask and the audio features.

14. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the method of any one of claims 1 to 12 when executing instructions stored in the memory.

15. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Multi-modal speech separation method, training method and related device

    CN115620723A

  • Audio-visual voice separation method and device, electronic equipment and storage medium

    CN116129929A