Method executed by electronic equipment, electronic equipment and storage medium

By generating a target sound mask and analyzing the audio signal frame by frame, the problem of inaccurate target sound extraction in video restoration is solved, achieving precise audio signal processing and improving the user experience.

CN122073115APending Publication Date: 2026-05-22BEIJING SAMSUNG TELECOM R&D CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING SAMSUNG TELECOM R&D CENT
Filing Date
2024-11-22
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

In video restoration operations, existing technologies struggle to accurately remove or extract target-related sounds from video-related audio, especially when the target is moving, obscured, or overlaps with other sound sources, leading to inaccurate sound processing.

Method used

By generating a target sound mask based on target information, audio signals, and direction information in the video, analyzing the audio signal frame by frame, obtaining the target's global sound features and motion trend features, updating the target's bimodal mask, and accurately extracting and removing the target sound.

Benefits of technology

It enables precise extraction and removal of target audio during video restoration, improving user experience and ensuring the accuracy and consistency of audio signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122073115A_ABST
    Figure CN122073115A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method executed by electronic equipment, the electronic equipment and a storage medium, and relates to the field of artificial intelligence. The method comprises the following steps: obtaining a target sound mask of a target at each moment based on image-related information of the target in a first video, a first audio signal corresponding to the first video and direction information of the first audio signal; and based on the target sound mask of the target at each moment and the first audio signal, a second audio signal is obtained, and the second audio signal does not include sound related to the target. Optionally, the method performed by the electronic device may be performed using an artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of signal processing technology, and more specifically, to a method for processing audio signals performed by an electronic device, an electronic device, and a storage medium. Background Technology

[0002] Currently, video restoration operations can fill in damaged target areas in a video (e.g., using potentially existing content, such as an image consistent with the background), and can also remove selected areas or target objects, then fill them with content that is consistent in both time and space. However, video restoration operations typically do not process the audio corresponding to the video; sounds related to the removed target are not eliminated. This is because it is impossible to directly determine the sounds related to the target, and the target may move or be occluded, which further increases the difficulty of eliminating or extracting the sounds related to the target.

[0003] How to accurately remove or extract target-related sounds from video-related audio during video restoration operations to meet user needs is a technical problem that those skilled in the art have been working hard to study. Summary of the Invention

[0004] In order to at least solve the above-mentioned problems existing in the prior art, the present invention provides a method performed by an electronic device, an electronic device, and a storage medium.

[0005] According to a first aspect of the embodiments of this application, a method executed by an electronic device is provided, comprising: obtaining a target sound mask of the target at various times based on image-related information of a target in a first video, a first audio signal corresponding to the first video, and direction information of the first audio signal; and obtaining a second audio signal based on the target sound mask of the target at various times and the first audio signal, wherein the second audio signal does not include sound related to the target.

[0006] Optionally, the image-related information includes the target visual mask, depth information, and optical flow information of the target.

[0007] Optionally, based on image-related information of the target in the first video, a first audio signal corresponding to the first video, and direction information of the first audio signal, a target sound mask of the target at each moment is obtained, including: obtaining a first mask for each audio frame of the first audio signal based on the image-related information and the direction information; and obtaining the target sound mask of the target in each audio frame of the first audio signal based on the first mask for each audio frame and the encoding features of the first audio signal.

[0008] Optionally, obtaining a first mask for each audio frame of the first audio signal based on the image-related information and the direction information includes: obtaining spatial distribution features of the sound signal for each audio frame by normalizing and encoding the direction information; obtaining spatial distribution features of the target for each audio frame by encoding the target visual mask and depth information of the target in the image-related information; obtaining a first feature for each audio frame based on the spatial distribution features of the sound signal and the spatial distribution features of the target; and obtaining the first mask for each audio frame based on the first feature and the optical flow information of the target in the image-related information.

[0009] Optionally, based on the spatial distribution features of the sound signal and the spatial distribution features of the target, a first feature is obtained for each audio frame, including: obtaining the first feature for each audio frame by performing feature processing on the spatial distribution features of the sound signal and the spatial distribution features of the target, wherein the first feature represents the position of the target in space and the direction of the sound contained in the target visual mask.

[0010] Optionally, obtaining the first mask for each audio frame based on the first feature and the optical flow information of the target in the image-related information includes: determining the spatial motion trend of the target for each audio frame based on the first feature and the optical flow information; determining the sound source motion trend within the visual mask of the target for each audio frame based on the first feature; and obtaining the first mask for each audio frame based on the determination results of the spatial motion trend and the sound source motion trend.

[0011] Optionally, determining the spatial motion trend of the target for each audio frame based on the first feature and the optical flow information includes: determining the spatial motion trend of the target in each sub-part of space for each audio frame based on visual information, according to the features of the first feature and the optical flow information.

[0012] Optionally, determining the sound source motion trend within the target visual mask for each audio frame based on the first feature includes: determining the sound source motion trend of the target in each sub-part of space for each audio frame based on sound information, according to the first feature.

[0013] Optionally, obtaining the target sound mask of the target in each audio frame of the first audio signal based on the first mask for each audio frame and the encoding features of the first audio signal includes: obtaining global sound features and global motion trend features of the target based on the first mask for each audio frame and the encoding features of the first audio signal, wherein the global sound features represent the features of all sounds related to the target, and the global motion trend features represent the motion trajectory information of the target in the first video; and determining the target sound mask of the target in each audio frame based on the global sound features and the global motion trend features.

[0014] Optionally, determining the target sound mask of the target in each audio frame based on the global sound features and the global motion trend features includes: updating the first mask for each audio frame based on the global motion trend features; and determining the target sound mask of the target in each audio frame based on the encoding features of the first audio signal, the global sound features, and the updated first mask for each audio frame.

[0015] Optionally, determining the target sound mask of the target in each audio frame based on the encoding features of the first audio signal, the global sound features, and the updated first mask for each audio frame includes: eliminating non-target sound features from the encoding features of each audio frame of the first audio signal based on the global sound features and the updated first mask for each audio frame; and determining the target sound mask of the target in each audio frame based on the encoding features of each audio frame after eliminating non-target sound features.

[0016] Optionally, obtaining a second audio signal based on the target sound mask of the target at each time point and the first audio signal includes: removing sound features related to the target from the encoding features of the first audio signal based on the target sound mask of the target in each audio frame to obtain non-target sound signal features of the first audio signal; and obtaining the second audio signal based on the non-target sound signal features.

[0017] Optionally, obtaining a second audio signal based on the non-target sound signal features includes: repairing the non-target sound signal features based on the updated first mask for each audio frame to obtain updated non-target sound signal features; and obtaining the second audio signal by decoding the updated non-target sound signal features.

[0018] Optionally, the updated first mask is obtained by: obtaining global motion trend features based on the first mask for each audio frame and the encoding features of the first audio signal; and updating the first mask for each audio frame based on the global motion trend features.

[0019] Optionally, updating the first mask for each audio frame based on the global motion trend features includes: adjusting at least one of the spatial motion trend and the sound source motion trend in the first mask of the current audio frame by comparing the motion trend of the target in the current audio frame with the global motion trend features and calculating the trend consistency.

[0020] Optionally, repairing the non-target sound signal features based on the updated first mask for each audio frame to obtain the updated non-target sound signal features includes: obtaining the room impact response of the non-target when it is not occluded by the target based on the updated first mask for each audio frame and the non-target sound signal features; and repairing the non-target sound signal features based on the room impact response to obtain the updated non-target sound signal features.

[0021] Optionally, obtaining the room impact response of the non-target when it is not occluded by the target, based on the updated first mask for each audio frame and the non-target sound signal features, includes: selecting signal features of a plurality of audio frames before and / or after the non-target is occluded by the target from the non-target sound signal features based on the updated first mask for each audio frame; and obtaining the room impact response of the non-target when it is not occluded by the target based on the signal features corresponding to the plurality of audio frames.

[0022] According to a second aspect of the present application, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform a method performed by the electronic device as described above.

[0023] According to a third aspect of the present application, a computer-readable storage medium for storing instructions is provided, wherein when the instructions are executed by at least one processor, they cause the at least one processor to perform a method performed by an electronic device as described above.

[0024] The beneficial effects of the technical solutions provided in this application will be explained in the following text in conjunction with specific optional embodiments, or can be learned from the description of the embodiments, or can be learned through the implementation of the embodiments. Attached Figure Description

[0025] To more clearly and easily illustrate and understand the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0026] Figure 1 A diagram illustrating a scheme for synchronously processing audio during video restoration is shown.

[0027] Figure 2 This is a diagram showing a beam formed in the direction of the target mask based on the target mask and the audio signal.

[0028] Figure 3A This is an illustration showing an example of a target being obscured or moved off-screen.

[0029] Figure 3B This is a diagram illustrating an example where the target overlaps with other sound sources.

[0030] Figure 4A This is a flowchart illustrating a method performed by an electronic device according to an exemplary embodiment of this application.

[0031] Figure 4B This is a schematic diagram illustrating a network structure corresponding to a method performed by an electronic device according to an exemplary embodiment of this application.

[0032] Figure 4C It is a diagram showing the changes in the spatial distribution characteristics of the sound signal and the target signal at two different times.

[0033] Figure 5A This is a flowchart illustrating an exemplary embodiment of the present application of obtaining a target bimodal mask based on target-related image information and orientation information.

[0034] Figure 5B This is a schematic diagram illustrating the process by which an auditory-visual feature analysis module generates a target bimodal mask according to an exemplary embodiment of this application.

[0035] Figure 6A This is a diagram showing an example of DOA information.

[0036] Figure 6B It shows the spatial distribution characteristics of the sound signal X spa An example illustration.

[0037] Figure 7A A diagram illustrating an example of a target spatial distribution feature according to an exemplary embodiment of this application is shown.

[0038] Figure 7B A spatial diagram of visual information is shown.

[0039] Figure 7C A diagram illustrating an example of discontinuous sound from a target is shown.

[0040] Figure 7D The diagram shows an example where the target's location overlaps with other objects.

[0041] Figure 8 This is a diagram illustrating an example of the spatial distribution characteristics of a sound signal and the spatial distribution characteristics of a target, as well as examples of the target's bimodal characteristics, according to an exemplary embodiment of this application.

[0042] Figure 9 This is a diagram illustrating an example of obtaining spatial motion trends based on visual information and sound source motion trends based on sound information, according to an exemplary embodiment of this application.

[0043] Figure 10A This is a flowchart illustrating the process of obtaining a target sound mask for a target in each audio frame of a first audio signal according to an exemplary embodiment of this application.

[0044] Figure 10B This is a schematic diagram illustrating the process by which a dual-modal, dual-stage sound extraction module obtains a target sound mask for a target in each audio frame of a first audio signal, according to an exemplary embodiment of this application.

[0045] Figure 11 This is a block diagram illustrating an encoder module according to an exemplary embodiment of this application.

[0046] Figure 12 This is a network flowchart illustrating an encoder module according to an exemplary embodiment of this application.

[0047] Figure 13 This is a schematic diagram illustrating the process by which the global information analysis module, according to an exemplary embodiment of this application, processes the encoded features of the target bimodal mask and the first audio signal.

[0048] Figure 14 This is a diagram illustrating an example of obtaining an updated target bimodal mask according to an exemplary embodiment of this application.

[0049] Figure 15 This is a diagram illustrating an example of obtaining a target sound mask in the current audio frame according to an exemplary embodiment of this application.

[0050] Figure 16 The diagram shows a network structure corresponding to a method performed by an electronic device according to another exemplary embodiment of this application.

[0051] Figure 17 This is a schematic diagram illustrating the structure of a repair module according to an exemplary embodiment of this application.

[0052] Figure 18 This is a block diagram illustrating a decoder module according to an exemplary embodiment of this application.

[0053] Figure 19 A schematic diagram illustrates the application of the method performed by an electronic device according to this application in a scenario where the target overlaps with other sound sources during video recording.

[0054] Figure 20 This diagram illustrates a method performed by an electronic device according to this application in a scenario where the target is not in the screen during video recording.

[0055] Figure 21 This is a schematic diagram illustrating the structure of an electronic device to which an exemplary embodiment of this application applies. Detailed Implementation

[0056] The following description, with reference to the accompanying drawings, is provided to aid in a thorough understanding of the various embodiments of this disclosure as defined by the claims and their equivalents. This description includes various specific details to aid understanding but should be considered exemplary only. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of this disclosure. Furthermore, for clarity and brevity, descriptions of well-known functions and structures may be omitted.

[0057] The terms and wording used in the following description and claims are not limited to their dictionary meanings, but are merely used by the inventors to enable a clear and consistent understanding of this disclosure. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of this disclosure is for illustrative purposes only and not for limiting the purpose of this disclosure as defined in the appended claims and their equivalents.

[0058] It should be understood that the singular forms of “a,” “an,” and “the” can also include plural references unless the context clearly indicates otherwise. Thus, for example, the reference to “component surface” includes referring to one or more such surfaces. When we say that an element is “connected” or “coupled” to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element are connected through an intermediate element. Furthermore, the use of “connected” or “coupled” herein can include wireless connections or wireless couplings.

[0059] The terms “comprising” or “may include” refer to the presence of a corresponding disclosed function, operation, or component that may be used in the various embodiments of this disclosure, rather than limiting the presence of one or more additional functions, operations, or features. Furthermore, the terms “comprising” or “having” may be interpreted as indicating certain characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof, but should not be construed as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof.

[0060] The term "or" as used in the various embodiments of this disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not explicitly defined, the multiple items may refer to one, more, or all of the multiple items. For example, the description "parameter A includes A1, A2, A3" can be implemented as parameter A includes A1 or A2 or A3, or it can be implemented as parameter A includes at least two of the three items A1, A2, and A3.

[0061] Unless otherwise defined, all terms used in this disclosure (including technical or scientific terms) have the same meaning as understood by one of those skilled in the art to which this disclosure pertains. Common terms as defined in dictionaries are to be interpreted as having a meaning consistent with the context in the relevant technical field and should not be interpreted ideally or overly formally unless expressly defined in this disclosure.

[0062] At least some of the functions of the device or electronic device provided in this disclosure embodiment can be implemented by an AI model, such as implementing at least one module of a plurality of modules of the device or electronic device by an AI model. AI-related functions can be executed by non-volatile memory, volatile memory, and a processor.

[0063] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as central processing unit (CPU), application processor (AP), etc., or pure graphics processing unit, such as graphics processing unit (GPU), vision processing unit (VPU), and / or AI-specific processors, such as neural processing unit (NPU).

[0064] The one or more processors control the processing of input data based on predefined operating rules or artificial intelligence (AI) models stored in non-volatile and volatile memory. These predefined operating rules or AI models are provided through training or learning.

[0065] Here, "providing through learning" refers to obtaining predefined operating rules or an AI model with desired characteristics by applying a learning algorithm to multiple learning datasets. This learning can be performed within the device or electronic device itself, in which the AI ​​is executed according to the embodiment, and / or can be implemented via a separate server / system.

[0066] AI models can contain multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network computations by calculating the input data of that layer (such as the computation results of the previous layer and / or the input data of the AI ​​model) and the multiple weight values ​​of the current layer. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q-networks.

[0067] A learning algorithm is a method of training a predetermined target device (e.g., a robot) using multiple learning data sets to enable, allow, or control the target device to make determinations or predictions. Examples of such learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0068] The methods provided in this disclosure may relate to one or more fields in the technical fields of speech, language, image, video, or data intelligence.

[0069] Optionally, in the context of speech or language, in the method performed by an electronic device according to this disclosure, a speech signal as an analog signal may be received via a speech input device (e.g., a microphone), and the speech portion may be converted into computer-readable text using an Automatic Speech Recognition (ASR) model. The user's utterance intent can be obtained by interpreting the converted text using a Natural Language Understanding (NLU) model. The ASR model or NLU model may be an artificial intelligence model. The artificial intelligence model may be processed by a dedicated artificial intelligence processor designed in a hardware architecture specified for processing the artificial intelligence model. Language understanding is a technique for recognizing and applying / processing human language / text, including, for example, natural language processing, machine translation, dialogue systems, question answering, or speech recognition / synthesis.

[0070] Optionally, when dealing with the field of images or videos, in the method performed by an electronic device according to this disclosure, output data can be obtained by using image data as input data for an artificial intelligence model. The methods of this disclosure can relate to the field of visual understanding in artificial intelligence technology, which is a technology for recognizing and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.

[0071] Optionally, in the field of data intelligence processing, in the method performed by an electronic device according to this disclosure, during the reasoning or prediction phase, an artificial intelligence model can be used to perform prediction by using real-time input data. The processor of the electronic device can perform preprocessing operations on the data to transform it into a form suitable for use as input to the artificial intelligence model. Reasoning and prediction are techniques for making logical inferences and predictions by determining information, including, for example, knowledge-based reasoning, optimization prediction, preference-based planning, or recommendation.

[0072] In this application, the artificial intelligence model can be obtained through training. Here, "obtained through training" means obtaining a predefined operational rule or artificial intelligence model configured to perform desired features (or objectives) by training a basic artificial intelligence model with multiple training data using a training algorithm. The artificial intelligence model may include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values, and neural network computation is performed by calculating the results of the previous layer and the multiple weight values.

[0073] Figure 1 A diagram illustrating a scheme for synchronously processing audio during video restoration is shown.

[0074] like Figure 1 As shown, the video restoration module repairs the input video, for example, by removing targets from the video. This scheme can utilize visual information provided by the video restoration module to obtain a target mask (or location information) of the removed target, and then use this target mask to determine which direction the sound should be eliminated. In an exemplary embodiment, the target mask can be obtained by marking the target to be removed (i.e., a pedestrian).

[0075] Then, the scheme uses a neural network-based beamforming module to form a beam pointing towards the target mask based on the target mask and audio signal, thereby extracting sound from the specified direction while suppressing sound from other directions, thus achieving the extraction of the target's sound. Figure 2 As shown in the diagram. Ultimately, the system can remove the extracted target sound from the audio signal.

[0076] However, the above-mentioned method of directly processing audio synchronously during video restoration cannot accurately extract or delete the target's audio. The processed audio may retain the target's audio or incorrectly delete audio other than the target's audio, as in the following three cases:

[0077] (1) In most cases, the target mask is continuous between each image frame in the video. However, when the target is blocked or moves off the screen, the video restoration module cannot provide target mask information. This will result in the loss of the target's visual mask information, making it impossible to extract the target's sound. Therefore, when the target is blocked or moves off the screen, the target's sound cannot be removed. Figure 3A As shown in the image.

[0078] (2) When the target overlaps with other sound sources, the direction of the sound from the overlapping sound sources is the same as or similar to the direction of the target. This will cause the above scheme to incorrectly extract the sound from the overlapping sound sources, such as... Figure 3B As shown in the diagram. Furthermore, considering that the target's sound is mostly discontinuous, there will be times when the target's sound is absent. When the target overlaps with other sound sources at those times, the above method will incorrectly extract the sound from those other sources.

[0079] (3) In real-world scenarios, the target will produce a variety of sounds. For example, when the target is running, the target's sounds include not only talking but also footsteps or other sounds caused by movement. Using the above methods or other methods, it is impossible to obtain all the target's sounds, and therefore it is impossible to completely remove the target's sounds.

[0080] Therefore, this application proposes a method executed by an electronic device that can extract and remove the sound of a target when removing the target in video restoration, thereby improving the user experience. Specifically, this method first utilizes the spatial information of the target (including the target's azimuth, direction of movement, and speed) and the spatial distribution of the sound signal obtained from the video restoration module to analyze and estimate the target's bimodal mask in each audio frame. Then, by analyzing the entire input audio signal frame by frame, the global sound features and global motion trend features of the target are obtained. Based on the target's global motion trend features and the target's bimodal mask, a more accurate target bimodal mask is updated. The target sound mask for each audio frame is obtained by analyzing the features of each audio frame, the global sound features, and the updated target bimodal mask. Here, the target sound mask represents the portion of the target sound features in the audio features of the audio frame, which can also be understood as the percentage of information occupied by the target's sound in each audio frame of the first audio signal. Finally, other sound sources are restored by analyzing the sound wave propagation path (especially the sound of sound sources blocked by the target), thereby simulating the sound propagation and auditory perception of other sound sources when the target does not exist in the actual scene.

[0081] The following description of several optional embodiments illustrates the technical solutions of this disclosure and the technical effects produced by these solutions. It should be noted that the following embodiments can be referenced, learned from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.

[0082] Figure 4A This is a flowchart illustrating a method performed by an electronic device according to an exemplary embodiment of this application. Figure 4B This is a schematic diagram illustrating a network structure corresponding to a method performed by an electronic device according to an exemplary embodiment of this application.

[0083] The following is a combination of... Figure 4B Describe the function of each module shown.

[0084] like Figure 4B As shown, the first video is the video input to the video restoration module. The video restoration module can restore the first video according to operation instructions targeting a target (e.g., removal instructions or restoration instructions), such as removing the target. The encoder module performs feature encoding (e.g., performing discrete Fourier transform) on the first audio signal corresponding to the first video to obtain the encoded features of the first audio signal. These encoded features are high-dimensional feature vectors used to represent information in different dimensions of the speech.

[0085] The auditory-visual feature analysis module can obtain image-related information about the target from the video restoration module. In this application, "image-related information about the target" can also be referred to as "spatial information about the target." This image-related information may include the target's visual mask, depth information, and optical flow information. The target visual mask may represent the area where the user-selected removed target is located; the target's depth information may represent the distance between the target and the camera in the image / video; and the target's optical flow information may represent the target's direction and speed of motion. Furthermore, by utilizing the target visual mask and depth information, the target's spatial position and distance can be obtained. The auditory-visual feature analysis module can also obtain directional information of the first audio signal at different times, such as direction of arrival (DOA) information, from external sources. The auditory-visual feature analysis module analyzes this information to obtain the spatial distribution characteristics of the target's sound signal and the target's spatial distribution characteristics, for example... Figure 4CThe changes in the spatial distribution of the sound signal (e.g., the target represented by the solid box moves from the upper left to the top) and the spatial distribution characteristics of the target signal (e.g., the target mask represented by the dashed box moves from the upper left to the top) are shown between two consecutive moments. The auditory and visual feature analysis module finally obtains the target bimodal mask by jointly analyzing these features. In this application, the "target bimodal mask" can also be referred to as the "first mask". The target bimodal mask contains the spatial position information of the target, the motion trend of the mask, and the direction of the sound source motion.

[0086] The bimodal-bi-stage sound extraction module obtains the target sound mask for each audio frame using a bi-stage analysis approach. Simply put, this module first analyzes the entire first audio signal frame-by-frame to obtain the target's global sound features and global motion trend features. Then, it uses the obtained global motion trend features to correct and update the target bimodal mask for each audio frame, resulting in a more accurate target bimodal mask for each audio frame. Finally, based on the target's global sound features and the target bimodal mask for each audio frame, it analyzes the encoding features of the first audio signal to obtain the target sound mask for each audio frame. The target sound mask for each audio frame represents the percentage of information occupied by the target's sound in each audio frame, thus also allowing the calculation of the percentage of non-target sounds in each audio frame of the first audio signal (the sum of these two percentages is 1).

[0087] The decoder module decodes the encoded features of the first audio signal after removing the target sound features from the encoded features using the target sound mask, in order to obtain the second audio signal.

[0088] like Figure 4A As shown, in step S410, based on image-related information of the target in the first video, the first audio signal corresponding to the first video, and the direction information of the first audio signal, the target sound mask at each moment is obtained. In the following description of this application, "each moment" is described as the time corresponding to each audio frame of the first audio signal. However, this application is not limited to this. "Each moment" can also be the time corresponding to every two audios of the first audio signal, or it can be the time corresponding to each video frame of the first video. This application does not specifically limit this.

[0089] Specifically, step S410 may include: obtaining a target bimodal mask for each audio frame of the first audio signal based on the image-related information and the orientation information; and obtaining a target sound mask of the target in each audio frame of the first audio signal based on the target bimodal mask for each audio frame and the encoding features of the first audio signal. The following will refer to... Figure 5A and Figure 5B Describe the process of obtaining the target bimodal mask.

[0090] Figure 5A This is a flowchart illustrating an exemplary embodiment of the present application of obtaining a target bimodal mask based on target-related image information and orientation information. Figure 5B This is a schematic diagram illustrating the process by which an auditory-visual feature analysis module generates a target bimodal mask according to an exemplary embodiment of this application.

[0091] like Figure 5A As shown, in step S510, the spatial distribution characteristics of the sound signal are obtained for each audio frame of the first audio signal by normalizing and encoding the directional information of the first audio signal.

[0092] Specifically, such as Figure 5B As shown, the auditory-visual feature analysis module may include a target visual spatial encoding module, an audio spatial encoding module, a bimodal feature analysis module, and a motion trend analysis module. Each of these modules can be implemented by a network capable of feature processing, such as a convolutional network, an attention network, or a recurrent network. Step S510 can be executed by the audio spatial encoding module. The audio spatial encoding module can obtain the directional information of the first audio signal at different times (e.g., each audio frame) from, for example, an external source. In the following description, the directional information of the first audio signal at different times is taken as DOA information, which can be obtained by processing signals from multiple microphones using a MUSIC method, such as... Figure 6A As shown, but this application is not limited to this, the directional information of the first audio signal at different times can also be other directional information. The audio spatial coding module can use a neural network to normalize and encode the DOA information of the first audio signal in each audio frame, thereby mapping it into a feature space and obtaining the spatial distribution feature X of the sound signal for each audio frame. spa Among them, the spatial distribution characteristics of sound signals can represent the signal intensity and probability of existence in each direction in a space from 0 degrees to 360 degrees. By analyzing the DOA information frame by frame, the direction of motion of the sound source can be obtained. Figure 6B It shows the spatial distribution characteristics of the sound signal X spa An example diagram illustrates that, since the spatial distribution characteristics of the sound signal are obtained through DOA information analysis, however, the number of microphones on electronic devices is usually limited (e.g., typically 2-4 microphones), making it impossible to accurately locate the sound source; only a rough sound direction can be obtained. Therefore, the positions of each sound source are... Figure 6B The spatial distribution characteristics of the sound signal shown X spaThe representation is in the form of blocks or regions, rather than points. Furthermore, in this application, the audio spatial coding module can be implemented using a convolutional network; however, this application is not limited to this, and it can also be implemented using other networks with feature processing capabilities (such as recurrent networks, attention networks, etc.).

[0093] In step S520, the target spatial distribution features are obtained for each audio frame by encoding the target visual mask and depth information of the target in the image-related information.

[0094] Specifically, step S520 can be executed by the target visual spatial encoding module. For example... Figure 5B As shown, the target visual space encoding module can obtain the target visual mask and target depth information corresponding to each audio frame from an external source (e.g., a video restoration module). By using a neural network to encode the target visual mask and target depth information, it maps them to the sound spatial distribution, thereby obtaining the target spatial distribution feature K for each audio frame. spa ,like Figure 7A As shown in the diagram. In other words, the target visual spatial encoding module can determine the spatial location of the target for each audio frame by utilizing the target visual mask and the target's depth information. Furthermore, in this application, the target visual spatial encoding module can be implemented using any network capable of feature processing, such as convolutional networks, attention networks, or recurrent networks.

[0095] In detail, because visual information (e.g., target visual mask and target depth information) has a different feature distribution than sound, for example, visual information only includes the area directly in front of the electronic device (e.g., ... Figure 7B As shown in the image, visual information contains 360-degree omnidirectional information. Furthermore, the feature distribution of visual information differs from that of sound information because the visual feature distribution is based on feature vectors obtained from image / video algorithms, while the sound feature distribution is based on feature vectors obtained from audio algorithms; these two reside in different feature spaces. Additionally, while the visual information of the target in the first video is continuous (i.e., there is continuity between image frames), the sound of the target is often discontinuous. For example, in the case of a pedestrian, the sound may be intermittent and spaced out (e.g.,...). Figure 7C As shown in the image, the target may leave the screen or be obscured by other objects (i.e., non-targets), but the target's sound will still be heard. Furthermore, the target's position may overlap with other objects (i.e., non-targets), and the sound direction of the overlapping objects will be the same as the target's sound direction (e.g., [image of target]). Figure 7DAs shown in the diagram, these factors will lead to inconsistencies between the information expressed by visual features and sound features. Therefore, in this application, the target visual space encoding module utilizes a neural network to encode the target visual mask and the target's depth information, thereby mapping them to the sound spatial distribution, facilitating subsequent joint processing of multiple features. Furthermore, since the frame rate of video is typically 24–50 frames per second, while the frame rate of audio is typically 50 frames per second, the time span of video feature processing and audio feature processing is also consistent.

[0096] In step S530, based on the spatial distribution characteristics of the sound signal and the spatial distribution characteristics of the target, a target bimodal feature is obtained for each audio frame. In this application, the "target bimodal feature" can also be referred to as the "first feature".

[0097] Specifically, step S530 can be executed by the bimodal feature analysis module, which maps the visual features of the target to the spatial features of the sound distribution, thereby obtaining the target's bimodal features for each audio frame. For example... Figure 5B As shown, the bimodal feature analysis module can obtain the target spatial distribution features K for each audio frame from the target visual spatial coding module. spa And obtain the spatial distribution features X of the sound signal for each audio frame from the audio spatial coding module. spa Then, the spatial distribution characteristics X of the sound signal can be analyzed. spa and target spatial distribution characteristics K spa Feature processing is performed to obtain the target bimodal features F for each audio frame. dual This feature processing can be, for example, convolution processing, fusion processing, or splicing processing. In this application, for each audio frame, the target bimodal feature F... dual This indicates the target's spatial location and the direction of sound contained in, for example, a target visual mask obtained from a video restoration module. Furthermore, the "sound contained in the target visual mask" is not necessarily the target's sound; for example, when a sound source is located precisely at the spatial location corresponding to the target visual mask, the "sound contained in the target visual mask" will include the sound of that sound source. Figure 8 As shown in (a), there are two sound sources at the spatial location corresponding to the target visual mask. One of the sound sources may be a sound source other than the target. This is because, due to sound interference, the sound localization obtained from the DOA information is not accurate enough and does not match the target spatial distribution characteristics K. spa Similar sound source features will also be considered as candidate sound sources for the target. Furthermore, since this application only needs to consider the sound of the target, features K that are far from the target space in the feature space will be discarded. spa The sound information, such as Figure 8(b) only shows the distribution characteristics K in relation to the target space. spa The sound information at the corresponding spatial location.

[0098] In step S540, based on the target bimodal features and the optical flow information of the target in the image-related information, the target bimodal mask is obtained for each audio frame.

[0099] Specifically, step S540 can be executed by the motion trend analysis module. For example... Figure 5B As shown, the motion trend analysis module can obtain the target bimodal features F from the bimodal feature analysis module. dual The motion trend analysis module can analyze the target's dual-modal characteristics F. dual The target's spatial motion trend is determined for each audio frame based on the target's optical flow information (in other words, based on the target's bimodal features F). dual (Analyzing the target's spatial motion trend frame-by-frame using optical flow information), based on the target's dual-modal characteristics F dual For each audio frame, determine the motion trend of the sound source within the target visual mask (in other words, based on the target bimodal features F). dual Analyze the motion trend of sound sources within the target visual mask frame by frame, and based on the determination (or analysis) results of the spatial motion trend and the determination (or analysis) results of the sound source motion trend, obtain the target bimodal mask M for each audio frame. dual .

[0100] Specifically, the step of determining the spatial motion trend of the target for each audio frame based on the target bimodal features and the optical flow information may include: determining the spatial motion trend of the target based on the target bimodal features F. dual Based on the characteristics of the optical flow information, the spatial motion trend of the target in each sub-part of space for each audio frame is determined using visual information; in other words, by analyzing the target's bimodal features F... dual By analyzing the changes in the optical flow information of the target between two adjacent audio frames, the spatial motion trend (also known as "visual motion trend") of the target in each sub-part of space for each time moment is determined based on visual information. For example, the motion trend analysis module analyzes the target's bimodal features F between the current audio frame and the previous audio frame. dual By analyzing the changes in the optical flow information of the target, we can obtain the spatial motion trend of the target within each subspace of the target visual mask (that is, each sub-part of the target in space, such as the upper left, upper right, lower left, and lower right sub-parts) at the current audio frame (or the current moment corresponding to the current audio frame), based on visual information. Figure 9In the example shown, each subspace has a relatively consistent spatial motion trend (all to the right). However, due to the different instantaneous motion directions of each subspace, the presence of interference, and / or the bias of the algorithm estimation, the spatial motion trend of each subspace will be slightly different. For example, although the spatial motion trends of the upper left subspace and the lower left subspace are both to the right, the spatial motion trend of the upper left subspace is to the lower right, while the spatial motion trend of the lower left subspace is to the upper right.

[0101] Furthermore, based on the target bimodal features, the step of determining the sound source motion trend within the target visual mask for each audio frame may include: according to the target bimodal features F dual This involves determining the sound source motion trend of the target in each sub-part of space for each audio frame based on sound information; in other words, analyzing the target's bimodal features F. dual By analyzing the changes and correlations between two adjacent audio frames, the sound source motion trend (also referred to as "sound-based motion trend") of the target in each sub-part of space for each time moment is determined. For example, the motion trend analysis module analyzes the target's bimodal features F between the current audio frame and the previous audio frame. dual By analyzing the changes and correlations, we can obtain the sound source motion trends based on sound information in each subspace of the target's spatial distribution (i.e., each sub-part of the target in space, such as the upper left, upper right, lower left, and lower right sub-parts) at the current audio frame (or the current moment corresponding to the current audio frame), such as... Figure 9 In the example shown, each subspace has a relatively consistent spatial motion trend (all to the right). However, due to the different instantaneous motion directions of each subspace, the presence of interference, and / or the bias of the algorithm estimation, the spatial motion trend of each subspace will be slightly different. For example, although the spatial motion trends of the upper left subspace and the lower left subspace are both to the right, the spatial motion trend of the upper left subspace is more upward than that of the lower left subspace.

[0102] Based on the above motion trend analysis, the motion trend analysis module can obtain the target dual-modal mask M for each audio frame. dual ,like Figure 9 As shown in the image.

[0103] Based on the above references Figure 5AIn the process of obtaining the target bimodal mask, the auditory-visual feature analysis module processes visual and auditory information simultaneously. In this application, visual information allows the neural network to focus only on the auditory information within or near the target's visual mask. Simultaneously, auditory information ensures that the neural network does not miss the target's auditory features even when visual information is lost. Furthermore, the spatial motion trend based on visual information and the sound source motion trend based on auditory information enable the neural network to obtain a more accurate target trajectory, thereby extracting more accurate auditory features.

[0104] Return to reference Figure 4A and Figure 4B In step S410, after obtaining the target bimodal mask, the bimodal-bi-stage sound extraction module can obtain the target sound mask of the target in each audio frame of the first audio signal based on the target bimodal mask for each audio frame and the encoding features of the first audio signal.

[0105] The following will refer to Figure 10A and Figure 10B Describe the process of obtaining the target sound mask for each audio frame of the first audio signal.

[0106] Figure 10A This is a flowchart illustrating the process of obtaining a target sound mask for a target in each audio frame of a first audio signal according to an exemplary embodiment of this application. Figure 10B This is a schematic diagram illustrating the process by which a dual-modal, dual-stage sound extraction module, according to an exemplary embodiment of this application, obtains the target sound mask of a target in each audio frame of a first audio signal. For example... Figure 10B As shown, the bimodal-bi-stage sound extraction module may include a global information analysis module, a mask update module, and an audio mask estimation module.

[0107] like Figure 10A As shown, in step S1010, based on the target bimodal mask for each audio frame and the encoding features of the first audio signal, global sound features and global motion trend features of the target are obtained, wherein the global sound features represent the features of all sounds related to the target, and the global motion trend features represent the motion trajectory information of the target in the first video.

[0108] Specifically, step S1010 can be executed by the global information analysis module. For example... Figure 10B As shown, the global information analysis module can obtain the target bimodal mask M corresponding to each audio frame of the first audio signal from the auditory and visual feature analysis module. dual The encoder module obtains the encoded features of the first audio signal (also known as the mixed audio features X). mixThen, the global information analysis module can analyze the target bimodal mask M for each audio frame. dual Analyze all the encoded features of the first audio signal to obtain the global sound features S of the target. global and global motion trend characteristics P global For example, a target bimodal mask M for each audio frame of the first audio signal can be obtained through a neural network. dual The encoded features of the first audio signal are convolved and fused to compare the spatial information and acoustic characteristics between the target and the interference, thereby obtaining the global acoustic features S of the target. global and global motion trend characteristics P global In this application, when analyzing the entire first audio signal to obtain the global sound features S of the target. global and global motion trend characteristics P global In this case, only one global sound feature S can be obtained in the end. global A global motion trend feature P global In this application, the global sound feature S global Features representing all sounds related to the target, such as the target's voice, footsteps, and the sound of clothes rubbing together, etc., and global motion trend features P. global This indicates the target's motion trajectory information throughout the entire first video.

[0109] In another exemplary embodiment, when obtaining global sound features and global motion trend features, the global information analysis module can divide the first audio signal into multiple audio segments, and obtain a corresponding global sound feature and a global motion trend feature for each of the multiple audio segments. Specifically, the global information analysis module can decide whether to divide the first audio signal into multiple audio segments, or whether to increase or decrease the length of each audio segment divided from the first audio signal, based on at least one of the following: the performance of the previously obtained second audio signal, the actual usage scenario, the performance of the electronic device, etc. For example, if the previously obtained second audio signal still contains sound related to the target (i.e., the sound related to the target has not been completely removed), the global information analysis module can increase the length of each audio segment divided from the first audio signal (correspondingly, reduce the number of multiple audio segments), thereby ensuring the accuracy of the global sound features and global motion trend features obtained for each audio segment. For example, if the actual usage scenario is complex, involves multiple sound sources, or the overlap between the sound from other sources and the target sound exceeds a predetermined threshold (i.e., high overlap), the global information analysis module can increase the length of each audio signal segment divided from the first audio signal. This ensures the accuracy of the global sound features and global motion trend features obtained for each audio signal segment (i.e., guarantees the accuracy of global information). For another example, if the electronic device requires the processing delay of the audio signal to be less than a certain time, the global information analysis module can determine the length of each audio signal segment divided from the first audio signal based on this required delay.

[0110] Furthermore, when the first audio signal is divided into multiple audio segments, the global sound features and global motion trend features obtained for any one of the audio segments can be used as the initial global sound features and global motion trend features for the next audio segment. These initial global sound features and global motion trend features are then updated by analyzing the next audio segment, thereby obtaining the global sound features and global motion trend features for the next audio segment. However, this application is not limited to this; instead of using the global sound features and global motion trend features obtained for the previous audio segment as initial values, the global sound features and global motion trend features can be obtained directly for the next audio segment according to step S1010.

[0111] The above description mentions that the global information analysis module obtains the encoding features of the first audio signal from the encoder module; correspondingly, Figure 4AThe method shown may also include a step of obtaining coded features corresponding to the input first audio signal, and this step may include: extracting a feature vector from the first audio signal; and obtaining the coded features of the first audio signal by performing feature encoding on the extracted feature vector. See below for further details. Figure 11 and Figure 12 This describes the process by which the encoder module encodes the first audio signal to obtain encoded features.

[0112] Figure 11 This is a block diagram illustrating an encoder module according to an exemplary embodiment of this application. Figure 12 This is a network flowchart illustrating an encoder module according to an exemplary embodiment of this application. The encoder module obtains a high-dimensional feature vector by encoding an input first audio signal. Figure 11 As shown, the encoder module may include a feature extraction module, a sub-feature segmentation module, and multiple sub-encoders.

[0113] like Figure 11 As shown, the feature extraction module extracts features from the input first audio signal to obtain a feature vector in another dimension. For example, the Short-Time Fourier Transform (STFT) (e.g., 512-point STFT) can be used to perform feature extraction, that is, the first audio signal is framed, windowed, and subjected to STFT to obtain features in the frequency domain.

[0114] For example, for a first audio signal with a sampling rate of 16k and a duration of n seconds, there are L = n * 16000 sampling points. After performing an STFT with a window length of W = s_n sampling points (i.e., the number of sampling points per audio frame is s_n, and the overlap area between audio frames is s_n / 2 (i.e., 50% overlap), which is also a frame shift of W / 2), the number of frames k is k = L / (s_n / 2) - 1, and the number of frequency points in each audio frame is f = s_n / 2. By extracting the real and imaginary parts of the frequency domain, a feature vector with dimension [k, f] can be obtained. For example, for a first audio signal with a sampling rate of 16k and a duration of 4s, after performing an STFT with a window length of W = 512 sampling points (i.e., a frame shift of 256 sampling points), the number of frames is 249, and the number of frequency points f in each audio frame is s_n / 2 = 512 / 2 = 256. Each frequency point is represented by a real part and an imaginary part. Therefore, a feature vector with a dimension of [249, 256] can be obtained.

[0115] The above example uses STFT to perform feature extraction, but this application is not limited to this. Other feature extraction methods can also be used, such as using Convolutional Neural Networks (CNN) for feature extraction.

[0116] like Figure 12 As shown, the sub-feature segmentation module performs frequency band segmentation to divide the feature vector F extracted by the sub-feature segmentation module into multiple first sub-band feature vectors. For example, the 16kHz frequency band is divided into N sub-bands. Considering performance and model complexity, N can be equal to 4, 5, or 6 in one example; however, this application is not limited to this. For example, as shown... Figure 12 As shown, the 16kHz frequency band can be divided into 4 sub-bands. For the frequency domain data obtained in the previous step, the data of 256 frequency points of each audio frame (i.e., the extracted feature vectors) can be divided into 4 first sub-band feature vectors f1, f2, f3 and f4. The frequency points contained in each first sub-band feature vector are {1~32}, {33~64}, {65~128} and {129~256}, respectively, and the corresponding frequencies are 0~2kHz, 2kHz~4kHz, 4kHz-8kHz and 8kHz~16kHz, respectively.

[0117] like Figure 12 As shown, after dividing the extracted feature vector F into multiple first sub-band feature vectors, multiple second sub-band feature vectors are obtained by encoding each of the multiple first sub-band feature vectors using the corresponding sub-band encoder. The encoded features of the first audio signal include the multiple second sub-band feature vectors. Figure 11 and Figure 12 As shown, the number of first sub-band feature vectors, N, is 4. Correspondingly, there are N = 4 sub-encoders. Each first sub-band feature vector is input into a corresponding sub-encoder (e.g., a 2D convolutional neural network (2D-CNN)) for encoding, obtaining four second sub-band feature vectors x1, x2, x3, and x4. These multiple second sub-band feature vectors can be collectively referred to as the encoded features of the first audio signal. This application achieves parallel encoding, reduces model complexity, and improves model processing speed by dividing the extracted feature vectors into multiple first sub-band feature vectors and using different sub-encoders to encode the corresponding first sub-band feature vectors.

[0118] However, this application is not limited to this. In another exemplary embodiment of this application, the encoded features of the first audio signal can be obtained by directly encoding the extracted feature vectors using an encoder module without frequency band division. That is, a higher-dimensional encoded feature is obtained by encoding the full-band features using only one encoder. In the following description, the sound vectors or encoded features mentioned refer to the vectors or encoded features of a certain sub-band.

[0119] Return to reference Figure 10A and Figure 10BThe global information analysis module performs feature processing on the encoded features of the target bimodal mask and the first audio signal for each audio frame based on the following two principles, so that the visual information and sound information of the target are consistent: (1) When the target bimodal mask M dual When there are no sound-related features in the audio frame (where “sound-related features” are not features obtained by directly encoding the sound signal, but rather such as sound direction information, motion information, etc.), the global information analysis module can know that the sound components of the current audio frame are all non-target sounds (i.e., none of them are target sounds), so these non-targets can be marked as interference sources; (2) when the target dual-mode mask M dual When sound-related features exist in the audio, considering that the interfering sound source may overlap with the target's location, the sound components in the current audio frame can be marked as potentially being the target's sound.

[0120] After the global information analysis module processes each audio frame of the first audio signal, the target's global motion trend characteristics P are... global Information can be accumulated. Compared to the previous few audio frames, the global motion trend feature P obtained after processing all audio frames of the first audio signal is significantly improved. global This can represent a relatively smooth motion trajectory (or motion trend) of the target. Furthermore, in this application, the global motion trend feature P... global The feature size does not change over time, nor does it increase with the amount of information. In other words, the motion information of the target is continuously compressed into a feature space.

[0121] Furthermore, when the global information analysis module processes each audio frame of the first audio signal, the target's global sound features S global The information changes (i.e., is constantly updated). Because the location of the interfering sound source overlaps with the target, the sound generated by the interfering sound source in the target's visual mask may be incorrectly labeled as the target's sound in the previous moment (e.g., the previous audio frame). This means that the sound features of the interfering sound source overlapping with the target are incorrectly added to the target's global sound features S updated after processing the previous audio frame. global However, through frame-by-frame audio processing, when the target or interfering sound source moves and separates, the global information analysis module can correct the target and thus obtain the correct sound characteristics.

[0122] Figure 13 The diagram illustrates the process by which the global information analysis module, according to an exemplary embodiment of this application, performs feature processing on the encoded features of the target bimodal mask and the first audio signal.

[0123] like Figure 13As shown, for the current audio frame, after performing global information analysis based on the target bimodal mask (specifically, the spatial motion trend based on visual information) and the encoded features of the current audio frame, the global information analysis module can obtain the updated (i.e., updated using the current audio frame) global sound features S of the target up to the current audio frame. global And the global motion trend features P of the target accumulated (or updated) up to the current audio frame. global .from Figure 13 As can be seen from this, the global motion trend features P of the target accumulated up to the current audio frame... global The global motion trajectory represented is not very smooth, and the sound features of interfering sound sources overlapping with the target are incorrectly added to the target's global sound features S updated up to the current audio frame. global However, in the next audio frame, the target or interfering sound source moves and separates. Therefore, for that next audio frame, after performing global information analysis based on the target's bimodal mask and the encoded features of that next audio frame, the global information analysis module can obtain the updated global sound features S of the target up to that next audio frame. global And the global motion trend features P of the target accumulated (or updated) up to the next audio frame. global .from Figure 13 As can be seen, the global motion trend features P of the target accumulated up to the next audio frame... global The global motion trajectory represented is smoother, and the global sound feature S was incorrectly added to the audio frame preceding the next audio frame. global The sound features of the sound source in the image are not within the mask in the next audio frame, therefore the sound source is marked as an interfering sound source, and the global sound features S of the target are updated up to the next audio frame. global The sound characteristics of the interfering sound source are removed from the audio.

[0124] Return to reference Figure 10A In step S1020, based on the global sound features and the global motion trend features, the target sound mask of the target in each audio frame is determined.

[0125] Specifically, step S1020 may include: updating the target bimodal mask for each audio frame based on the global motion trend features.

[0126] In one exemplary embodiment of this application, step S1020 may be performed by the mask update module. For example... Figure 10B As shown, the mask update module can obtain the global motion trend feature P corresponding to the first audio signal from the global information analysis module. globalAnd obtain the target bimodal mask M for each audio frame in the first audio signal from the auditory and visual feature analysis module. dual Then, based on the global motion trend feature P corresponding to the first audio signal... global Update the target bimodal mask M for each audio frame frame by frame. dual This allows us to obtain the updated target bimodal mask for each audio frame. Updated target bimodal mask It can represent the target's spatial location, movement trend, and sound characteristics more accurately.

[0127] Specifically, due to sound interference, the target motion direction obtained based on sound characteristics may be deviated. Simultaneously, the optical flow characteristics of each pixel estimated visually will also deviate, causing the target's motion direction to jitter in each frame, and the represented target motion trend will be inaccurate. Therefore, the mask update module updates the target bimodal mask audio-level by audio based on global motion trend features. In an exemplary embodiment of this application, the step of updating the target bimodal mask for each audio frame based on the global motion trend features may include: adjusting at least one of the spatial motion trend and sound source motion trend in the target bimodal mask of the current audio frame by comparing the target's motion trend in the current audio frame with the global motion trend features and calculating trend consistency.

[0128] like Figure 14 As shown, the mask update module can update the mask by comparing the motion trend of the current audio frame with the global motion trend feature P. global Perform feature processing to compare and calculate trend consistency between them, thereby improving the target bimodal mask M of the current audio frame. dual At least one of the spatial motion trend and the acoustic motion trend is fine-tuned to obtain an updated target bimodal mask. However, this application is not limited thereto. In another exemplary embodiment of this application, the mask update module can calculate the motion trend of the current audio frame and the global motion trend feature P using a formula. global The comparison and trend consistency calculation between them are used to achieve the target bimodal mask M for the current audio frame. dual Fine-tuning of at least one of the spatial motion trend and the sound source motion trend, and thus obtaining an updated target bimodal mask. Updated target bimodal mask It can more accurately represent the spatial information of the target.

[0129] In addition, step S1020 may also include: determining the target sound mask of the target in each audio frame based on the coding features of the first audio signal, global sound features, and the updated target bimodal mask for each audio frame.

[0130] Specifically, the operation of determining the target sound mask for each audio frame based on the encoded features of the first audio signal, global sound features, and the updated target bimodal mask for each audio frame can be performed by... Figure 10B The audio mask estimation module performs this operation. It obtains the updated target bimodal mask for each audio frame from the mask update module. The audio mask estimation module can obtain accurate spatial information of the target. Simultaneously, based on the target's global acoustic features obtained from the global information analysis module, the audio mask estimation module can determine the target's acoustic characteristics and category. Therefore, when determining the target's acoustic mask in each audio frame, the audio mask estimation module can base its calculations on the target's global acoustic features S. global and the target bimodal mask updated for each audio frame Non-target sound features are eliminated from the encoded features of each audio frame of the first audio signal, and the target sound mask of the target is determined (or estimated) based on the encoded features of each audio frame after the elimination of non-target sound features.

[0131] like Figure 15 As shown, for the current audio frame, the audio mask estimation module can estimate the target bimodal mask based on the updated target bimodal mask for the current audio frame. More attention is paid to the spatial features corresponding to the mask, and based on the global acoustic features S of the target. global The sound features of interfering sound sources are removed from the encoded features of the current audio frame, thereby obtaining the target sound mask M of the target in the current audio frame. aduio Similarly, the audio mask estimation module can obtain the target sound mask for each audio frame of the first audio signal.

[0132] Based on the above references Figures 10A to 15 In the description, the bimodal-bi-stage sound extraction module first analyzes the global sound information (i.e., the target bimodal mask M). dual This allows us to obtain the target's global motion trajectory, thereby correcting the target's motion trend in each audio frame and obtaining an accurate sound mask for the target, avoiding the erroneous extraction of sound features from other interfering sound sources.

[0133] In this application, a recurrent neural network (RNN) can be used to implement the bimodal-bi-stage sound extraction module. However, this application is not limited to this and other networks with temporal processing capabilities (such as CNN, attention network, etc.) can also be used to implement the bimodal-bi-stage sound extraction module.

[0134] Return to reference Figure 4A After obtaining the target sound mask at each moment, in step S420, a second audio signal is obtained based on the target sound mask at each moment and the first audio signal, wherein the second audio signal does not include sounds related to the target. In this application, "sounds related to the target" can be any sound produced by the target, such as sounds made by the mouth, sounds produced by body movement (e.g., walking sounds, clapping sounds, rubbing sounds of clothes, etc.).

[0135] Specifically, the step of obtaining the second audio signal based on the target sound mask of the target at each time and the first audio signal may include: removing the sound features related to the target from the coding features of the first audio signal based on the target sound mask of the target at each audio frame to obtain the non-target sound signal features of the first audio signal; and obtaining the second audio signal based on the non-target sound signal features.

[0136] like Figure 4B As shown, by performing feature processing on the target sound mask and the encoded features of the first audio signal in each audio frame, the features of the sound related to the target can be removed from the encoded features of the first audio signal, thereby obtaining the non-target sound signal features of the first audio signal.

[0137] In one exemplary embodiment of this application, after obtaining the non-target sound signal features of the first audio signal, the non-target sound signal features can be directly decoded using a decoder to obtain the second audio signal y. other In other words, in this embodiment, the audio signal obtained after removing the target-related sound from the first audio signal can be directly output as the second audio signal.

[0138] In another exemplary embodiment of this application, after obtaining the non-target sound signal features of the first audio signal, the non-target sound signal features of the first audio signal can be repaired to obtain updated non-target sound signal features. Then, by decoding the updated non-target sound signal features, a second audio signal in which the non-target sound is enhanced or repaired can be obtained, thereby improving the user experience. Specifically, in a real-world scenario, when a target blocks the sound of a sound source, after removing the target, the sound of the sound source will no longer be blocked and will propagate directly to the electronic device in the direction blocked by the target. At this time, the sound intensity picked up by the electronic device will increase, and the user's auditory perception will also change. Therefore, for a video shot of a scene where a target blocks a sound source, after removing the target from the video, by enhancing or repairing the remaining non-target sound, a better auditory experience can be provided to the user. Therefore, in this alternative exemplary embodiment of the present application, the step of obtaining the second audio signal based on the non-target sound signal features may include: repairing the non-target sound signal features based on the updated target bimodal mask for each audio frame to obtain the updated non-target sound signal features; and obtaining the second audio signal by decoding the updated non-target sound signal features. (Refer to the following...) Figure 16 This will be described in detail.

[0139] Figure 16 The diagram shows a network structure corresponding to a method performed by an electronic device according to another exemplary embodiment of this application. Figure 17 This is a schematic diagram illustrating the structure of a repair module according to an exemplary embodiment of this application. Figure 16 The modules shown, excluding the repair module, are similar to... Figure 4B The modules shown are the same, so they will not be described again here.

[0140] like Figure 16 As shown, the repair module can obtain the non-target sound signal features X of each audio frame of the first audio signal. othter Furthermore, the updated target bimodal mask for each audio frame can be obtained from the bimodal-two-stage sound extraction module. Then, based on the updated target bimodal mask for each audio frame... For each audio frame, the non-target sound signal feature X othter Perform feature processing to obtain updated features of the non-target sound signal. Then the decoder module analyzes the updated non-target sound signal features. Decoding yields the second audio signal y. other See below for reference. Figure 17To describe in detail the characteristics of the updated non-target sound signal. The process.

[0141] like Figure 17 As shown, the repair module may include a sound propagation path analysis module and an audio repair module, each of which can be implemented by a network capable of feature processing, such as a convolutional network, attention network, or recursive network. In summary, the repair module performs repair operations by utilizing the Room Impulse Response (RIR), where RIR may include capturing and analyzing the acoustic properties of a room or environment to measure and model how sound waves interact with space (including reflection, reverberation, and echo). However, when the sound source is occluded by a target, the learned normal RIR can be utilized. com (i.e., RIR when the sound source is not blocked by the target) com This tool can be used to fix the sound of a sound source after the target has been removed from the video.

[0142] Specifically, the sound propagation path analysis module can be based on the updated target bimodal mask for each audio frame. Non-target sound signal features X othter Obtain the RIR of a non-target when it is not occluded by a target. com .

[0143] In one exemplary embodiment of this application, the step of obtaining the room impact response when the non-target is not occluded by the target may include: selecting signal features of a plurality of audio frames before and / or after the non-target is occluded by the target, based on the updated target bimodal mask for each audio frame; and obtaining the RIR of the non-target when it is not occluded by the target based on the signal features corresponding to the plurality of audio frames. In this application, in order to reduce the impact of changes (e.g., movement) of other objects in the space on the sound that needs to be repaired, when analyzing the RIR, only a plurality of audio frames adjacent to the moment when the non-target (i.e., a sound source other than the target) is occluded by the target may be considered. For example, the RIR may be analyzed by selecting a first plurality of audio frames before being occluded by the target and / or a second plurality of audio frames after being occluded by the target and subsequently no longer occluded by the target, based on the updated target bimodal mask for each audio frame, thereby obtaining the RIR when the sound source is not occluded by the target. In this application, depending on the actual situation, an appropriate number of the aforementioned first and / or second audio frames can be selected based on the updated target bimodal mask for each audio frame. For example, when the movement of non-targets (i.e., other sound sources or objects) is relatively slow, the number of the first and / or second audio frames selected based on the updated target bimodal mask for each audio frame can be larger, thereby obtaining a more accurate RIR. When the movement of non-targets (i.e., other sound sources or objects) is relatively fast, the number of the first and / or second audio frames selected based on the updated target bimodal mask for each audio frame can be smaller to reduce the impact of these non-targets on the RIR analysis operation. By analyzing the signal characteristics of these selected audio frames, the normal RIR of the positions of these other sound sources or objects in the current environment can be obtained. com .

[0144] The RIR of the non-target when it is not occluded by the target was obtained. com In the future, the audio restoration module can be based on the RIR obtained from the sound propagation path analysis module. com The non-target sound signal features are repaired to obtain updated non-target sound signal features. Specifically, the audio repair module can repair the non-target sound signal features X of each audio frame. othter and RIR com Perform feature processing to obtain updated non-target sound signal features for each audio frame.

[0145] Furthermore, in another exemplary embodiment of this application, after removing the target from the video, a new object (e.g., a cat) can be added to the video after the target has been removed. Accordingly, the audio restoration module can use a similar method to add the sound of the object to the second audio signal according to the category of the object.

[0146] The repair module generated updated non-target sound signal characteristics. Subsequently, the decoder module can use the updated non-target sound signal characteristics of each audio frame. Feature decoding is performed to obtain the second audio signal, i.e., to recover the time-domain signal. If Figure 4A and Figure 16 The encoder module in the middle adopts Figure 11 and Figure 12 The structure shown indicates that the decoder module can employ multiple decoders to implement decoding. See below for further details. Figure 18 This will be described in detail.

[0147] Figure 18 This is a block diagram illustrating a decoder module according to an exemplary embodiment of this application.

[0148] like Figure 18 As shown, the decoder module includes multiple sub-decoders, a feature merging module, and a time-domain signal recovery module. Figure 18 In the example shown, the updated non-target sound signal features obtained from the repair module It includes multiple sub-band features, and each sub-decoder performs feature processing on its corresponding sub-band features. Then, the feature merging module merges the processed sub-band features to facilitate subsequent feature transformation processing. Afterward, the time-domain signal recovery module performs audio signal recovery operations on the merged features to obtain the processed audio signal, i.e., the second audio signal y. other For example, the time-domain signal recovery module can use inverse short-time Fourier transform to perform sound signal recovery operations. However, this application is not limited to this. The time-domain signal recovery module can also use other feature transformation methods, such as using a CNN network to perform sound signal recovery operations. In this application, if the encoder module uses short-time Fourier transform for feature extraction, then in the decoder module, the time-domain signal recovery module can use inverse short-time Fourier transform to perform sound signal recovery operations. Correspondingly, if the encoder module uses a CNN network for feature extraction, then in the decoder module, the time-domain signal recovery module can use a CNN network to perform sound signal recovery operations.

[0149] The above description, with reference to the accompanying drawings, illustrates a method for obtaining a desired second audio signal by extracting and removing the target sound from an input first audio signal (i.e., a mixed audio signal). This method can be applied to various scenarios requiring video restoration, enabling the sound of the restored video to more ideally reflect the sound environment within the restored video. For example, it can be applied to mobile phone video restoration, extracting, separating, or eliminating sound within a target or region. The following examples illustrate two application scenarios of the method performed by an electronic device as described above; however, actual usage scenarios are not limited to these two scenarios.

[0150] Figure 19 The diagram illustrates the application of the method performed by an electronic device according to this application in a scenario where the target overlaps with other sound sources during video recording.

[0151] like Figure 19 As shown, during video recording, when a moving car (as the target to be removed in video restoration, the area where it is located is marked as a mask) overlaps with the position of a pedestrian at a certain moment, after applying the above-described method performed by an electronic device according to this application, only the sound of the car can be removed from the recorded video, while the sound of the pedestrian can be restored accordingly.

[0152] Specifically, firstly, the user can record a video of the aforementioned scenario. During video restoration, the user can select the car as the target to be removed. By utilizing the method described in this application, which is performed by an electronic device, the car's sound can be removed while the pedestrian's sound is preserved. Furthermore, when a pedestrian in the video is obscured by a car, the pedestrian's sound will be processed and restored accordingly, thereby improving the user's auditory experience.

[0153] Figure 20 A schematic diagram illustrates the application of the method performed by an electronic device according to this application in a scenario where the target is not in the screen during video recording.

[0154] like Figure 20 As shown, when a moving car (the area where the car is located, which is marked as a mask, is driving off the screen during video recording) is applied, the sound of the car can still be removed from the recorded video after applying the above-described method performed by the electronic device.

[0155] Specifically, firstly, the user can record a video of the above scene. When performing video repair, the user can select the car as the target to be removed. By using the method performed by the electronic device described in this application, the sound of the car can be removed from the entire video, even if the car drives off the screen, thereby improving the user's auditory experience.

[0156] The method performed by an electronic device as proposed in this application can determine the sound of a target based on its spatial location (e.g., target visual mask), movement direction, movement speed, and sound information, and extract or eliminate the target's sound. Furthermore, considering the spatial impact of the target's removal on other sound sources in the video, the method performed by the electronic device can also repair the sounds of these other sound sources. Moreover, considering the filling of other objects after removing the target image, the method performed by the electronic device can add the sound of that type of object to the video's audio. The method proposed in this application is applicable not only to audio restoration but also to speech enhancement and speech analysis.

[0157] This disclosure also provides an electronic device including at least one processor, and optionally, at least one transceiver coupled to the at least one processor and / or at least one memory, wherein the at least one processor is configured to perform the steps of the method provided in any optional embodiment of this disclosure.

[0158] Figure 21 The diagram shows a structural schematic of an electronic device to which an embodiment of the present invention applies, such as... Figure 21 As shown, Figure 21 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, each of the processor 4001, memory 4003, and transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this disclosure. Optionally, the electronic device may be a first network node, a second network node, or a third network node.

[0159] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0160] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 21 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0161] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.

[0162] The memory 4003 is used to store computer programs or executable instructions that execute the embodiments of this disclosure, and is controlled by the processor 4001 to execute them. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0163] This disclosure provides a computer-readable storage medium storing a computer program or instructions that, when executed by at least one processor, can perform or implement the steps and corresponding content of the aforementioned method embodiments.

[0164] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.

[0165] The terms “first,” “second,” “third,” “fourth,” “1,” “2,” etc. (if present) in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in a sequence other than that shown in the figures or text.

[0166] It should be understood that although arrows indicate various operation steps in the flowcharts of the embodiments of this disclosure, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of the embodiments of this disclosure, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured as required, and the embodiments of this disclosure do not limit this.

[0167] The above text and accompanying drawings are provided as examples only to help the reader understand this disclosure. They are not intended and should not be construed as limiting the scope of this disclosure in any way. Although certain embodiments and examples have been provided, it will be apparent to those skilled in the art, based on the content disclosed herein, that changes can be made to the illustrated embodiments and examples, and other similar implementations based on the technical concept of this disclosure can be adopted without departing from the scope of this disclosure, and these modifications and modifications are also within the protection scope of the embodiments of this disclosure.

Claims

1. A method performed by an electronic device, comprising: Based on the image-related information of the target in the first video, the first audio signal corresponding to the first video, and the direction information of the first audio signal, the target sound mask of the target at each moment is obtained; Based on the target sound mask and the first audio signal at each time point, a second audio signal is obtained, wherein the second audio signal does not include sound related to the target.

2. The method according to claim 1, wherein, The image-related information includes the target's visual mask, depth information, and optical flow information.

3. The method according to claim 1, wherein, Based on image-related information of the target in the first video, the first audio signal corresponding to the first video, and the direction information of the first audio signal, the target sound mask at each moment is obtained, including: Based on the image-related information and the direction information, a first mask is obtained for each audio frame of the first audio signal; and Based on the first mask for each audio frame and the encoding features of the first audio signal, the target sound mask for each audio frame of the first audio signal is obtained.

4. The method according to claim 3, wherein, Based on the image-related information and the direction information, a first mask is obtained for each audio frame of the first audio signal, including: By normalizing and encoding the directional information, the spatial distribution characteristics of the sound signal are obtained for each audio frame; By encoding the target visual mask and depth information of the target in the image-related information, the target spatial distribution features are obtained for each audio frame; Based on the spatial distribution characteristics of the sound signal and the spatial distribution characteristics of the target, a first feature is obtained for each audio frame; Based on the first feature and the optical flow information of the target in the image-related information, the first mask is obtained for each audio frame.

5. The method according to claim 4, wherein, Based on the spatial distribution features of the sound signal and the spatial distribution features of the target, a first feature is obtained for each audio frame, including: By performing feature processing on the spatial distribution features of the sound signal and the spatial distribution features of the target, the first feature is obtained for each audio frame, wherein the first feature represents the position of the target in space and the direction of the sound contained in the target visual mask.

6. The method according to claim 4, wherein, Based on the first feature and the optical flow information of the target in the image-related information, the first mask is obtained for each audio frame, including: Based on the first feature and the optical flow information, the spatial motion trend of the target is determined for each audio frame; Based on the first feature, the motion trend of the sound source within the target visual mask is determined for each audio frame; Based on the determination results of the spatial motion trend and the sound source motion trend, the first mask is obtained for each audio frame.

7. The method according to claim 6, wherein, Based on the first feature and the optical flow information, the spatial motion trend of the target is determined for each audio frame, including: Based on the first feature and the features of the optical flow information, the spatial motion trend of the target in each sub-part of space for each audio frame is determined based on visual information.

8. The method according to claim 6, wherein, Based on the first feature, the motion trend of the sound source within the target visual mask is determined for each audio frame, including: Based on the first feature, the sound source motion trend of the target in each sub-part of space for each audio frame is determined based on sound information.

9. The method according to claim 3, wherein, Based on the first mask for each audio frame and the encoding features of the first audio signal, the target sound mask for each audio frame of the first audio signal is obtained, including: Based on the first mask and the encoding features of the first audio signal for each audio frame, global sound features and global motion trend features of the target are obtained, wherein the global sound features represent the features of all sounds related to the target, and the global motion trend features represent the motion trajectory information of the target in the first video; Based on the global sound features and the global motion trend features, the target sound mask for the target is determined in each audio frame.

10. The method according to claim 9, wherein, Based on the global sound features and the global motion trend features, the target sound mask for each audio frame is determined, including: Based on the global motion trend features, the first mask is updated for each audio frame; Based on the encoding features of the first audio signal, the global sound features, and the updated first mask for each audio frame, the target sound mask for each audio frame is determined.

11. The method according to claim 10, wherein, Based on the encoded features of the first audio signal, the global sound features, and the updated first mask for each audio frame, the target sound mask of the target in each audio frame is determined, including: Based on the global sound features and the updated first mask for each audio frame, non-target sound features are eliminated from the encoded features of each audio frame of the first audio signal; Based on the encoded features of each audio frame after eliminating non-target sound features, the target sound mask for each audio frame is determined.

12. The method according to claim 3, wherein, Based on the target sound mask and the first audio signal at various times, a second audio signal is obtained, including: Based on the target, the target sound mask in each audio frame removes the sound features related to the target from the coding features of the first audio signal to obtain the non-target sound signal features of the first audio signal; A second audio signal is obtained based on the characteristics of the non-target sound signal.

13. The method according to claim 12, wherein, Based on the characteristics of the non-target sound signal, a second audio signal is obtained, including: The non-target sound signal features are repaired based on the updated first mask for each audio frame to obtain the updated non-target sound signal features. The second audio signal is obtained by decoding the updated features of the non-target sound signal.

14. The method according to claim 13, wherein, The updated first mask is obtained through the following operation: Global motion trend features are obtained based on the encoding features of the first mask and the first audio signal for each audio frame; Based on the global motion trend features, the first mask is updated for each audio frame.

15. The method according to claim 10 or 14, wherein, Based on the global motion trend features, the first mask is updated for each audio frame, including: By comparing the motion trend of the target in the current audio frame with the global motion trend features and calculating the trend consistency, at least one of the spatial motion trend and sound source motion trend in the first mask of the current audio frame is adjusted.

16. The method according to claim 13, wherein, The non-target sound signal features are repaired based on the updated first mask for each audio frame to obtain the updated non-target sound signal features, including: Based on the updated first mask for each audio frame and the non-target sound signal features, the room impact response of the non-target when it is not obscured by the target is obtained; The non-target sound signal features are repaired based on the room impact response to obtain updated non-target sound signal features.

17. The method according to claim 16, wherein, Based on the updated first mask for each audio frame and the non-target sound signal features, the room impact response of the non-target when it is not obscured by the target is obtained, including: Based on the updated first mask for each audio frame, signal features of a plurality of audio frames before and / or after the non-target is occluded by the target are selected from the non-target sound signal features. Based on the signal characteristics corresponding to the multiple audio frames, the room impact response of the non-target when it is not obstructed by the target is obtained.

18. An electronic device comprising: At least one processor; as well as At least one memory that stores computer-executable instructions. Wherein, when the computer-executable instructions are executed by the at least one processor, they cause the at least one processor to perform the method as described in any one of claims 1 to 17.

19. A computer-readable storage medium for storing instructions, wherein, When the instruction is executed by at least one processor, it causes the at least one processor to perform the method as described in any one of claims 1 to 17.