Method executed by electronic equipment, electronic equipment and storage medium

By adaptively adjusting the number and position of virtual microphones, the problem of excessively wide beamwidth caused by the limited number of microphones in mobile phones is solved, achieving more accurate target sound extraction and interference suppression.

CN121600947APending Publication Date: 2026-03-03BEIJING SAMSUNG TELECOM R&D CENT +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411139749.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

In existing technologies, due to the limited number of microphones in mobile phones, the beamwidth is wide, making it difficult to accurately extract the target sound that the user is interested in, and it is easily affected by interference sounds.

Method used

By adaptively creating virtual microphones of varying numbers and positions based on DOA and target direction information, a narrower beam is formed, reducing sound interference outside the target area and improving sound extraction performance.

Benefits of technology

It enables adaptive adjustment of the number and position of microphones under different scenarios and postures to form a narrower beam, thereby improving the accuracy of target sound extraction and the ability to suppress interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600947A_ABST
    Figure CN121600947A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method executed by electronic equipment, the electronic equipment and a storage medium, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring an input first audio signal; based on direction of arrival (DOA) information of a first audio signal, target direction information and spatial information of a physical microphone (RM) of the electronic equipment, obtaining the number and positions of virtual microphones (VM); a second audio signal is obtained based on the first audio signal and the number and location of VMs. Optionally, the method performed by the electronic device may be performed using an artificial intelligence model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of signal processing technology, and more specifically, to a method for processing audio signals performed by an electronic device, an electronic device, and a storage medium. Background Technology

[0002] Currently, audio scaling technology can select sound from the desired direction while reducing sound from other directions. Audio scaling can be used in conjunction with video (e.g., video captured by a camera), where audio can be scaled simultaneously with the video image. However, due to hardware limitations preventing the placement of too many microphones on mobile phones, current phones typically have only 3-4 microphones. With existing technology, fewer microphones result in a wider beamwidth. If the beamwidth exceeds the area of ​​interest to the user, some interfering sound will be captured, thus failing to accurately deliver the desired target sound.

[0003] How to accurately provide the desired sound to users and meet their needs is a technical problem that those skilled in the art have been working hard to study. Summary of the Invention

[0004] In order to at least solve the above-mentioned problems existing in the prior art, the present invention provides a method performed by an electronic device, an electronic device, and a storage medium.

[0005] According to a first aspect of the embodiments of this application, a method performed by an electronic device is provided, comprising: obtaining an input first audio signal; obtaining the number and position of virtual microphones (VMs) based on the direction of arrival (DOA) information, target direction information, and spatial information of the physical microphones (RM) of the electronic device; and obtaining a second audio signal based on the first audio signal and the number and position of the VMs.

[0006] Optionally, the step of obtaining the number and location of virtual microphones (VMs) based on the direction of arrival (DOA) information of the first audio signal, the target direction information, and the spatial information of the physical microphone (RM) of the electronic device includes: for each frequency unit under each time unit, performing the following operations: obtaining a first signal spatial distribution feature based on the DOA information and the target direction information; and determining the number and location of VMs based on the first signal spatial distribution feature and the spatial information of the RM.

[0007] Optionally, the step of obtaining the first signal spatial distribution feature based on the DOA information and the target direction information includes: obtaining a third feature vector by feature mapping of the DOA information; obtaining a fourth feature vector by encoding the target direction information; obtaining a second signal spatial distribution feature based on the third feature vector and the fourth feature vector; and obtaining the first signal spatial distribution feature based on the historical statistical information of the first audio signal and the second signal spatial distribution feature.

[0008] Optionally, the historical statistical information includes speech signal feature information and noise signal feature information of the last frame in the previous time unit. The step of obtaining the first signal spatial distribution feature based on the historical statistical information of the first audio signal and the second signal spatial distribution feature includes: calculating the correlation of each RM with respect to the audio signal within the target image region based on the speech signal feature information and the second signal spatial distribution feature of the last frame in the previous time unit, and calculating the correlation of each RM with respect to the audio signal outside the target image region based on the noise signal feature information and the second signal spatial distribution feature of the last frame in the previous time unit, to obtain the first signal spatial distribution feature.

[0009] Optionally, the step of determining the number and location of VMs based on the spatial distribution characteristics of the first signal and the spatial information of the RM includes: determining the number and location of VMs based on the spatial distribution characteristics of the first signal, the spatial information of the RM, and the state information of the electronic device.

[0010] Optionally, the step of determining the number and location of VMs based on the spatial distribution characteristics of the first signal, the spatial information of the RM, and the state information of the electronic device includes: encoding the state information of the electronic device into a state vector; determining the microphone aperture based on the spatial information of the RM; and determining the number and location of VMs based on the spatial distribution characteristics of the first signal, the microphone aperture, and the state vector.

[0011] Optionally, the number and position of VMs determined for each frequency unit under each time unit constitute a three-dimensional VM position feature vector in terms of time direction, frequency direction, and spatial direction. The step of obtaining the second audio signal based on the first audio signal and the number and position of VMs includes: obtaining a first feature vector corresponding to the first audio signal; and obtaining the second audio signal based on the first feature vector, the VM position feature vector, and the spatial information of the RM.

[0012] Optionally, the step of obtaining the second audio signal based on the first feature vector, the VM position feature vector, and the spatial information of the RM includes: determining the relative position information between the VM and the RM based on the VM position feature vector and the spatial information of the RM; determining the directivity of each VM based on the relative position information between the VM and the RM and the directivity of the RM; obtaining a second feature vector based on the relative position information between the VM and the RM, the first feature vector, and the directivity of each VM; and obtaining the second audio signal based on the first feature vector and the second feature vector.

[0013] Optionally, the step of obtaining the second audio signal based on the first feature vector and the second feature vector includes: obtaining a mask related to the second audio signal within the target image region based on the first feature vector and the second feature vector; and obtaining the second audio signal based on the mask and the first feature vector.

[0014] Optionally, the step of obtaining a mask related to a second audio signal within a target image region based on a first feature vector and a second feature vector includes: merging the first feature vector and the second feature vector along a spatial direction to obtain a merged feature vector; and performing feature processing sequentially in a first direction, a second direction, and a third direction on the merged feature vector to obtain the mask, wherein the first direction, the second direction, and the third direction are one of a spatial direction, a temporal direction, and a frequency direction, respectively.

[0015] Optionally, the step of performing feature processing in the frequency direction for each time unit includes: for each frequency unit, performing the following operations: based on the initial hidden state of each RM in the current frequency unit, processing the features corresponding to the current frequency unit in the input feature vector corresponding to each RM to obtain the processed features and the processed hidden state corresponding to each RM, wherein the processed hidden state corresponding to each RM is used as the initial hidden state of each RM in the next frequency unit; based on the initial hidden state of each VM in the current frequency unit, processing the features corresponding to the current frequency unit in the input feature vector corresponding to each VM to obtain the processed features and the processed hidden state corresponding to each VM; obtaining the global hidden state by performing feature processing on the processed hidden state corresponding to each VM; obtaining the initial hidden state of each VM in the next frequency unit by performing feature processing on the feature vector corresponding to the next frequency unit in the VM position feature vector and the global hidden state.

[0016] Optionally, the step of performing feature processing in the temporal direction for each frequency unit includes: for each frame in each time unit, performing the following operations: based on the initial hidden state of each RM in the current frame, processing the features corresponding to the current frame in the input feature vector corresponding to each RM to obtain the processed features and the processed hidden state corresponding to each RM, wherein the processed hidden state corresponding to each RM is used as the initial hidden state of each RM in the next frame; based on the initial hidden state of each VM in the current frame, processing the features corresponding to the current frame in the input feature vector corresponding to each VM to obtain the processed features and the processed hidden state corresponding to each VM; obtaining the global hidden state by performing feature processing on the processed hidden state corresponding to each VM; obtaining the initial hidden state of each VM in the next frame by performing feature processing on the feature vector corresponding to the next frame in the VM position feature vector and the global hidden state.

[0017] Optionally, the step of performing feature processing in the spatial direction for each time unit includes: for each RM, performing the following operations: based on the initial hidden state of the current RM in each frequency unit, processing the features corresponding to each frequency unit in the input feature vector corresponding to the current RM to obtain the processed features and the processed hidden state corresponding to each frequency unit, wherein the processed hidden state corresponding to each frequency unit is used as the initial hidden state of each frequency unit in the next RM; for each VM, performing the following operations: based on the initial hidden state of the current VM in each frequency unit, processing the features corresponding to each frequency unit in the input feature vector corresponding to the current VM to obtain the processed features and the processed hidden state corresponding to each frequency unit, obtaining the global hidden state by performing feature processing on the processed hidden state corresponding to each frequency unit, and obtaining the initial hidden state of the next VM in each frequency unit by performing feature processing on the feature vector corresponding to the next VM in the VM position feature vector and the global hidden state.

[0018] Optionally, a frequency unit is a frequency point or a sub-band, wherein a sub-band includes multiple frequency points.

[0019] Optionally, a time unit is one or more frames.

[0020] Optionally, the step of obtaining the second audio signal based on the mask and the first feature vector includes: obtaining a fifth feature vector based on the mask and the first feature vector; and obtaining the second audio signal by performing an audio signal recovery operation on the fifth feature vector.

[0021] Optionally, the step of obtaining the fifth feature vector based on the mask and the first feature vector includes: processing each sub-band mask in the mask and the corresponding second sub-band feature vector in the first feature vector to obtain multiple third sub-band feature vectors; performing feature transformation on the multiple third sub-band feature vectors to obtain multiple predicted features; and merging the multiple predicted features to obtain the fifth feature vector.

[0022] According to a second aspect of the present application, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform a method performed by the electronic device as described above.

[0023] According to a third aspect of the present application, a computer-readable storage medium for storing instructions is provided, wherein when the instructions are executed by at least one processor, they cause the at least one processor to perform a method performed by an electronic device as described above.

[0024] The beneficial effects of the technical solutions provided in this application will be explained in the following text in conjunction with specific optional embodiments, or can be learned from the description of the embodiments, or can be learned through the implementation of the embodiments. Attached Figure Description

[0025] To more clearly and easily illustrate and understand the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0026] Figure 1 This is a schematic diagram illustrating beamforming technology based on signal processing.

[0027] Figure 2 This is a schematic diagram illustrating the beamwidth formed using different numbers of microphones.

[0028] Figure 3 This is a schematic diagram illustrating the system architecture of a neural network-based VM estimation technique.

[0029] Figure 4 This is a flowchart illustrating a method performed by an electronic device according to an exemplary embodiment of this application.

[0030] Figure 5 This is a schematic diagram illustrating a method performed by an electronic device according to an exemplary embodiment of this application.

[0031] Figure 6This is a flowchart illustrating the process of obtaining the number and location of VMs for each frequency unit under each time unit according to an exemplary embodiment of this application.

[0032] Figure 7A This is a schematic diagram illustrating the structure of a VM analysis and feature generation module according to an exemplary embodiment of this application.

[0033] Figure 7B This is a diagram illustrating an example of spatial signal distribution.

[0034] Figure 7C This is a schematic diagram illustrating a method performed by an electronic device according to an exemplary embodiment of this application.

[0035] Figure 8 This is a diagram illustrating an example of obtaining DOA information using the Multiple Signal Classification (MUSIC) algorithm.

[0036] Figure 9 This is a schematic diagram illustrating the process of obtaining the spatial distribution characteristics of a first signal according to an exemplary embodiment of this application.

[0037] Figure 10 This is a diagram illustrating examples of first and second signal spatial distribution features according to exemplary embodiments of this application.

[0038] Figure 11 This is a schematic diagram illustrating the correlation between two consecutive frames of audio signals.

[0039] Figure 12A This is a schematic diagram illustrating the effect of microphone aperture on beamwidth.

[0040] Figure 12B This is a schematic diagram illustrating the effect of the folded state of a foldable electronic device on the microphone aperture.

[0041] Figure 13 This is a schematic diagram illustrating the beam formed by the microphone array.

[0042] Figure 14 This is a flowchart illustrating a process for obtaining a second audio signal based on a first audio signal and the number and location of VMs, according to an exemplary embodiment of this application.

[0043] Figure 15 This is a block diagram illustrating an encoder module according to an exemplary embodiment of this application.

[0044] Figure 16 This is a network flowchart illustrating an encoder module according to an exemplary embodiment of this application.

[0045] Figure 17 This is a schematic diagram illustrating the process of determining the directionality of a VM.

[0046] Figure 18 This is a schematic diagram illustrating the process of performing feature processing in the frequency direction.

[0047] Figure 19 This is a system block diagram illustrating a sound extraction module according to an exemplary embodiment of this application.

[0048] Figure 20 This is a schematic diagram illustrating a process of performing feature processing in the frequency direction according to an exemplary embodiment of this application.

[0049] Figure 21 This is a schematic diagram illustrating a process of performing feature processing in the time direction according to an exemplary embodiment of this application.

[0050] Figure 22 This is a schematic diagram illustrating a process of performing feature processing in a spatial direction according to an exemplary embodiment of this application.

[0051] Figure 23 This is a block diagram illustrating a decoder module according to an exemplary embodiment of this application.

[0052] Figure 24 This is a network flowchart illustrating a decoder module according to an exemplary embodiment of this application.

[0053] Figure 25 This is a schematic diagram illustrating the process of training a sound extraction module according to an exemplary embodiment of this application.

[0054] Figure 26 This is a schematic diagram illustrating the process of jointly training a third neural network and a VM analysis and feature generation module according to an exemplary embodiment of this application.

[0055] Figure 27 This is a schematic diagram illustrating the application of the method performed by an electronic device according to this application in a scenario of recording video.

[0056] Figure 28 This is a schematic diagram illustrating the structure of an electronic device to which an exemplary embodiment of this application applies. Detailed Implementation

[0057] The following description, with reference to the accompanying drawings, is provided to aid in a thorough understanding of the various embodiments of this disclosure as defined by the claims and their equivalents. This description includes various specific details to aid understanding but should be considered exemplary only. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of this disclosure. Furthermore, for clarity and brevity, descriptions of well-known functions and structures may be omitted.

[0058] The terms and wording used in the following description and claims are not limited to their dictionary meanings, but are merely used by the inventors to enable a clear and consistent understanding of this disclosure. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of this disclosure is for illustrative purposes only and not for limiting the purpose of this disclosure as defined in the appended claims and their equivalents.

[0059] It should be understood that the singular forms of “a,” “an,” and “the” can also include plural references unless the context clearly indicates otherwise. Thus, for example, the reference to “component surface” includes referring to one or more such surfaces. When we say that an element is “connected” or “coupled” to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element are connected through an intermediate element. Furthermore, the use of “connected” or “coupled” herein can include wireless connections or wireless couplings.

[0060] The terms “comprising” or “may include” refer to the presence of a corresponding disclosed function, operation, or component that may be used in the various embodiments of this disclosure, rather than limiting the presence of one or more additional functions, operations, or features. Furthermore, the terms “comprising” or “having” may be interpreted as indicating certain characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof, but should not be construed as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof.

[0061] The term "or" as used in the various embodiments of this disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not explicitly defined, the multiple items may refer to one, more, or all of the multiple items. For example, the description "parameter A includes A1, A2, A3" can be implemented as parameter A includes A1 or A2 or A3, or it can be implemented as parameter A includes at least two of the three items A1, A2, and A3.

[0062] Unless otherwise defined, all terms used in this disclosure (including technical or scientific terms) have the same meaning as understood by one of those skilled in the art to which this disclosure pertains. Common terms as defined in dictionaries are to be interpreted as having a meaning consistent with the context in the relevant technical field and should not be interpreted ideally or overly formally unless expressly defined in this disclosure.

[0063] At least some of the functions of the device or electronic device provided in this disclosure embodiment can be implemented by an AI model, such as implementing at least one module of a plurality of modules of the device or electronic device by an AI model. AI-related functions can be executed by non-volatile memory, volatile memory, and a processor.

[0064] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as central processing unit (CPU), application processor (AP), etc., or pure graphics processing unit, such as graphics processing unit (GPU), vision processing unit (VPU), and / or AI-specific processors, such as neural processing unit (NPU).

[0065] The one or more processors control the processing of input data based on predefined operating rules or artificial intelligence (AI) models stored in non-volatile and volatile memory. These predefined operating rules or AI models are provided through training or learning.

[0066] Here, "providing through learning" refers to obtaining predefined operating rules or an AI model with desired characteristics by applying a learning algorithm to multiple learning datasets. This learning can be performed within the device or electronic device itself, in which the AI ​​is executed according to the embodiment, and / or can be implemented via a separate server / system.

[0067] AI models can contain multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network computations by calculating the input data of that layer (such as the computation results of the previous layer and / or the input data of the AI ​​model) and the multiple weight values ​​of the current layer. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q-networks.

[0068] A learning algorithm is a method of training a predetermined target device (e.g., a robot) using multiple learning data sets to enable, allow, or control the target device to make determinations or predictions. Examples of such learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0069] The methods provided in this disclosure may relate to one or more fields in the technical fields of speech, language, image, video, or data intelligence.

[0070] Optionally, in the context of speech or language, in the method performed by an electronic device according to this disclosure, a speech signal as an analog signal may be received via a speech input device (e.g., a microphone), and the speech portion may be converted into computer-readable text using an Automatic Speech Recognition (ASR) model. The user's utterance intent can be obtained by interpreting the converted text using a Natural Language Understanding (NLU) model. The ASR model or NLU model may be an artificial intelligence model. The artificial intelligence model may be processed by a dedicated artificial intelligence processor designed in a hardware architecture specified for processing the artificial intelligence model. Language understanding is a technique for recognizing and applying / processing human language / text, including, for example, natural language processing, machine translation, dialogue systems, question answering, or speech recognition / synthesis.

[0071] Optionally, when dealing with the field of images or videos, in the method performed by an electronic device according to this disclosure, output data can be obtained by using image data as input data for an artificial intelligence model. The methods of this disclosure can relate to the field of visual understanding in artificial intelligence technology, which is a technology for recognizing and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.

[0072] Optionally, in the field of data intelligence processing, in the method performed by an electronic device according to this disclosure, during the reasoning or prediction phase, an artificial intelligence model can be used to perform prediction by using real-time input data. The processor of the electronic device can perform preprocessing operations on the data to transform it into a form suitable for use as input to the artificial intelligence model. Reasoning and prediction are techniques for making logical inferences and predictions by determining information, including, for example, knowledge-based reasoning, optimization prediction, preference-based planning, or recommendation.

[0073] In this application, the artificial intelligence model can be obtained through training. Here, "obtained through training" means obtaining a predefined operational rule or artificial intelligence model configured to perform desired features (or objectives) by training a basic artificial intelligence model with multiple training data using a training algorithm. The artificial intelligence model may include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values, and neural network computation is performed by calculating the results of the previous layer and the multiple weight values.

[0074] There are many types of regional sound extraction techniques, mainly divided into beamforming techniques based on signal processing and virtual microphone (VM) estimation techniques based on neural networks.

[0075] Beamforming technology based on signal processing can be considered a type of spatial filter. For example... Figure 1 As shown, the technique can use the following equation (1) to process the signals received by multiple microphones according to the weighting vector w to obtain an enhanced target signal y.

[0076]

[0077] Where k represents time, z i This represents the signal received by the i-th microphone.

[0078] However, for beamforming technology based on signal processing, the hardware layout of mobile phones limits the number of microphones that can be placed on the phone. For example... Figure 2 As shown, fewer microphones result in a wider beamwidth. Box 201 indicates that an array with four microphones forms a beamwidth of 60 degrees, while an array with three microphones will form an even wider beamwidth. If the beamwidth exceeds the area of ​​interest for the user, interfering sounds will be captured, and the desired target sound cannot be accurately delivered. Therefore, the challenge of using signal processing-based beamforming technology on mobile phones with fewer microphones needs to be considered.

[0079] Furthermore, neural network-based VM estimation techniques directly estimate VM signals in the time domain. This technique relies on a supervised learning framework and builds a model that can predict VM signals based on observations from a real microphone (RM). Specifically, such as... Figure 3 As shown, under the constraint of the received signal at actual location 2, the microphone signal at location 2 is estimated using the time-domain signals of RM 1 and RM 3, where the positions of all microphones are preset during model training. Figure 3 As shown, the encoder performs feature encoding on the signals ① and ③ received by the input RM 1 and RM 3, the feature processing module processes the encoded features, and the decoder estimates the signal ② of VM 2, thereby realizing the virtual generation of the signal ② received by VM 2 based on the signals ① and ③ received by RM 1 and RM 3.

[0080] However, for neural network-based VM estimation techniques, since the number and location of VMs are fixed, while the locations of the target and interference sources in a real-world scenario are dynamically changing, interference will leak in if the interference source is very close to the target source because the beamwidth is not narrow enough (due to power consumption, not many VMs can be used). Furthermore, this VM estimation technique does not consider changes in device orientation. When the device orientation changes, the relative positions of the RMs may also change (for example, after folding a foldable phone, each RM will be closer to each other). Correspondingly, the microphone array composed of the RMs and the fixed-generation VMs will also change, leading to changes in beamwidth and direction, making it difficult to accurately extract the target sound. Additionally, this VM estimation technique does not consider the variation in beamwidth at different frequencies under the same array conditions. For example, high-frequency beams are typically narrower, while low-frequency beams are wider. Since the interference source may be located within the low-frequency beam region, the low-frequency components may contain more interference.

[0081] To this end, this application proposes a method executed by an electronic device that can adaptively virtualize different numbers and positions of microphones for different frequencies at different times based on changes in the scene, changes in interference sources, and changes in the posture of the electronic device at different times. This results in a narrower beam, which can reduce or eliminate sound interference outside the target area while extracting sound within the target area (i.e., region of interest), thereby improving the sound extraction performance (such as accuracy) within the target area.

[0082] The following description of several optional embodiments illustrates the technical solutions of this disclosure and the technical effects produced by these solutions. It should be noted that the following embodiments can be referenced, learned from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.

[0083] Figure 4 This is a flowchart illustrating a method performed by an electronic device according to an exemplary embodiment of this application. Figure 5 This is a schematic diagram illustrating a method performed by an electronic device according to an exemplary embodiment of this application.

[0084] like Figure 4 As shown, in step S410, the first audio signal is obtained. In this application, the first audio signal may be a mixed audio signal picked up from the environment by the RM of the electronic device, which is a time-domain signal.

[0085] In step S420, based on the direction of arrival (DOA) information of the first audio signal, the target direction information, and the spatial information of the RM of the electronic device, the number and location of VMs are obtained. This operation can be performed by... Figure 5 The VM analysis and feature generation module executes this process, and the process of obtaining the number and location of VMs is performed for each frequency unit within each time unit. A frequency unit can be a frequency point or a sub-band, where a sub-band includes multiple frequency points. Furthermore, a time unit is one or more frames. The following will refer to... Figure 6 and Figure 7A This will be described in detail.

[0086] Figure 6 This is a flowchart illustrating the process of obtaining the number and location of VMs for each frequency unit under each time unit according to an exemplary embodiment of this application. Figure 7A This is a schematic diagram illustrating the structure of a VM analysis and feature generation module according to an exemplary embodiment of this application.

[0087] In step S610, based on the DOA information and target direction information, the first signal spatial distribution feature C is obtained. spa_1st In this application, the target direction information can be the direction of user interest (DOI) information, such as the direction specified by the user through operating the screen of an electronic device (e.g., zooming the screen, selecting a specific area on the screen), the direction the camera of the electronic device is pointing, etc. Furthermore, in this application, the signal spatial distribution can represent the specific direction of the target audio signal or interference signal in space, such as... Figure 7B As shown, it can determine the direction and width of the array beam. By arranging more VMs in the direction of the target audio signal or interference signal, a narrower beam can be obtained.

[0088] Specifically, firstly, the third feature vector is obtained by performing feature mapping on the DOA information. This operation can be performed by... Figure 7A The feature mapping module in the application performs this operation. For example, in this application, the third feature vector can be obtained by normalizing the DOA information and mapping the normalized DOA information to a feature distribution in the feature space. In an exemplary embodiment, this application may employ a Multiple Signal Classification (MUSIC) algorithm to obtain the DOA information, such as... Figure 8 As shown in the image. After obtaining DOA information, it can be accessed through... Figure 7AThe feature mapping module in the code normalizes the DOA information and maps the normalized DOA information to a feature distribution in the feature space, thereby obtaining the third feature vector (also known as the DOA mapping space feature), such as... Figure 9 As shown in the figure. This feature space represents the signal strength and probability of presence in each direction within a space ranging from 0 degrees to 360 degrees.

[0089] Then, the fourth feature vector is obtained by encoding the target direction information. In this application, when a user zooms in or out of the screen of an electronic device or selects a region of interest on the screen, target direction information can be obtained from the camera module of the electronic device. In this application, this can also be referred to as DOI information (or DOI parameters), that is, obtaining the direction (i.e., azimuth and elevation angles) and the region width (i.e., the width of the region of interest). After obtaining the target direction information, it can be further processed through... Figure 7A The encoding module in the module encodes the target orientation information to obtain the fourth feature vector (which can also be called the encoded vector for obtaining the target orientation information).

[0090] Subsequently, based on the third and fourth eigenvectors, the spatial distribution characteristics C of the second signal are obtained. spa_2nd For example, such as Figure 7A As shown, the spatial distribution feature C of the second signal can be obtained by multiplying the third feature vector obtained by the feature mapping module and the fourth feature vector obtained by the encoding module. spa_2nd (This can be simply referred to as the second space feature), such as Figure 10 As shown in the diagram. This process can also be described as feature labeling of the DOA mapping space features using the encoded vector of the target orientation information, such as... Figure 9 As shown in the figure. The spatial distribution characteristics C of the second signal. spa_2nd The spatial distribution of each audio signal is shown.

[0091] Then, based on the historical statistical information of the first audio signal and the spatial distribution characteristics of the second signal, the spatial distribution characteristics of the first signal are obtained. For example... Figure 11 As shown, for two consecutive audio frames (e.g., frame n and frame n+1), the extracted signals of frame n+1 and frame n have a high correlation, and the extracted signal of frame n has already been distinguished from the signal outside the region of interest by the algorithm of this application. Therefore, it can be achieved through... Figure 7A The update module in the middle uses historical statistical information to analyze the spatial distribution characteristics C of the second signal. spa_2nd The feature is updated to obtain a more accurate spatial distribution characteristic C of the first signal. spa_1st ,like Figure 9As shown in the diagram. Specifically, the historical statistical information includes the speech signal feature information and noise signal feature information of the last frame in the previous time unit. In this case, the step of obtaining the spatial distribution features of the first signal based on the historical statistical information of the first audio signal and the spatial distribution features of the second signal may include: obtaining the spatial distribution features of the first signal based on the speech signal feature information of the last frame in the previous time unit and the spatial distribution features of the second signal C. spa_2nd Calculate the correlation of each RM with respect to the audio signal within the target image region, based on the noise signal feature information of the last frame in the previous time unit and the second signal spatial distribution feature C. spa_2nd The correlation of each RM with respect to the audio signal outside the target image region is calculated to obtain the first signal spatial distribution feature C. spa_1st .like Figure 10 As shown in the figure, the spatial distribution characteristics C of the first signal spa_1st It can more accurately show the spatial distribution of each audio signal.

[0092] In one exemplary embodiment of this application, the speech signal feature information may be the speech signal covariance matrix of the last frame in the previous time unit, and the noise signal feature information may be the noise signal covariance matrix of the last frame in the previous time unit. Specifically, the extracted features (such as...) from the last frame of the previous time unit can be used... Figure 7C As shown in the diagram, speech signal feature information is obtained by calculating the speech signal covariance matrix. Noise signal feature information can also be obtained based on the extracted features of the last frame of the previous time unit; that is, by calculating the noise signal covariance matrix, and then using the speech signal covariance matrix and the second signal spatial distribution feature C... spa_2nd Calculate the correlation of each RM with respect to the audio signal within the target image region, based on the noise signal covariance matrix of the last frame in the previous time unit and the second signal spatial distribution feature C. spa_2nd The correlation of each RM with respect to the audio signal outside the target image region is calculated to obtain the first signal spatial distribution feature C. spa_1st However, the covariance matrix of the speech signal and the covariance matrix of the noise signal are only examples of the characteristic information of the speech signal and the characteristic information of the noise signal, and this application is not limited thereto.

[0093] In step S620, the number and location of VMs are determined based on the first signal spatial distribution characteristics and the spatial information of the RM. Specifically, step S620 may include: determining the number and location of VMs based on the first signal spatial distribution characteristics, the spatial information of the RM, and the state information of the electronic device. More specifically, the step of determining the number and location of VMs based on the first signal spatial distribution characteristics, the spatial information of the RM, and the state information of the electronic device may include: encoding the state information of the electronic device into a state vector; determining the microphone aperture according to the spatial information of the RM; and determining the number and location of VMs based on the first signal spatial distribution characteristics, the microphone aperture, and the state vector.

[0094] Specifically, the operations described above, which determine the number and location of VMs based on the spatial distribution characteristics of the first signal, the microphone aperture, and the state vector, can be implemented using a first neural network. In this application, the first neural network (also referred to as the VM analysis module) can be trained according to the following criteria to estimate the number of VMs:

[0095] (1) To prevent the first neural network from virtualizing the number of VMs to the maximum number each time, this application encodes the state information of the electronic device into a state vector, thereby limiting the number of VMs and reducing power consumption. In an exemplary embodiment of this application, the state information of the electronic device may include the remaining power and / or remaining computing resources of the electronic device.

[0096] (2) In the field of acoustics, the microphone aperture, as an electroacoustic transducer that converts acoustic signals into electrical signals, has its beam shape affected by its length (i.e., the maximum distance between receivers and transmitters) and the number of spatial sampling points (the number of receivers and transmitters). In other words, the length of the aperture, to a certain extent, determines the number of receivers and transmitters that can be deployed. Generally, the longer the microphone aperture, the more receivers and transmitters can be deployed, and the narrower the beam will be; conversely, the shorter the microphone aperture, the fewer receivers and transmitters that can be deployed, and the narrower the beam will be. Figure 12A As shown. For typical electronic devices, the distance between the RMs usually remains constant; however, for foldable electronic devices, the distance between the RMs changes depending on the folding state of the device. For example... Figure 12B As shown, for foldable electronic devices, when the electronic device changes from an unfolded state to a folded state, the change in the device's posture causes the relative distance between the microphone arrays (RMs) to decrease, which in turn reduces the microphone aperture. In this case, the number of RMs would theoretically decrease. Therefore, this application can obtain the posture of the electronic device, determine the spatial information of the RMs (e.g., the relative positions of the RMs) based on the posture of the electronic device, and then calculate the microphone aperture based on the spatial information of the RMs.

[0097] (3) For different frequencies, under the same microphone array conditions, the high-frequency beam is usually narrower and the low-frequency beam is usually wider. Therefore, the first neural network of this application is trained to allocate more VMs for low frequencies and fewer VMs for high frequencies. This is because a wider beam makes it easier to extract more audio signals outside the region of interest (which can be regarded as interference signals).

[0098] Furthermore, the first neural network can be trained to estimate the location of the VM based on the following criteria: calculating the direction and width of the beam formed by the microphone array based on the spatial distribution characteristics of the first signal (e.g., Figure 13 As shown in the diagram, the beam is used to cover the target audio signal within the region of interest; then, the location of the VMs is estimated based on the number of VMs, the direction of the beam, and the width.

[0099] In summary, during the inference phase, after obtaining the spatial distribution characteristics of the first signal, the microphone aperture, and the state vector, this information can be used to... Figure 7A The first neural network, trained in the model, predicts the number and location of virtual machines (VMs) for different frequency units at different time units (i.e., different moments). Since the number and location of VMs are predicted for each frequency unit within each time unit, the number and location of VMs determined for each frequency unit within each time unit constitute a three-dimensional VM position feature vector VM in the time, frequency, and spatial directions. pos In this application, the term "time unit" as used herein can refer to a "time interval" having a preset time length T. For example, a time unit can be the time of one frame, such as 16ms. Furthermore, to further reduce the complexity of the related neural network used in this application, the time unit can also be several milliseconds to several seconds; in other words, a time unit can be one or more frames. Moreover, within the same time unit, the number and position of VMs are not updated and remain unchanged. Additionally, the term "frequency unit" as used herein can refer to, for example, a frequency point or sub-band, and for each sub-band, the number and position of VMs are not updated and remain unchanged, as will be discussed later. Figure 14 A more detailed description of frequency points and sub-bands is provided. In a simplified example, such as... Figure 7C As shown, the VM analysis module can be run only once at the beginning of each time unit to obtain the VM location feature vector VM for each frequency unit under that time unit. pos In other words, the number and location of VMs may differ in different frequency units within the same time unit, but remain constant.

[0100] Return to reference Figure 4 In step S430, a second audio signal is obtained based on the first audio signal and the number and position of the VMs. (See below for further details.) Figure 14 This will be described in detail.

[0101] like Figure 14 As shown, in step S1410, a first feature vector corresponding to the first audio signal is obtained.

[0102] Specifically, the step of obtaining the first feature vector corresponding to the input first audio signal may include: extracting the feature vector from the first audio signal; and obtaining the first feature vector by performing feature encoding on the extracted feature vector.

[0103] First, feature extraction is performed on the first audio signal to obtain a feature vector in another dimension. For example, Short-Time Fourier Transform (STFT) (e.g., 512-point STFT) can be used to perform feature extraction, that is, the first audio signal is framed, windowed, and subjected to STFT to obtain features in the frequency domain.

[0104] For example, for a first audio signal with a sampling rate of 16k and a duration of n seconds, there are L = n * 16000 sampling points. After performing an STFT with a window length of W = s_n sampling points (i.e., the number of sampling points per frame is s_n, and the overlap area between frames is s_n / 2 (i.e., 50% overlap), which is also a frame shift of W / 2), the number of frames k is k = L / (s_n / 2) - 1, and the number of frequency points per frame is f = s_n / 2. By extracting the real and imaginary parts of the frequency domain, a feature vector with dimension [k, f] can be obtained. For example, for a first audio signal with a sampling rate of 16k and a duration of 4s, after performing an STFT with a window length of W = 512 sampling points (i.e., a frame shift of 256 sampling points), the number of frames is 249, and the number of frequency points f per frame is s_n / 2 = 512 / 2 = 256. Each frequency point is represented by a real part and an imaginary part. Therefore, a feature vector with a dimension of [249, 256] can be obtained.

[0105] The above example uses STFT to perform feature extraction, but this application is not limited to this. Other feature extraction methods can also be used, such as using Convolutional Neural Networks (CNN) for feature extraction.

[0106] After feature extraction to obtain feature vectors, the first feature vector is obtained by feature encoding the extracted feature vectors.

[0107] In one exemplary embodiment of this application, the extracted feature vector can be directly encoded using an encoder to obtain the first feature vector. Given that this application requires the virtualization of different microphone features for different frequencies, in another exemplary embodiment of this application, the frequency is divided into multiple sub-bands, and frequency points within the same sub-band use the same type of VM features (i.e., the same number and location of VMs) at the same time, thereby reducing computational complexity. The following refers to... Figure 15 and Figure 16 This will be described in detail. If the frequency is not divided into multiple sub-bands and a full-band processing method is used, the VM analysis and feature generation modules described later will virtualize the same number and location of VMs for different frequencies (i.e., frequency points).

[0108] Figure 15 This is a block diagram illustrating an encoder module according to an exemplary embodiment of this application. Figure 16 This is a network flowchart illustrating an encoder module according to an exemplary embodiment of this application.

[0109] like Figure 15 As shown, the encoder module includes a feature extraction module, a sub-feature partitioning module, and multiple sub-encoders. The feature extraction module is used to perform the operation of extracting feature vectors from the first audio signal as described above. Since this has been described in detail above, it will not be repeated here.

[0110] like Figure 16 As shown, the sub-feature segmentation module performs frequency band segmentation to divide the extracted feature vector F into multiple first sub-band feature vectors. For example, the 16kHz frequency band is divided into N sub-bands. Considering performance and model complexity, N can be equal to 4, 5, or 6 in one example; however, this application is not limited to this. For example, as shown... Figure 16 As shown, the 16kHz frequency band can be divided into 4 sub-bands. For the frequency domain data obtained in the previous step, the 256 frequency points of each frame (i.e., the extracted feature vectors) can be divided into 4 first sub-band feature vectors f1, f2, f3 and f4. The frequency points contained in each first sub-band feature vector are {1~32}, {33~64}, {65~128} and {129~256}, respectively, and the corresponding frequencies are 0~2kHz, 2kHz~4kHz, 4kHz-8kHz and 8kHz~16kHz, respectively.

[0111] After dividing the extracted feature vector F into multiple first sub-band feature vectors, each of the multiple first sub-band feature vectors is encoded using a corresponding sub-band encoder to obtain multiple second sub-band feature vectors, wherein the first feature vectors include the multiple second sub-band feature vectors. Figure 15 and Figure 16As shown, the number of first sub-band feature vectors, N, is 4. Correspondingly, there are N = 4 sub-encoders. Each first sub-band feature vector is input into a corresponding sub-encoder (e.g., a 2D convolutional neural network (2D-CNN)) for encoding, obtaining four second sub-band feature vectors x1, x2, x3, and x4. These multiple second sub-band feature vectors can be collectively referred to as the first feature vectors. This application achieves parallel encoding, reduces complexity, and improves the model's processing speed by dividing the extracted feature vectors into multiple first sub-band feature vectors and using different sub-encoders to encode the corresponding first sub-band feature vectors.

[0112] Return to reference Figure 14 In step S1420, a second audio signal is obtained based on the first feature vector, the VM position feature vector, and the spatial information of the RM.

[0113] The VM position feature vector determined in step S620 (i.e., the number and position of VMs determined for each frequency unit in each time unit) does not consider the directivity of the RM. The directivity of the RM indicates that the microphone picks up sound from different directions, which determines the microphone's sensitivity and signal reception capability at different angles. Microphone directivity can be divided into omnidirectional, cardioid, wide cardioid, and supercardioid. An omnidirectional microphone can pick up audio signals from all directions without distinction; a cardioid microphone can mainly pick up audio signals from the front; a wide cardioid microphone can pick up audio signals from all directions similar to an omnidirectional microphone, but its ability to pick up audio signals from the rear is slightly weaker than that of an omnidirectional microphone; a supercardioid microphone can mainly pick up audio signals from both the front and rear, but its ability to pick up audio signals from the rear is slightly weaker than that of the front. Electronic devices such as mobile phones typically include three microphone arrays (RMs) with different directional properties. The top and bottom RMs are usually omnidirectional or widecardioid, while the rear RM is usually cardioid or supercardioid. Therefore, to enable the microphone array to pick up more signals within the region of interest (ROI) and strong interference signals outside the RIO, this application adds directional properties to the determined VM based on the relative position between the microphone array (VM) and the RMs. This enhances the distinction between signals within the RIO and interference signals outside the RIO, thereby improving signal extraction within the RIO.

[0114] Specifically, step S1420 may include: determining the relative position information between the VM and the RM based on the VM position feature vector and the spatial information of the RM; determining the directional of each VM based on the relative position information between the VM and the RM and the directional of the RM; obtaining a second feature vector based on the relative position information between the VM and the RM, the first feature vector and the directional of each VM; and obtaining a second audio signal based on the first feature vector and the second feature vector.

[0115] In one exemplary embodiment of this application, such as Figure 7A As shown, the relative positional information between each VM and each RM can be determined by the first sub-network (i.e., the directivity processing module) in the second neural network (i.e., the VM feature generation module) based on the VM positional feature vector and the spatial information of the RM. Subsequently, the first sub-network can determine the relative positional information between each VM and each RM, as well as the directivity P of the RM closest to each VM, based on the relative positional information between each VM and each RM and the directivity P of the RM closest to each VM. RM The directionality of each VM is determined, specifically by the first subnetwork based on the relative position information between each VM and each RM, and the directionality P of the RM nearest to each VM. RM To predetermine the directivity P of each VM VM And based on the signal direction, the directivity P of each VM is determined. VM The directionality of each VM is determined by adjusting its orientation to point more towards signals or interference within the region of interest. Figure 17 As shown in the image.

[0116] After determining the target of each VM, such as Figure 7A As shown, based on the relative position information between the VM and the RM, the first feature vector, and the directional characteristics of each VM, the second feature vector is obtained through the second sub-network (i.e., the amplitude and phase processing module) in the second neural network. Specifically, the second sub-network can calculate the phase difference and amplitude attenuation between the RM and the VM based on the relative position information between the VM and the RM and the first feature vector, and then generate the second feature vector based on the directional characteristics of each VM and the phase difference and amplitude attenuation between the RM and the VM.

[0117] After generating the second feature vector, the second audio signal can be obtained based on the first and second feature vectors. This will be described in detail below.

[0118] First, based on the first and second feature vectors, a mask related to the second audio signal within the target image region is obtained. This operation can be performed by... Figure 5 The sound extraction module in the program is executed.

[0119] Specifically, because the number and location of VMs typically vary across different time and frequency units, the second feature vector is discontinuous in the temporal, spatial (also known as microphone) or frequency directions. This makes it difficult for neural networks to transmit hidden states when processing these features. For example, as... Figure 18 As shown, for RM1, the hidden state H 1_rm Initialized to 0, the hidden state H is used by the neural network. 1_rm After processing the features of the frequency unit f=1, the hidden state H 1_rm The hidden state H is updated; thereafter, the neural network can utilize the updated hidden state H. 1_rm Process the next feature (i.e., the feature at frequency unit f = 2), thus completing the transmission of feature information from frequency unit f = 1 to frequency unit f = 2. However, as Figure 18 As shown, for VMs, neural networks have difficulty transmitting hidden states when processing features. For example, for VM2, the hidden state H... 2_vm Initialized to 0, the hidden state H is used by the neural network. 2_vm After processing the features of the frequency unit f=1, the hidden state H 2_vm The hidden state H is updated; however, VM2 does not exist at frequency unit f=2, and therefore no feature data associated with VM2 exists. 2_vm It cannot be transmitted smoothly. Subsequently, VM2 exists at frequency unit f=4, and corresponding feature data related to VM2 exists. If the previous hidden state H is used directly... 2_vm This prevents the neural network from learning feature information at frequency units f=2 and f=3; similarly, for VM1, since VM1 does not exist at frequency unit f=1, there is no feature data associated with VM1, therefore no new hidden state is generated, and the initial hidden state H remains unchanged. 1_vm If f=0, it will not be updated, which causes the loss of feature information from the frequency unit f=1 during processing, thus affecting the performance of feature processing.

[0120] Since the VM position feature vector is used when generating the second feature vector, the second feature vector contains spatial information. This application uses the following two operations to process the correlation between the VM position feature vector of the current step (e.g., the current frequency unit) and the hidden state of the previous step (the previous frequency unit) to achieve normal transmission of the hidden state during network feature processing: (1) After processing the features corresponding to the previous step in the feature vector corresponding to each VM, the global hidden state is obtained by performing feature processing on the hidden state of all VMs. The fusion weight of each hidden state can be learned by the network or by weighted summation. For example, the closer a VM exists in the previous step is to a VM existing in the current step, the greater the weight of the VM existing in the previous step. However, this application is not limited to this. (2) When processing the features corresponding to the current step in the feature vector corresponding to each VM, for a VM, the hidden state of the VM in the current step is updated by calculating the correlation between the position feature of the VM in the VM position feature vector and the spatial feature contained in the global hidden state.

[0121] Based on this principle, this application processes the feature vectors in the time direction, spatial direction (also known as microphone direction), and frequency direction, thereby enabling the normal transmission of hidden states during network feature processing and improving the accuracy of the final prediction results.

[0122] In an exemplary embodiment of this application, the step of obtaining a mask related to a second audio signal within a target image region based on a first feature vector and a second feature vector may include: merging the first feature vector X and the second feature vector X_vm along a spatial direction to obtain a merged feature vector; performing feature processing sequentially in a first direction, a second direction, and a third direction on the merged feature vector to obtain a mask, wherein the first direction, the second direction, and the third direction are one of a spatial direction, a temporal direction, and a frequency direction, respectively; furthermore, the first direction, the second direction, and the third direction are all different from each other, that is, when performing feature processing in the first direction, the input feature vector is the merged feature vector; when performing feature processing in the second direction, the input feature vector is the feature vector obtained after performing feature processing in the first direction; when performing feature processing in the third direction, the input feature vector is the feature vector obtained after performing feature processing in the second direction, and the final feature vector obtained after performing feature processing in the third direction can be represented as a mask. In other words, feature processing can be performed in any of the following orders: (1) spatial direction, temporal direction, and frequency direction; (2) spatial direction, frequency direction, and temporal direction; (3) temporal direction, spatial direction, and frequency direction; (4) temporal direction, frequency direction, and spatial direction; (5) frequency direction, spatial direction, and temporal direction; (6) frequency direction, temporal direction, and spatial direction. See below for further details. Figures 19 to 22 The process of performing feature processing in the order of frequency, time, and space is described in detail.

[0123] Figure 19 This is a system block diagram illustrating a sound extraction module according to an exemplary embodiment of this application. Figure 20 This is a schematic diagram illustrating the process of performing feature processing in the frequency direction for each time unit according to an exemplary embodiment of this application.

[0124] In this application, a frequency unit can be a frequency point or a sub-band. When a frequency unit is a sub-band, the sub-band can include multiple frequency points. The following description uses a frequency unit as a sub-band as an example. In this application, the sound extraction module can also be called a third neural network, including a first feature extraction module, a second feature extraction module, a third feature extraction module, a first fusion module, a second fusion module, a third fusion module, a first update module, a second update module, and a third update module. The first feature processing module to the third feature processing module can adopt a recurrent neural network (RNN), but this application is not limited to this. Other networks with temporal processing capabilities (such as CNN, attention network, long short-term memory network (LSTM), etc.) can also be used. Similarly, the first fusion module to the third fusion module and the first update module to the third update module can adopt CNN, RNN, etc.

[0125] exist Figure 19 and Figure 20 middle, express Figure 19 After the first feature processing module or the second feature processing module processes the features corresponding to the k-th subband in the feature vector corresponding to each VM, the hidden state set of the network is obtained, where N represents the number of VMs; This represents the initial hidden layer state corresponding to the i-th VM before processing the feature corresponding to the k-th frequency unit in the feature vector corresponding to the i-th VM; express Figure 19 The third feature processing module processes the features corresponding to each subband in the feature vector corresponding to the i-th VM, and then sets the hidden state of the network. This indicates that at time t, the subband f = k (which can also be written as f) k The location feature vector of the i-th VM.

[0126] like Figure 19 and Figure 20 As shown, when processing the feature vectors corresponding to the RMs, for each sub-band, based on the initial hidden state of each RM in the current sub-band, the features corresponding to the current sub-band in the input feature vectors corresponding to each RM are processed to obtain the processed features and the processed hidden state corresponding to each RM. The processed hidden state corresponding to each RM is used as the initial hidden state of each RM in the next sub-band. This processing operation can be performed by... Figure 19The first feature processing module in the third neural network is executed. Since feature processing is performed first in the frequency direction, the input feature vector corresponding to each RM is the first feature vector in the merged feature vector, where, as mentioned above, the merged feature vector is obtained by merging the first feature vector and the second feature vector along the spatial direction. Figure 19 neutralization Figure 20 The example shown is a 3D form of merged feature vectors, but this application is not limited to this; the merged feature vectors can also be 4D, 5D, or other higher-dimensional feature vectors. The process of processing the feature vectors of the RM is similar to the above-mentioned reference. Figure 18 The process described is the same, so it will not be repeated here.

[0127] When processing the feature vectors corresponding to VMs, the input feature vector corresponding to each VM is the second feature vector from the merged feature vectors. For example... Figure 19 and Figure 20 As shown, for each sub-band, the following operations are performed: based on the initial hidden state of each VM in the current sub-band, the features corresponding to the current sub-band in the input feature vector corresponding to each VM are processed to obtain the processed features and the processed hidden state corresponding to each VM; by performing feature processing on the processed hidden state corresponding to each VM, the global hidden state is obtained; by performing feature processing on the feature vector corresponding to the next sub-band in the VM position feature vector and the global hidden state, the initial hidden state of each VM in the next sub-band is obtained.

[0128] like Figure 20 As shown, for the current subband f=1 (i.e., f1), the first feature processing module in the third neural network is based on the initial hidden state of VM2 in the current subband f=1. The features corresponding to the current subband f=1 in the input feature vector corresponding to VM2 are processed to obtain the processed features corresponding to the current subband f=1 and the processed hidden state corresponding to VM2. The features corresponding to the current subband f=1 in the input feature vector corresponding to each VM are processed in a similar manner (if there is no feature corresponding to the current subband f=1 in the feature vector corresponding to a VM, no processing is performed), thereby obtaining the processed features and the processed hidden state corresponding to each VM. (Where i represents the VM number). Then, the processed hidden states corresponding to each VM are processed using the first fusion module in the third neural network. Perform feature processing to obtain the global hidden state H global_vmFor example, the feature processing performed by the first fusion module can be a convolutional operation performed by a CNN, but this application is not limited to this; it can also be feature processing performed by a network with an attention mechanism. Subsequently, the feature vector corresponding to the next subband f=2 (i.e., f2) in the VM position feature vector is updated using the first update module in the third neural network. and the global hidden state H global_vm Feature processing is performed to obtain the initial hidden state of each VM in the next subband f=2. exist Figure 20 Since only VM1 and VM4 exist in the next subband f=2, the initial hidden states of VM1 and VM4 in the next subband f=2 can be obtained. and Then, after performing feature processing on subband f=1, feature processing on subband f=2 can be performed in a similar manner. For example, the first feature processing module can be based on the initial hidden state. and The features corresponding to the current subband f=2 in the input feature vectors corresponding to VM1 and VM4 are processed to obtain the processed features corresponding to the current subband f=2 and the processed hidden state corresponding to VM1. and the hidden state corresponding to VM4 By processing the feature vectors corresponding to each VM in the above manner, the normal propagation of hidden states can be achieved during network feature processing.

[0129] By reference Figure 20 The process described above yields the feature vector corresponding to each VM after processing in the frequency direction, and then, as... Figure 19 As shown, the obtained feature vectors processed in the frequency direction for each VM are input into the second feature processing module, and then feature processing is performed in the time direction. See below for reference. Figure 21 and Figure 19 This will be described.

[0130] Figure 21 This is a schematic diagram illustrating the process of performing feature processing in the time direction for each frequency unit according to an exemplary embodiment of this application.

[0131] like Figure 19 and Figure 21As shown, when processing the feature vectors corresponding to the RM, for each frame (e.g., a frame is 16ms), based on the initial hidden state of each RM in the current frame, the features corresponding to the current frame in the input feature vectors corresponding to each RM are processed to obtain the processed features and the processed hidden state corresponding to each RM. The processed hidden state corresponding to each RM is used as the initial hidden state of each RM in the next frame.

[0132] For example, such as Figure 21 As shown, for RM1, the second feature processing module in the third neural network is based on the initial hidden state H in the first frame. 1_rm The features corresponding to the first frame in the input feature vector corresponding to RM1 are processed to obtain the processed features and the updated hidden state H. 1_rm The updated hidden state is used as the initial hidden state of RM1 in the next frame.

[0133] When processing the feature vectors corresponding to the VM, such as Figure 19 and Figure 21 As shown, for each frame, the following operations are performed: based on the initial hidden state of each VM in the current frame, the features corresponding to the current frame in the input feature vector corresponding to each VM are processed to obtain the processed features and the processed hidden state corresponding to each VM; by performing feature processing on the processed hidden state corresponding to each VM, the global hidden state is obtained; by performing feature processing on the feature vector corresponding to the next frame in the VM position feature vector and the global hidden state, the initial hidden state of each VM in the next frame is obtained.

[0134] like Figure 21 As shown, for the first frame, the second feature processing module is based on the initial hidden state of VM2 in the first frame. The features corresponding to the first frame in the input feature vector corresponding to VM2 are processed to obtain the processed features and the processed hidden state corresponding to VM2. The features corresponding to the first frame in the input feature vector corresponding to each VM are processed in a similar manner (if the feature vector corresponding to a VM does not contain a feature corresponding to the first frame, then no processing is performed), thereby obtaining the processed features and the processed hidden state corresponding to each VM. (where i represents the VM number), for example, such as Figure 21 As shown, since only the feature vectors corresponding to VM2, VM5, and VM6 contain features corresponding to the first frame, the hidden state obtained in the current step is only the hidden state corresponding to VM2. Hidden state corresponding to VM5 and the hidden state corresponding to VM6 Then, the processed hidden state is processed by utilizing the second fusion module in the third neural network. and Perform feature processing to obtain the global hidden state H global_vm For example, the feature processing performed by the second fusion module can be a convolutional operation performed by a CNN, but this application is not limited to this; it can also be feature processing performed by a network with an attention mechanism. Subsequently, the feature vector corresponding to the next frame (i.e., the 2nd frame) in the VM position feature vector is updated using the second update module in the third neural network. and the global hidden state H global_vm Feature processing is performed to obtain the initial hidden state of each VM in frame 2. exist Figure 21 Since only VM1 and VM4 exist in the second frame, the initial hidden state of VM1 and VM4 in the second frame can be obtained through the update operation performed by the second update module. and Then, after performing feature processing on the first frame, feature processing can be performed on the second frame in a similar manner. For example, the second feature processing module can be based on the initial hidden state. and The features corresponding to the second frame in the input feature vectors corresponding to VM1 and VM4 are processed to obtain the processed features corresponding to the second frame and the processed hidden state corresponding to VM1. and the hidden state corresponding to VM4 By processing the feature vectors corresponding to each VM in the above manner, the normal propagation of hidden states can be achieved during network feature processing.

[0135] By reference Figure 21 The above process describes how to obtain the feature vector corresponding to each VM after processing in the time direction, and then, as... Figure 19 As shown, the obtained feature vectors, processed in the temporal direction and corresponding to each VM, are input into the third feature processing module, and then feature processing is performed in the spatial direction. See below for reference. Figure 22 and Figure 19 This will be described.

[0136] Figure 22 This is a schematic diagram illustrating the process of performing feature processing in the spatial direction for each time unit according to an exemplary embodiment of this application.

[0137] like Figure 19 and Figure 22 As shown, when processing the feature vector corresponding to the RM, for each RM, the following operations are performed: based on the initial hidden state of the current RM under each sub-band, the features corresponding to each sub-band in the input feature vector corresponding to the current RM are processed to obtain the processed features and the processed hidden state corresponding to each sub-band, wherein the processed hidden state corresponding to each sub-band is used as the initial hidden state of each sub-band under the next RM.

[0138] For example, such as Figure 22 As shown, for RM1, the third feature processing module in the third neural network processes the features corresponding to each sub-band in the input feature vector corresponding to RM1 based on the initial hidden state of RM1 under each sub-band, obtaining the processed features and the processed hidden state corresponding to each sub-band. The processed hidden state corresponding to each sub-band is used as the initial hidden state of each sub-band under RM2. Feature processing is performed for RM2 and RM3 in a similar manner.

[0139] When processing the feature vectors corresponding to the VM, such as Figure 19 and Figure 22 As shown, for each VM, the following operations are performed: based on the initial hidden state of the current VM in each sub-band, the features corresponding to each sub-band in the input feature vector corresponding to the current VM are processed to obtain the processed features and the processed hidden state corresponding to each sub-band. By performing feature processing on the processed hidden state corresponding to each sub-band, the global hidden state is obtained. And by performing feature processing on the feature vector corresponding to the next VM in the VM position feature vector and the global hidden state, the initial hidden state of the next VM in each sub-band is obtained.

[0140] like Figure 22 As shown, for VM1, the third feature processing module is based on the initial hidden state of VM1 under subband f=2 and subband f=3. and The features corresponding to subbands f=2 and f=3 in the input feature vector corresponding to VM1 are processed to obtain the processed features and the processed hidden states corresponding to subbands f=2 and f=3. and (If the feature vector corresponding to a certain VM1 does not contain a feature corresponding to a certain subband, then no processing is performed on that subband.) Then, the processed hidden state is processed using the third fusion module in the third neural network. and Perform feature processing to obtain the global hidden state H global_vmFor example, the feature processing performed by the third fusion module can be a convolutional operation performed by a CNN, but this application is not limited to this; it can also be feature processing performed by a network with an attention mechanism. Subsequently, the feature vector corresponding to the next VM (i.e., VM2) in the VM location feature vector is updated using the third update module in the third neural network. and the global hidden state H global_vm Feature processing is performed to obtain the initial hidden state of VM2 in each subband. exist Figure 22 Since at time t, the feature vector corresponding to VM2 only contains features corresponding to subband f=1, subband f=4, and subband f=5, the initial hidden states of VM2 under subband f=1, subband f=4, and subband f=5 can be obtained through the update operation performed by the third update module. and Then, after performing feature processing on VM1, feature processing on VM2 can be performed in a similar manner. For example, the third feature processing module can be based on the initial hidden state. and The features corresponding to subband f=1, subband f=4, and subband f=5 in the input feature vector corresponding to VM2 are processed to obtain the processed features corresponding to VM2 and the processed hidden state corresponding to subband f=1. Hidden state corresponding to subband f=4 and the hidden state corresponding to subband f=5 By processing the feature vectors corresponding to each VM in the above manner, the normal propagation of hidden states can be achieved during network feature processing.

[0141] Finally, by performing feature processing in the frequency, time, and spatial directions as described above, a mask associated with the second audio signal within the target image region can be obtained, wherein the mask can represent the percentage of information occupied by the second audio signal (i.e., the target audio signal) in the first audio signal.

[0142] After obtaining the mask, the second audio signal is obtained based on the mask and the first feature vector. This operation can be performed by... Figure 5 The decoder module in the process executes the steps. Specifically, the step of obtaining the second audio signal based on the mask and the first feature vector includes: obtaining a fifth feature vector based on the mask and the first feature vector; and obtaining the second audio signal by performing an audio signal recovery operation on the fifth feature vector.

[0143] In one exemplary embodiment of this application, if Figure 5The encoder module directly encodes the extracted feature vector to obtain the first feature vector. In the decoder module, the mask can be directly multiplied by the first feature vector to obtain the fifth feature vector. Then, the obtained fifth feature vector is decoded to recover the target audio signal, i.e., the second audio signal.

[0144] In another exemplary embodiment of this application, if Figure 5 The encoder module in the middle adopts Figure 15 The structure shown divides the first feature vector into multiple first sub-band feature vectors. By encoding these multiple first sub-band feature vectors using multiple sub-band encoders, the first feature vector is obtained. In the decoder module, multiple sub-decoders can be used accordingly to perform decoding. See below for further details. Figure 23 and Figure 24 This will be described in detail.

[0145] Figure 23 This is a block diagram illustrating a decoder module according to an exemplary embodiment of this application. Figure 24 This is a network flowchart illustrating a decoder module according to an exemplary embodiment of this application.

[0146] like Figure 23 As shown, the decoder module includes multiple sub-decoders, a feature merging module, and a time-domain signal recovery module.

[0147] exist Figure 24 In the example shown, the obtained first feature vector includes four second sub-band feature vectors x1, x2, x3, and x4. In this case, the i-th sub-band mask M in the mask is processed by the i-th sub-decoder. i and the i-th second sub-band feature vector x in the first feature vector i The process involves processing (e.g., performing element-wise multiplication) to obtain the i-th third sub-band feature vector. Then, feature transformation is performed on the four obtained third sub-feature vectors to obtain four predicted features y1, y2, y3, and y4. This feature transformation operation can be implemented using a linear fully connected layer; however, this application is not limited to this and can also be implemented using a network capable of feature transformation (e.g., CNN). Subsequently, a feature merging module performs feature merging (which can also be understood as band merging) on ​​the four predicted features y1, y2, y3, and y4 to obtain the merged feature y. k This is the fifth eigenvector, used to facilitate subsequent feature transformation operations.

[0148] In other words, the step of obtaining the fifth feature vector based on the mask and the first feature vector includes: processing each sub-band mask in the mask and the corresponding second sub-band feature vector in the first feature vector to obtain multiple third sub-band feature vectors; performing feature transformation on the multiple third sub-band feature vectors to obtain multiple predicted features; and merging the multiple predicted features to obtain the fifth feature vector.

[0149] After obtaining the fifth eigenvector, it can be obtained through methods such as... Figure 23 The temporal signal recovery module performs an audio signal recovery operation on the fifth feature vector to obtain the target audio signal, i.e., the second audio signal. For example, the temporal signal recovery module can use inverse short-time Fourier transform to perform the audio signal recovery operation. However, this application is not limited to this. The temporal signal recovery module can also use other feature transformation methods, such as using a CNN network to perform the audio signal recovery operation. In this application, if the encoder module uses short-time Fourier transform for feature extraction, then in the decoder module, the temporal signal recovery module can use inverse short-time Fourier transform to realize the audio signal recovery operation on the fifth feature vector. Correspondingly, if the encoder module uses a CNN network for feature extraction, then in the decoder module, the temporal signal recovery module can use a CNN network to realize the audio signal recovery operation on the fifth feature vector.

[0150] The method described above with reference to the accompanying drawings for extracting the user-desired target audio signal (second audio signal) from the input mixed audio signal (i.e., the first audio signal) uses multiple neural networks (i.e., the first neural network, the second neural network, and the third neural network), which can be obtained through offline training. This will be described in detail below.

[0151] Figure 25 This is a schematic diagram illustrating the process of training a sound extraction module (i.e., a third neural network) according to an exemplary embodiment of this application. Simply put, the sound extraction module is trained using simulated features generated by a simulator.

[0152] First, the locations of the virtual machines (VMs) are obtained using expert guidance based on the positions of the target audio signal and the interference signal. Then, based on these VM locations, the feature vectors of the VMs are simulated using a simulator to obtain the feature vectors of the VMs. Additionally, the signals of the RMs (e.g., signals of RM1, RM2, and RM3) are encoded using an encoder module to obtain the feature vectors of the RMs. Subsequently, the obtained VM feature vectors (simulate_vm) and RM feature vectors are input into a third neural network (i.e., the sound extraction module) for feature processing to obtain a mask. After performing a dot product operation between the mask and the RM feature vectors, the result of the dot product operation is input into a decoder module for feature decoding to extract the audio signal. Based on the extracted audio signal and the target audio signal, the maximum signal-to-noise ratio (SNR) is calculated as the objective function. The third neural network is iteratively trained by adjusting its model parameters until the objective function SNR obtained through the above process converges.

[0153] After obtaining the trained third neural network through the above training process, this trained third neural network is then jointly trained with the VM analysis and feature generation module (i.e., the first and second neural networks). See below for reference. Figure 26 This will be described in detail.

[0154] Figure 26 This is a schematic diagram illustrating the process of jointly training a third neural network and a VM analysis and feature generation module (wherein the VM analysis and feature generation module includes a first neural network and a second neural network) according to an exemplary embodiment of this application.

[0155] Specifically, an encoder module encodes the signals of the RM (e.g., signals of RM1, RM2, and RM3) to obtain the feature vector of the RM (i.e., the first feature vector). Additionally, the obtained feature vector of the RM, DOA information, target direction information (i.e., DOI information), and spatial information of the RM are input into a VM analysis and feature generation module to obtain the feature vector of the VM (i.e., the second feature vector). Furthermore, the information input to the VM analysis and feature generation module also includes at least one of the following: electronic device status information, historical statistical information, RM directivity, and electronic device status information. Specifically, for each frequency unit within each time unit, the following operations are performed: based on the DOA information and target direction information, a first signal spatial distribution feature is obtained; based on the first signal spatial distribution feature and the spatial information of the RM, the number and location of VMs are determined, wherein the number and location of VMs obtained in each frequency unit within each time unit constitute the VM location feature vector. Subsequently, the relative positional information between the VM and RM is determined based on the VM's positional feature vector and the RM's spatial characteristics. The directional nature of each VM is determined based on this relative positional information and the RM's directional nature. Finally, the feature vector of each VM is obtained based on the relative positional information, the RM's feature vector, and the directional nature of each VM. Then, the VM's feature vector and the RM's feature vector are input into the third neural network (i.e., the sound extraction module) to obtain a mask. After performing a dot product operation between the mask and the RM's feature vector, the result is input into the decoder module for feature decoding to extract the audio signal. During training, the SNR is calculated as the first objective function based on the extracted audio signal and the target audio signal. The minimum mean square error (MSE) is calculated as the second objective function based on the VM's feature vector obtained through the VM analysis and feature generation module and the VM's feature vector simulate_vm obtained through simulator simulation. The neural networks are iteratively trained by adjusting the model parameters of the first and second neural networks while simultaneously fine-tuning the model parameters of the third neural network until the first and second objective functions obtained through the above process converge.

[0156] The above description, with reference to the accompanying drawings, outlines a method for extracting a user-desired target audio signal (second audio signal) from an input mixed audio signal (i.e., a first audio signal). This method can be applied to various scenarios requiring voice focusing, enabling video or audio recorders to clearly pick up sounds of interest to the user (such as distant sounds or information from a specific direction). The following examples illustrate three application scenarios of the method performed by an electronic device as described above; however, actual usage scenarios are not limited to these three.

[0157] Figure 27This is a schematic diagram illustrating the application of the method performed by an electronic device according to this application in a video recording scenario. In this example, the user can select an object of interest to record and control the range of the recorded object.

[0158] First, the user activates the audio zoom function in video mode. Then, the user selects a region of interest (ROI) (also called a target image region) on the screen by performing a touch operation (e.g., a swipe). Subsequently, a method executed by the electronic device obtains the target orientation information (i.e., the orientation and width of the ROI), also known as DOI information. Specifically, the video processing module acquires the coordinate values ​​of the user-selected ROI, then calculates the width of the ROI based on the coordinate values ​​and the zoom state, and subsequently calculates the orientation of the ROI based on the current zoom state and the angular offset value relative to the front. Based on this, the target sound (i.e., the second audio signal) within the ROI region can be extracted using the method executed by the electronic device proposed in this application. In this example, the user can accurately select the ROI and accurately extract the target sound within that region for enhancement or other processing. That is, the method executed by the electronic device described above is applicable not only to sound extraction but also to speech enhancement and speech separation.

[0159] Furthermore, in video recording scenarios, the schematic diagrams of the methods performed by electronic devices according to this application can also be applied. For example, a user can record a singer's voice at a concert by zooming in and out of the screen, or record music from a street performer in a noisy environment by zooming in and out of the screen.

[0160] First, the user activates the audio zoom function in video mode. Then, the user zooms in or out of the screen (i.e., zooms in or out of the audio pickup area). A method executed by the electronic device then obtains target orientation information (i.e., the direction and width of the region of interest), where the direction of the region of interest can be assumed to be directly in front (i.e., facing forward). Based on this, the target sound (i.e., the second audio signal) within the region of interest can be extracted using the method executed by the electronic device proposed in this application, and used as the audio signal for the user during video recording. In this example, by zooming in or out of the screen, the method executed by the electronic device proposed in this application can obtain target orientation information, thereby accurately extracting the sound within the screen while eliminating interference from outside the screen.

[0161] Furthermore, the schematic diagram of the method executed by the electronic device according to this application can also be applied in scenarios involving recording voice memos or making calls. For example, when a user makes a voice memo, the way the user holds the phone may change, and the relative position of the user's mouth to the phone may also differ. By applying the method executed by the electronic device proposed in this application, the user's voice can be accurately extracted by following the changes in the relative position between the user's mouth and the phone. This allows the user to record their voice while blocking out background noise, even in noisy environments.

[0162] Specifically, first, the user activates the voice zoom function in video mode. Then, target direction information is obtained, in this case, the direction of the user's voice relative to the phone. Subsequently, the user's voice (i.e., the second audio signal) can be extracted using the method performed by the electronic device proposed in this application, while suppressing ambient noise.

[0163] This disclosure also provides an electronic device including at least one processor, and optionally, at least one transceiver coupled to the at least one processor and / or at least one memory, wherein the at least one processor is configured to perform the steps of the method provided in any optional embodiment of this disclosure.

[0164] Figure 28 The diagram shows a structural schematic of an electronic device to which an embodiment of the present invention applies, such as... Figure 28 As shown, Figure 28 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, each of the processor 4001, memory 4003, and transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this disclosure. Optionally, the electronic device may be a first network node, a second network node, or a third network node.

[0165] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0166] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 28 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0167] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.

[0168] The memory 4003 is used to store computer programs or executable instructions that execute the embodiments of this disclosure, and is controlled by the processor 4001 to execute them. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0169] This disclosure provides a computer-readable storage medium storing a computer program or instructions that, when executed by at least one processor, can perform or implement the steps and corresponding content of the aforementioned method embodiments.

[0170] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.

[0171] The terms “first,” “second,” “third,” “fourth,” “1,” “2,” etc. (if present) in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in a sequence other than that shown in the figures or text.

[0172] It should be understood that although arrows indicate various operation steps in the flowcharts of the embodiments of this disclosure, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of the embodiments of this disclosure, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured as required, and the embodiments of this disclosure do not limit this.

[0173] The above text and accompanying drawings are provided as examples only to help the reader understand this disclosure. They are not intended and should not be construed as limiting the scope of this disclosure in any way. Although certain embodiments and examples have been provided, it will be apparent to those skilled in the art, based on the content disclosed herein, that changes can be made to the illustrated embodiments and examples, and other similar implementations based on the technical concept of this disclosure can be adopted without departing from the scope of this disclosure, and these modifications and modifications are also within the protection scope of the embodiments of this disclosure.

Claims

1. A method performed by an electronic device, comprising: Obtain the first input audio signal; Based on the direction of arrival (DOA) information of the first audio signal, the target direction information, and the spatial information of the physical microphone (RM) of the electronic device, the number and location of virtual microphones (VMs) are obtained. A second audio signal is obtained based on the first audio signal and the number and position of VMs.

2. The method according to claim 1, wherein, The steps for obtaining the number and location of virtual microphones (VMs) based on the direction of arrival (DOA) information of the first audio signal, the target direction information, and the spatial information of the physical microphone (RM) of the electronic device include: performing the following operations for each frequency unit within each time unit: Based on the DOA information and the target direction information, the spatial distribution characteristics of the first signal are obtained; Based on the spatial distribution characteristics of the first signal and the spatial information of the RM, the number and location of VMs are determined.

3. The method according to claim 2, wherein, The steps for obtaining the spatial distribution characteristics of the first signal based on the DOA information and the target direction information include: A third feature vector is obtained by feature mapping the DOA information; The fourth feature vector is obtained by encoding the target direction information; Based on the third and fourth feature vectors, the spatial distribution characteristics of the second signal are obtained; Based on the historical statistical information of the first audio signal and the spatial distribution characteristics of the second signal, the spatial distribution characteristics of the first signal are obtained.

4. The method according to claim 3, wherein, The historical statistical information includes the speech signal feature information and noise signal feature information of the last frame in the previous time unit. The steps for obtaining the spatial characteristics of the first signal based on historical statistical information of the first audio signal and spatial distribution characteristics of the second signal include: The first signal spatial distribution feature is obtained by calculating the correlation of each RM with respect to the audio signal within the target image region based on the speech signal feature information and the second signal spatial distribution feature of the last frame in the previous time unit, and by calculating the correlation of each RM with respect to the audio signal outside the target image region based on the noise signal feature information and the second signal spatial distribution feature of the last frame in the previous time unit.

5. The method according to claim 2, wherein, The steps for determining the number and location of VMs based on the spatial distribution characteristics of the first signal and the spatial information of the RM include: Based on the spatial distribution characteristics of the first signal, the spatial information of the RM, and the state information of the electronic device, the number and location of the VM are determined.

6. The method according to claim 5, wherein, The steps for determining the number and location of VMs based on the spatial distribution characteristics of the first signal, the spatial information of the RM, and the state information of the electronic device include: The state information of the electronic device is encoded into a state vector; The microphone aperture is determined based on the spatial information provided by the RM; The number and location of VMs are determined based on the spatial distribution characteristics of the first signal, the microphone aperture, and the state vector.

7. The method according to claim 2, wherein, The number and location of VMs determined for each frequency unit within each time unit constitute a three-dimensional VM position feature vector in terms of time, frequency, and space. The step of obtaining the second audio signal based on the first audio signal and the number and location of VMs includes: Obtain the first feature vector corresponding to the first audio signal; and The second audio signal is obtained based on the first feature vector, the VM position feature vector, and the spatial information of the RM.

8. The method according to claim 7, wherein, The steps for obtaining the second audio signal based on the first feature vector, the VM position feature vector, and the spatial information of the RM include: Based on the VM location feature vector and the RM spatial information, the relative position information between the VM and the RM is determined; Based on the relative position information between the VM and the RM and the directionality of the RM, the directionality of each VM is determined; Based on the relative positional information between VM and RM, the first feature vector, and the directionality of each VM, the second feature vector is obtained; The second audio signal is obtained based on the first and second feature vectors.

9. The method according to claim 8, wherein, The steps for obtaining the second audio signal based on the first and second feature vectors include: Based on the first feature vector and the second feature vector, a mask related to the second audio signal within the target image region is obtained; The second audio signal is obtained based on the mask and the first feature vector.

10. The method according to claim 9, wherein, The steps for obtaining a mask related to a second audio signal within a target image region based on a first feature vector and a second feature vector include: By merging the first and second eigenvectors along the spatial direction, a merged eigenvector is obtained; For the merged feature vector, feature processing is performed sequentially in the first direction, the second direction, and the third direction to obtain the mask. Among them, the first direction, the second direction, and the third direction are one of the spatial direction, the temporal direction, and the frequency direction, respectively.

11. The method according to claim 10, wherein, The steps for performing feature processing in the frequency direction for each time unit include: For each frequency unit, performing the following operations: Based on the initial hidden state of each RM in the current frequency unit, the features corresponding to the current frequency unit in the input feature vector corresponding to each RM are processed to obtain the processed features and the processed hidden state corresponding to each RM. The processed hidden state corresponding to each RM is used as the initial hidden state of each RM in the next frequency unit. Based on the initial hidden state of each VM in the current frequency unit, the features corresponding to the current frequency unit in the input feature vector corresponding to each VM are processed to obtain the processed features and the processed hidden state corresponding to each VM. The global hidden state is obtained by performing feature processing on the processed hidden state corresponding to each VM. By performing feature processing on the feature vector corresponding to the next frequency unit in the VM location feature vector and the global hidden state, the initial hidden state of each VM in the next frequency unit is obtained.

12. The method according to claim 10, wherein, The steps for performing feature processing in the time direction for each frequency unit include: for each frame in each time unit, performing the following operations: Based on the initial hidden state of each RM in the current frame, the features corresponding to the current frame in the input feature vector corresponding to each RM are processed to obtain the processed features and the processed hidden state corresponding to each RM. The processed hidden state corresponding to each RM is used as the initial hidden state of each RM in the next frame. Based on the initial hidden state of each VM in the current frame, the features corresponding to the current frame in the input feature vector corresponding to each VM are processed to obtain the processed features and the processed hidden state corresponding to each VM. The global hidden state is obtained by performing feature processing on the processed hidden state corresponding to each VM. By performing feature processing on the feature vector corresponding to the next frame in the VM position feature vector and the global hidden state, the initial hidden state of each VM in the next frame is obtained.

13. The method according to claim 10, wherein, The steps for performing feature processing in the spatial direction for each time unit include: For each RM, perform the following operations: Based on the initial hidden state of the current RM under each frequency unit, process the features corresponding to each frequency unit in the input feature vector corresponding to the current RM to obtain the processed features and the processed hidden state corresponding to each frequency unit. The processed hidden state corresponding to each frequency unit is used as the initial hidden state of each frequency unit under the next RM. For each VM, perform the following operations: Based on the initial hidden state of the current VM in each frequency unit, process the features corresponding to each frequency unit in the input feature vector corresponding to the current VM to obtain the processed features and the processed hidden state corresponding to each frequency unit. By performing feature processing on the processed hidden state corresponding to each frequency unit, obtain the global hidden state. And by performing feature processing on the feature vector corresponding to the next VM in the VM position feature vector and the global hidden state, obtain the initial hidden state of the next VM in each frequency unit.

14. The method according to any one of claims 2-5 and 7-13, wherein, A frequency unit is a frequency point or a sub-band, where a sub-band includes multiple frequency points.

15. The method according to any one of claims 2-5 and 7-13, wherein, A time unit is one or more frames.

16. The method according to claim 9, wherein, The steps for obtaining the second audio signal based on the mask and the first feature vector include: Based on the mask and the first feature vector, the fifth feature vector is obtained; The second audio signal is obtained by performing an audio signal recovery operation on the fifth feature vector.

17. The method according to claim 16, wherein, The steps for obtaining the fifth feature vector based on the mask and the first feature vector include: By processing each sub-band mask in the mask and the corresponding second sub-band feature vector in the first feature vector, multiple third sub-band feature vectors are obtained. By performing feature transformation on the multiple third sub-band feature vectors, multiple predicted features are obtained; By merging the multiple predicted features, a fifth feature vector is obtained.

18. An electronic device comprising: At least one processor; as well as At least one memory that stores computer-executable instructions. Wherein, when the computer-executable instructions are executed by the at least one processor, they cause the at least one processor to perform the method as described in any one of claims 1 to 17.

19. A computer-readable storage medium for storing instructions, wherein, When the instruction is executed by at least one processor, it causes the at least one processor to perform the method as described in any one of claims 1 to 17.

Citation Information

Cited By

  • Sound detection method

    CN122116940A