Method performed by electronic apparatus, electronic apparatus, and storage medium

The method dynamically adjusts virtual microphones using neural networks to optimize beamforming on mobile devices, addressing interference and improving sound extraction accuracy by adapting to changing scenarios and device postures.

WO2026043358A1PCT designated stage Publication Date: 2026-02-26SAMSUNG ELECTRONICS CO LTD

Patent Information

Application Number
PCT/KR2025/099340
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-19
Filing Date
2025-02-06
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Existing audio zoom technologies on mobile devices with limited microphones suffer from wide beam widths, leading to interference from unwanted sounds and inaccurate extraction of desired sounds due to the hardware layout constraints.

Method used

A method that dynamically adjusts the number and positions of virtual microphones based on direction of arrival information, target direction, and spatial information of real microphones, using a neural network to optimize beamforming for accurate sound extraction.

Benefits of technology

Adaptive adjustment of virtual microphones enhances sound extraction accuracy by narrowing the beam width to focus on target sounds while reducing interference, even with changing scenarios and device postures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025099340_26022026_PF_FP_ABST
    Figure KR2025099340_26022026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a method performed by an electronic apparatus, the electronic apparatus, and a storage medium, which involves the field of artificial intelligence. The method includes: obtaining an input first audio signal; obtaining a number and positions of virtual microphones (VMs) based on direction of arrival (DOA) information of the first audio signal, target direction information and spatial information of real microphones (RMs) of the electronic apparatus; and obtaining a second audio signal based on the first audio signal and the number and positions of the VMs. Alternatively, the above method executed by electronic apparatus may be executed by using artificial intelligence models.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD PERFORMED BY ELECTRONIC APPARATUS, ELECTRONIC APPARATUS, AND STORAGE MEDIUM

[0001] The present application relates to the technical field of the signal processing, and in particular, the present application relates to a method for processing an audio signal executed by an electronic apparatus, the electronic apparatus, and a storage medium.

[0002] Currently, audio zoom technology can select sound from a desired direction while reducing sounds from other directions. The audio zoom technology may be used with a video (for example, a camera captures a video), and an audio may be zoomed simultaneously with a video image. However, due to a hardware layout, it is impossible to arrange too many microphones on a mobile phone, thus, at present, the number of microphones on the mobile phone is relatively small (usually there are 3~4 microphones). Under the existing technology, the fewer the microphones, the wider the beam width. If the beam width exceeds an area of interest to a user, some interfering sounds will be extracted, and further a target sound that the user wants cannot be provided accurately.

[0003] How to accurately provide the target sound that the user wants to meet the user's needs is a technical problem that those skilled in the art are continuously working on.

[0004] In order to at least resolve the above problem existing in the prior art, the present disclosure provides a method executed by an electronic apparatus, the electronic apparatus, and a storage medium.

[0005] According to a first aspect of embodiments of the present application, there is provided a method performed by an electronic apparatus, including: obtaining an input first audio signal; obtaining a number and positions of virtual microphones (VMs) based on direction of arrival (DOA) information of the first audio signal, target direction information and spatial information of real microphones (RMs) of the electronic apparatus; and obtaining a second audio signal based on the first audio signal and the number and positions of the VMs.

[0006] In an embodiment, the obtaining the number and positions of the VMs based on the DOA information of the first audio signal, the target direction information and the spatial information of the RMs of the electronic apparatus includes: for each frequency unit at each time unit, performing following operations: obtaining a first signal spatial distribution feature based on the DOA information and the target direction information; and determining the number and positions of the VMs based on the first signal spatial distribution feature and the spatial information of the RMs.

[0007] In an embodiment, the obtaining the first signal spatial distribution feature based on the DOA information and the target direction information includes: obtaining a third feature vector by feature mapping the DOA information; obtaining a fourth feature vector by encoding the target direction information; obtaining a second signal spatial distribution feature based on the third feature vector and the fourth feature vector; and obtaining the first signal spatial distribution feature based on historical statistical information of the first audio signal and the second signal spatial distribution feature.

[0008] In an embodiment, the historical statistical information includes voice signal feature information and noise signal feature information of a last frame in a previous time unit, wherein the obtaining the first signal spatial distribution feature based on the historical statistical information of the first audio signal and the second signal spatial distribution feature includes: obtaining the first signal spatial distribution feature, by calculating correlation of audio signals inside a target image area for each RM based on the voice signal feature information of the last frame in the previous time unit and the second signal spatial distribution feature, and calculating correlation of audio signals outside the target image area for each RM based on the noise signal feature information of the last frame in the previous time unit and the second signal spatial distribution feature.

[0009] In an embodiment, the determining the number and positions of the VMs based on the first signal spatial distribution feature and the spatial information of the RMs includes: determining the number and positions of the VMs based on the first signal spatial distribution feature, the spatial information of the RMs and status information of the electronic apparatus.

[0010] In an embodiment, the determining the number and positions of the VMs based on the first signal spatial distribution feature, the spatial information of the RMs and the status information of the electronic apparatus includes: encoding the status information of the electronic apparatus into a status vector; determining a microphone aperture according to the spatial information of the RMs; and determining the number and positions of the VMs based on the first signal spatial distribution feature, the microphone aperture and the status vector.

[0011] In an embodiment, the number and positions of the VMs determined for each frequency unit at each time unit constitute a three-dimensional (3D) VM position feature vector in a time direction, a frequency direction and a space direction, wherein the obtaining the second audio signal based on the first audio signal and the number and positions of the VMs includes: obtaining a first feature vector corresponding to the first audio signal; and obtaining the second audio signal based on the first feature vector, the VM position feature vector and the spatial information of the RMs.

[0012] In an embodiment, the obtaining the second audio signal based on the first feature vector, the VM position feature vector and the spatial information of the RMs includes: determining a relative position information between the VMs and the RMs based on the VM position feature vector and the spatial information of the RMs; determining a directivity of each VM based on the relative position information between the VMs and the RMs and directivity of the RMs; obtaining the second feature vector based on the relative position information between the VMs and the RMs, the first feature vector and the directivity of each VM; and obtaining the second audio signal based on the first feature vector and the second feature vector.

[0013] In an embodiment, the obtaining the second audio signal based on the first feature vector and the second feature vector includes: obtaining a mask related to the second audio signal inside the target image area based on the first feature vector and the second feature vector; and obtaining the second audio signal based on the mask and the first feature vector.

[0014] In an embodiment, the obtaining the mask related to the second audio signal inside the target image area based on the first feature vector and the second feature vector includes: obtaining a merged feature vector by merging the first feature vector and the second feature vector along the space direction; and obtaining the mask by performing feature processing on the merged feature vector in a first direction, a second direction and a third direction in sequence, wherein the first direction, the second direction and the third direction are one of a space direction, a time direction and a frequency direction, respectively.

[0015] In an embodiment, the performing the feature processing in the frequency direction for each time unit includes: for each frequency unit, performing following operations: processing a feature corresponding to a current frequency unit in an input feature vector corresponding to each RM based on an initial hidden status of each RM at the current frequency unit, to obtain a processed feature and a processed hidden status corresponding to each RM, wherein the processed hidden status corresponding to each RM is used as an initial hidden status of each RM under a next frequency unit; processing a feature corresponding to the current frequency unit in an input feature vector corresponding to each VM based on an initial hidden status of each VM at the current frequency unit, to obtain a processed feature and a processed hidden status corresponding to each VM; obtaining a global hidden status by performing feature processing on the processed hidden status corresponding to each VM; and obtaining an initial hidden status of each VM at the next frequency unit by performing feature processing on a feature vector corresponding to the next frequency unit in the VM position feature vector and the global hidden status.

[0016] In an embodiment, the performing the feature processing in the time direction for each frequency unit includes: for each frame in each time unit, performing following operations: processing a feature corresponding to a current frame in an input feature vector corresponding to each RM based on an initial hidden status of each RM at the current frame, to obtain the processed feature and the processed hidden status corresponding to each RM, wherein the processed hidden status corresponding to each RM is used as an initial hidden status of each RM at a next frame; processing a feature corresponding to the current frame in an input feature vector corresponding to each VM based on an initial hidden status of each VM at the current frame, to obtain a processed feature and a processed hidden status corresponding to each VM; obtaining a global hidden status by performing feature processing on the processed hidden status corresponding to each VM; and obtaining an initial hidden status of each VM at the next frame by performing the feature processing on a feature vector corresponding to the next frame in the VM position feature vector and the global hidden status.

[0017] In an embodiment, the performing the feature processing in the space direction for each time unit includes: for each RM, performing following operations: processing a feature corresponding to each frequency unit in an input feature vector corresponding to a current RM based on an initial hidden status of the current RM at each frequency unit, to obtain a processed feature and a processed hidden status corresponding to each frequency unit, wherein the processed hidden status corresponding to each frequency unit is used as an initial hidden status of each frequency unit at a next RM; and for each VM, performing following operations: processing a feature corresponding to each frequency unit in an input feature vector corresponding to a current VM based on an initial hidden status of the current VM at each frequency unit, to obtain a processed feature and a processed hidden status corresponding to each frequency unit; obtaining a global hidden status by performing feature processing on the processed hidden status corresponding to each frequency unit;, and obtaining an initial hidden status of a next VM at each frequency unit by performing feature processing on a feature vector corresponding to a next VM in the VM position feature vector and the global hidden status.

[0018] In an embodiment, the frequency unit is a frequency point or a sub-band, wherein the sub-band includes a plurality of frequency points.

[0019] In an embodiment, the time unit is one or more frames.

[0020] In an embodiment, the obtaining the second audio signal based on the mask and the first feature vector includes: obtaining a fifth feature vector based on the mask and the first feature vector; and obtaining the second audio signal by performing an audio signal restoration operation on the fifth feature vector.

[0021] In an embodiment, the obtaining the fifth feature vector based on the mask and the first feature vector includes: obtaining a plurality of third sub-band feature vectors by processing each sub-band mask in the mask and a corresponding second sub-band feature vector in the first feature vector; obtaining a plurality of prediction features by performing feature conversion on the plurality of third sub-band feature vectors; and obtaining the fifth feature vector by performing feature merge on the plurality of prediction features.

[0022] According to a second aspect of the embodiments of the present application, there is provided an electronic apparatus, including: at least one processor; and at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, cause the at least one processor to perform the above-described method performed by the electronic apparatus.

[0023] According to a third aspect of the embodiments of the present application, there is provided a computer readable storage medium storing instructions which, when executed by at least one processor, cause the at least one processor to perform the above-described method performed by the electronic apparatus.

[0024] The advantageous effects brought by the technical solutions provided by the embodiments of the present application will be illustrated hereinafter with reference to the specific alternative embodiments, or may be learned from the description of the embodiments, or may be acquainted through implementation of the embodiments.

[0025] In order to illustrate and understand technical solutions of embodiments of the present application more clearly and easily, accompanying drawings that need to be used in the description of the embodiments of the present application will be briefly introduced below.

[0026] FIG. 1 is a schematic diagram illustrating beamforming technology based on signal processing.

[0027] FIG. 2 is a schematic diagram illustrating a beam width formed by using different number of microphones.

[0028] FIG. 3 is a structural schematic diagram illustrating a system of VM estimation technology based on a neural network.

[0029] FIG. 4 is a flowchart illustrating a method executed by an electronic apparatus according to an exemplary embodiment of the present application.

[0030] FIG. 5 is a schematic diagram illustrating a process of a method executed by an electronic apparatus according to an exemplary embodiment of the present application.

[0031] FIG. 6 is a flowchart illustrating a process of obtaining a number and positions of VMs for each frequency unit at each time unit according to an exemplary embodiment of the present application.

[0032] FIG. 7A is a structural schematic diagram illustrating a VM analysis and feature generation module according to an exemplary embodiment of the present application.

[0033] FIG. 7B is a diagram illustrating an example of spatial signal distribution.

[0034] FIG. 7C is a schematic diagram illustrating a process of a method perform by an electronic apparatus according to an exemplary embodiment of the present application.

[0035] FIG. 8 is a diagram illustrating an example of obtaining DOA information by using a Multiple Signal Classification (MUSIC) algorithm.

[0036] FIG. 9 is a schematic diagram illustrating a process of obtaining a first signal spatial distribution feature according to an exemplary embodiment of the present application.

[0037] FIG. 10 is a diagram illustrating examples of a first signal spatial distribution feature and a second spatial distribution feature according to an exemplary embodiment of the present application.

[0038] FIG. 11 is a schematic diagram illustrating correlation between two consecutive frames of audio signals.

[0039] FIG. 12A is a schematic diagram illustrating an affect of a microphone aperture on a beam width.

[0040] FIG. 12B is a schematic diagram illustrating an affect of a folding state of a foldable electronic apparatus on a microphone aperture.

[0041] FIG. 13 is a schematic diagram illustrating a beam formed by a microphone array.

[0042] FIG. 14 is a flowchart illustrating a process of obtaining a second audio signal based on a first audio signal and a number and positions of VMs according to an exemplary embodiment of the present application.

[0043] FIG. 15 is a block diagram illustrating an encoder module according to an exemplary embodiment of the present application.

[0044] FIG. 16 is a flowchart illustrating a network of an encoder module according to an exemplary embodiment of the present application.

[0045] FIG. 17 is a schematic diagram illustrating a process of determining directivity of VM.

[0046] FIG. 18 is a schematic diagram illustrating a process of performing feature processing in a frequency direction.

[0047] FIG. 19 is a system block diagram illustrating a sound extraction module according to an exemplary embodiment of the present application.

[0048] FIG. 20 is a schematic diagram illustrating a process of performing feature processing in a frequency direction according to an exemplary embodiment of the present application.

[0049] FIG. 21 is a schematic diagram illustrating a process of performing feature processing in a time direction according to an exemplary embodiment of the present application.

[0050] FIG. 22 is a schematic diagram illustrating a process of performing feature processing in a space direction according to an exemplary embodiment of the present application.

[0051] FIG. 23 is a block diagram illustrating a decoder module according to an exemplary embodiment of the present application.

[0052] FIG. 24 is a flowchart illustrating a network of a decoder module according to an exemplary embodiment of the present application.

[0053] FIG. 25 is a schematic diagram illustrating a process of training a sound extraction module according to an exemplary embodiment of the present application.

[0054] FIG. 26 is a schematic diagram illustrating a process of jointly training a third neural network and a VM analysis and feature generation module according to an exemplary embodiment of the present application.

[0055] FIG. 27 is a schematic diagram illustrating application of a method performed by an electronic apparatus according to the present application to a scenario of recording a video.

[0056] FIG. 28 is a structural schematic diagram illustrating an electronic apparatus applicable according to an exemplary embodiment of the present application.

[0057] The description is provided below with reference to the accompanying drawings to facilitate comprehensive understanding of various embodiments of the present disclosure as defined by the claims and the equivalents thereof. This description includes various specific details to help with understanding but should only be considered illustrative. Consequently, those ordinarily skilled in the art will realize that various embodiments described here can be varied and modified without departing from the scope and spirit of the present disclosure. In addition, the description of function and structure of the common knowledge can be omitted for clarity and conciseness.

[0058] The terms and expressions used in the description and claims below are not limited to their lexicographical meaning but are used only by the inventor to enable the clear and consistent understanding of the present disclosure. Therefore, it should be apparent to those skilled in the art that the following description of the various embodiments of the present disclosure is provided only for the purpose of the illustration without limiting the present disclosure as defined by the appended claims and their equivalents.

[0059] It will be understood that, unless specifically stated, the singular forms "one", "a", and "said" used herein may also include the plural form. Thus, for example, "component surface" refers to one or more such the surfaces. When it refers to one element as being "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or it may refer to a connection relationship between the one element and the other element established through an intermediate element. In addition, "connected" or "coupled" as used herein may include wirelessly connected or wirelessly coupled.

[0060] The terms "includes" and "may include" mean the presentation of the corresponding disclosed functions, operations, or components that can be used in various embodiments of the present disclosure, but do not limit the presentation of one or more additional functions, operations, or features. In addition, it should be understood that the terms "including" or "having" may be interpreted to mean certain features, numbers, steps, operations, components, assemblies or combinations thereof, but should not be interpreted to exclude the possibility of the existence of one or more of other features, numbers, steps, operations, components, assemblies and / or combinations thereof.

[0061] The term "or" as used in various embodiments of the present disclosure includes any listed term and all the combinations thereof. For example, "A or B" may include "A", may include "B", or may include both "A and B". When describing a plurality of (two or more) items, the plurality of items may refer to one, more, or all of the plurality of items if a relationship among the plurality of items is not explicitly defined. For example, for the description "a parameter A comprises A1, A2, A3", it may be implemented as parameter A comprising A1, A2 or A3, or as parameter A comprising at least two of the three items of the parameter A1, A2, A3.

[0062] All terms (including technical or scientific terms) used in the present disclosure have the same meaning as understood by those skilled in the art to which the present disclosure belongs, unless defined differently. Common terms as defined in dictionaries are interpreted to have a meaning consistent with the context in the relevant technology art and should not be interpreted in an idealized or overly formalistic manner, unless expressly so defined in the present disclosure.

[0063] At least part of the functions in a device or electronic apparatus provided in the embodiments of the present disclosure may be implemented through an AI model, such as, at least one of a plurality of modules of the device or electronic apparatus may be implemented through the AI model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor.

[0064] The processor may include one or more processors. At this time, the one or more processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, or may be a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU).

[0065] One or more processors control the processing of input data with a predefined operation rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.

[0066] Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or an AI model of a desired characteristic is made. The learning may be performed in a device or electronic apparatus itself in which AI according to an embodiment is performed, and / or may be implemented through a separate server / system.

[0067] The AI model may consist of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a neural network calculation by calculating between the input data of this layer (such as, a calculation result of the previous layer and / or the input data of the AI model) and the plurality of weight values of the current layer. Examples of neural networks include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial networks (GAN), and a deep Q-network.

[0068] The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0069] The methods provided in the present disclosure may involve one or more of technical fields such as speech, language, image, video, or data intelligence.

[0070] When involving the field of speech or language, in the method according to an embodiment of the present disclosure executed by electronic apparatus, a speech signal, which is an analog signal, may be received via speech input devices (e.g., a microphone), and the speech signal is converted into computer readable text using an automatic speech recognition (ASR) model. The user's intent of utterance may be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or NLU model may be an artificial intelligence model. The artificial intelligence model may be processed by an artificial intelligence-dedicated processor designed in a hardware structure specified for artificial intelligence model processing. Language understanding is a technique for recognizing and applying / processing human language / text and includes, e.g., natural language processing, machine translation, dialog system, question answering, or speech recognition / synthesis.

[0071] When involving the field of image or video, in the method according to an embodiment of the present disclosure executed by electronic apparatus, output data may be obtained by using image data as input data for an artificial intelligence model. The method of the present disclosure may be related to the field of visual understanding in the artificial intelligence technology, and the visual understanding is a technique for recognizing and processing things as human vision does and includes, e.g., object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.

[0072] When involving the field of data intelligence processing, in the method according to an embodiment of the present disclosure executed by electronic apparatus, in the reasoning or predicting stage, an artificial intelligence model can be used to perform predictions by using real-time input data. Processors of the electronic apparatus may perform a pre-processing operation on the data to convert into a form appropriate for use as an input for the artificial intelligence model. Reasoning and prediction is a technique of logically reasoning and predicting by determining information and includes, e.g., knowledge-based reasoning, optimization prediction, preference-based planning, or recommendation.

[0073] In the present application, the artificial intelligence model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training algorithm. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.

[0074] There are various kinds of regional sound extraction technology, which mainly include beamforming technology based on signal processing and virtual microphone (VM) estimation technology based on neural network.

[0075] The beamforming technology based on signal processing may be regarded as a spatial filter. As shown in FIG. 1, the beamforming technology may use following Equation (1) to process signals received by a plurality of microphones according to a weighted vector w to obtain an enhanced target signal y.

[0076] (1)

[0077] where k denotes a time, denotes a signal received by an ithmicrophone.

[0078] However, for the beamforming technology based on signal processing, it is impossible to arrange too many microphones on a mobile phone due to a hardware layout of the mobile phone. As shown in FIG. 2, the fewer the microphones, the wider the beam width, for example, a box indicated by 201 is used to represent that a beam width formed by an array containing 4 microphones of a plurality of microphones 20 has reached 60°, while a beam width formed by an array containing 3 microphones will be wider. For example, a width of a first beam 31 formed by an array containing 4 microphones of the plurality of microphones 20 is wider and a width of a second beam 32 formed by an array containing the plurality of microphones 20 including 8 microphones is narrower. 8 microphones 21, 22, 23, 24, 25, 26, 27 and 28 are shown in FIG. 2, but the number of microphones is not limited.

[0079] If the beam width exceeds an area of interest to a user, some interfering sounds will be extracted, and a target sound that the user wants cannot be provided accurately. For example, sound of a car 41 occurs beyond the area of interest to a user. Since the width of the first beam 31 contains a source of car sound, some interfering sound from the car 41 will be extracted. Since the width of the second beam 32 does not contain the source of car sound, some interfering sounds from the car 41 will be suppressed relatively well. For example, sound of children 42 occurs beyond the area of interest to a user. Since the width of the first beam 31 and the width of the second beam 32 do not contain a source of children sound, some interfering sounds from the children 42 will be suppressed well in the each of the cases. Thus, it needs to consider a challenge of a small number of microphones while using the beamforming technology based on signal processing on the mobile phone.

[0080] In addition, the VM estimation technology based on neural network estimates VM signals directly in a time domain. The technology relies on a supervised learning framework and builds a model that can predict VM signals based on observation results of real microphones (RMs). Specifically, as shown in FIG. 3, under constraint of a received signal k at an actual position 2, a microphone signal at the position 2 is estimated using time domain signals of RM1 at an actual position 1 and RM3 at an actual position 3, wherein positions of all microphones are preset while training a model. As shown in FIG. 3, an encoder performs feature encoding on input signals j and l received by RM1 and RM3, a feature processing module processes an encoded feature, and a decoder estimates a signal k' of VM2, so that the microphone signal at the position 2 is virtualized as the signal k' of VM2 according to the signals j and l received by RM1 and RM3.

[0081] However, for the VM estimation technology based on neural network, since the number and positions of VMs are fixed, and positions of a target source and an interference source are dynamically changing in real scenarios, if the interference source is very close to the target source, interference will leak in because the beam width isn't narrow enough (due to power consumption, too many VMs cannot be used). In addition, the VM estimation technology does not consider a change of a device posture. When the device posture changes, relative positions of RMs may also change (e.g., after folding a foldable mobile phone, each RM will become more closer), and correspondingly, microphone arrays formed by RMs and the fixedly generated VMs will also undergo changes, further resulting in changes in the beam width and direction, and making it difficult to accurately extract a target sound. In addition, the VM estimation technology does not consider that the beam width varies at different frequencies under same array conditions. For example, high-frequency beams are usually narrower, while low-frequency beams are wider, then low-frequency components may contain more interference, as the interference source may be located within the low-frequency beam's area.

[0082] Therefore, the present application provides a method executed by an electronic apparatus. The method can adaptively virtualize different number of microphones at different positions for different frequencies at different times according to changes in the scenario, the changes in the interference source and changes in the device posture at different times, so that the formed beam is narrower, thereby reducing or eliminating sound interference outside a target area (i.e, an area of interest), and improving sound extraction performance (such as accuracy) inside the target area while extracting a sound inside the target area.

[0083] Below, the technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure will be explained by describing several optional embodiments. It should be noted that, the following embodiments may be referred to, imitated or combined with each other, and the same term, similar features and similar implementation steps in different embodiments will not be described repeatedly.

[0084] FIG. 4 is a flowchart illustrating a method executed by an electronic apparatus according to an exemplary embodiment of the present application. FIG. 5 is a schematic diagram illustrating a process of the method executed by the electronic apparatus according to an exemplary embodiment of the present application.

[0085] As shown in FIG. 4, in step S410, an input first audio signal is obtained. In the present application, the first audio signal may be a mixed audio signal picked up from an environment by RMs of the electronic apparatus, which is a time domain signal.

[0086] In step S420, a number and positions of VMs are obtained based on direction of arrival (DOA) information of the first audio signal, target direction information, spatial information of RMs of the electronic apparatus. This operation may be performed by a VM analysis and feature generation module 510 in FIG. 5, and the process of obtaining the number and positions of VMs is performed for each frequency unit at each time unit, wherein the frequency unit may be a frequency point or a sub-band, and wherein the sub-band includes a plurality of frequency points, and the time unit is one or more frames. This will be described in detail with reference to FIG. 6 and FIG. 7A below.

[0087] FIG. 6 is a flowchart illustrating a process of obtaining a number and positions of VMs for each frequency unit at each time unit according to an exemplary embodiment of the present application. FIG. 7A is a structural schematic diagram illustrating the VM analysis and feature generation module according to an exemplary embodiment of the present application.

[0088] In step S610, a first signal spatial distribution feature is obtained based on the DOA information and the target direction information. In the present application, the target direction information may be direction of user interested (DOI) information, such as information of a direction specified by a user operating a screen of an electronic apparatus (e.g., a direction specified by an operation of zooming the screen or selecting a specific area on the screen), a direction aimed by a camera of the electronic apparatus, or the like. In addition, in the present application, signal spatial distribution may represent a specific direction of a target audio signal or an interference signal in a space, as shown in FIG. 7B, and it may decide a direction and a width of an array beam. A narrower beam may be obtained by arranging more VMs in the direction of the target audio signal or interference signal.

[0089] Specifically, at first, a third feature vector is obtained by performing feature mapping on the DOA information. This operation is performed by a feature mapping module 711 in FIG. 7A. For example, in the present application, the third feature vector v3may be obtained by normalizing the DOA information and mapping the normalized DOA information to feature distribution in a feature space. In an exemplary embodiment, the present application may obtain the DOA information by using a Multiple Signal Classification (MUSIC) algorithm, as shown in FIG. 8. After the DOA information is obtained, the DOA information is normalized and the normalized DOA information is mapped to the feature distribution in the feature space through the feature mapping module 711 in FIG. 7A to thereby obtain the third feature vector v3(also referred to as a DOA mapped space feature), as shown in FIG. 9. The feature space represents a signal strength and a probability of existence in each direction in a 0° to 360° space.

[0090] Then, a fourth feature vector is obtained by encoding the target direction information. In the present application, when a user zooms in or out a screen of an electronic apparatus or selects an area of interest on the screen, the target direction information may be obtained from a camera module of the electronic apparatus, and in the present application, it may also be referred to as DOI information (also referred to as a DOI parameter), that is, a direction (i.e., an azimuth and an elevation) and an area width (i.e., a width of the area of interest) are obtained. After the target direction information is obtained, the target direction information may be encoded by an encoding module 712 in FIG. 7A to obtain the fourth feature vector v4(which may also be referred to as an encoded vector of the target direction information).

[0091] Hereafter, a second signal spatial distribution feature is obtained based on the third feature vector and the fourth feature vector. For example, as shown in FIG. 7A, the second signal spatial distribution feature (which may be briefly referred to as a second spatial feature) may be obtained by multiplying the third feature vector v3obtained from the feature mapping module 711 and the fourth feature vector v4obtained from the encoding module 712, as shown in FIG. 10. The process may also be referred to as performing feature labeling on the DOA mapped space feature using the encoded vector of the target direction information, as shown in FIG. 9. The second signal spatial distribution feature illustrates spatial distribution of each audio signal in the space.

[0092] Afterwards, the first signal spatial distribution feature is obtained based on historical statistical information of the first audio signal and the second signal spatial distribution feature . As shown in FIG. 11, for two consecutive frames of audio signals (for example, a nth frame and a n+1th frame), extracted signals of the second frame and the first frame have high correlation, and the extracted signal of the first frame has been distinguished between a signal inside the area of interest and a signal outside the area of interest by the algorithm of the present application. Thus, the second spatial distribution feature may be updated through an update module 713 in FIG. 7A using the historical statistical information to further obtain the more accurate first signal spatial distribution feature , as shown in FIG. 9. Specifically, the historical statistical information includes voice signal feature information and noise signal feature information of a last frame in a previous time unit, and in this case, the obtaining the first signal spatial distribution feature based on the historical statistical information of the first audio signal and the second signal spatial distribution feature may include: obtaining the first signal spatial distribution feature , by calculating correlation of audio signals inside a target image area for each RM based on the voice signal feature information of the last frame in the previous time unit and the second signal spatial distribution feature , and calculating correlation of audio signals outside the target image area for each RM based on the noise signal feature information of the last frame in the previous time unit and the second signal spatial distribution feature . As shown in FIG. 10, the first signal spatial distribution feature may illustrate spatial distribution of each audio signal in the space more accurately than the second signal spatial distribution feature .

[0093] In an exemplary embodiment of the present application, the voice signal feature information may be a voice signal covariance matrix of a last frame in a previous time unit, and the noise signal feature information may be a noise signal covariance matrix of the last frame in the previous time unit. Specifically, the voice signal feature information may be obtained according to an extracted feature of the last frame in the previous time unit (as shown in FIG. 7C), that is, the voice signal covariance matrix is calculated, and the noise signal feature information may be obtained according to the extracted feature of the last frame in the previous time unit, that is, the noise signal covariance matrix is calculated. Then, correlation of audio signals inside the target image area for each RM is calculated according to the voice signal covariance matrix and the second signal spatial distribution feature , and correlation of audio signals outside the target image area for each RM is calculated based on the noise signal covariance matrix of the last frame in the previous time unit and the second signal spatial distribution feature , thereby obtaining the first signal spatial distribution feature . However, the voice signal covariance matrix and the noise signal covariance matrix are only one example of the voice signal feature information and the noise signal feature information, and the present application is not limited hereto.

[0094] In step S620, the number and positions of the VMs are determined based on the first signal spatial distribution feature and the spatial information of the RMs. Specifically, the step S620 may include: determining the number and positions of the VMs based on the first signal spatial distribution feature, the spatial information of the RMs and status information of the electronic apparatus. More specifically, the determining the number and positions of the VMs based on the first signal spatial distribution feature, the spatial information of the RMs and the status information of the electronic apparatus may include: encoding the status information of the electronic apparatus into a status vector; determining a microphone aperture according to the spatial information of the RMs; and determining the number and positions of the VMs based on the first signal spatial distribution feature, the microphone aperture and the status vector.

[0095] Specifically, the above operation of determining the number and positions of the VMs based on the first signal spatial distribution feature, the microphone aperture and the status vector may be realized through a first neural network. In the present application, the first neural network (which may also be referred to as a VM analysis module) may be trained according to following criteria for estimating the number of VMs:

[0096] (1) In order to prevent the first neural network from virtualizing the number of VMs into a maximum number each time, the present application encodes the status information of the electronic apparatus into a status vector to further limit the number of VMs and reduce power consumption. In an exemplary embodiment of the present application, the status information of the electronic device may include remaining power and / or remaining operating resources of the electronic apparatus.

[0097] (2) In acoustic field, a microphone aperture acts as an electroacoustic transducer that converts acoustic signals into electrical signals. A length of the aperture, which is a maximum distance between RMs, and a number of spatial sampling points on the aperture (the number of RMs) will affect a shape of the beam. That is to say, the length of the aperture decides the number of deployable RMs to a certain extent. Generally, the longer the microphone aperture length, the more the deployable VMs, and the beam will become narrower correspondingly; the shorter the microphone aperture length, the fewer the deployable VMs, and the beam will become wider correspondingly, as shown in FIG. 12A. For a general electronic apparatus, the distance between RMs usually does not vary, but for a foldable electronic apparatus, the distance between RMs will vary as the folding state of the electronic apparatus varies. As shown in FIG. 12B, for a foldable electronic apparatus, when the electronic apparatus changes from an unfolded state to a folded state, a change in a posture of the electronic apparatus causes a relative distance of RMs to become smaller, thereby further causing the microphone aperture to become smaller. In this case, the number of RMs will theoretically become less. Thus, the present application may obtain a posture of an electronic apparatus, determine spatial information of RMs (e.g., relative position of RMs) according to the posture of the electronic apparatus, and then calculate a microphone aperture according to the spatial information of RMs.

[0098] (3) For different frequencies, under same microphone array conditions, high-frequency beams are usually narrower, while low-frequency beams are usually wider. Thus, the first neural network of the present application is trained to allocate more VMs for low frequency and less VMs for high frequency. This is because more audio signals outside an area of interest (which may be regarded as interference signals) are easily extracted for the wider beam.

[0099] In addition, the first neural network may be trained according to following criteria to estimate a position of VM: a direction and width of a beam 1310 formed by a microphone array is calculated according to the first signal spatial distribution feature, as shown in FIG. 13, wherein the beam 1310 is used to cover target audio signals inside an area of interest; and then the position of VM is estimated according to the number of VMs and the direction and width of the beam 1310.

[0100] In conclusion, in a reasoning phase, after the first spatial distribution feature, the microphone aperture and the status vector are obtained, these information may be used to predict the number and positions of the VMs for different frequency units at different time units (i.e., at different times) through the trained first neural network 714 in FIG. 7A. Since the number and positions of the VMs are predicted for each frequency unit at each time unit, the number and positions of the VMs determined for each frequency unit at each time unit constitute a three-dimensional (3D) VM position feature vector in a time direction, a frequency direction and a space direction. In the present application, "the time unit" mentioned here may represent "a time interval" having a preset time length T, for example, the time unit may be a time of a frame, e.g., 16ms. In addition, in order to further reduce complexity of the related neural network used by the present application, the time unit may also be a few milliseconds to a few seconds, in other words, the time unit may be one or more frames. In addition, the number and positions of the VMs are not updated and remain unchanged within a same time unit. Besides, "the frequency unit" mentioned here may represent a frequency point or a sub-band, and the number and positions of the VMs are not updated and remain unchanged for each sub-band, wherein the frequency point and the sub-band will be described in detail subsequently with reference to FIG. 14. In a simplified example, as shown in FIG. 7C, the VM analysis module 710 may be run only once at the beginning of each time unit, and further the VM position feature vector for each frequency unit at the time unit is obtained. That is to say, the number and positions of the VMs may be different and remain unchanged at different frequency units within the time unit.

[0101] Returning to refer to FIG. 4, in step S430, a second audio signal is obtained based on the first audio signal and the number and positions of the VMs. This will be described in detail with reference to FIG. 14 below.

[0102] As shown in FIG. 14, in step S1410, a first feature vector corresponding to the first audio signal is obtained.

[0103] Specifically, the obtaining the first feature vector corresponding to the first audio signal may include: extracting a feature vector from the first audio signal; and obtaining the first feature vector by performing feature encoding on the extracted feature vector.

[0104] At first, feature extraction is performed on the first audio signal picked up to obtain a feature vector in another dimension. For example, Short-Time Fourier Transform (STFT) such as 512-point STFT may be used to perform feature extraction, that is, frame segmentation and windowing and STFT are performed on the first audio signal to obtain a feature on a frequency domain.

[0105] For example, for a first audio signal with a time length of n seconds at a 16k sampling rate, there is L=n*16000 sampling-points data. By performing the STFT with a window length of W=s_n sampling points (that is, a number of sampling points per frame is s_n, and an overlap area between frames is s_n / 2 (that is, 50% overlap), that is, a frame shift is W / 2), a number k of frames is k={L / (s_n / 2)}-1, and a number f of frequency points per frame is f=s_n / 2, from which real and imaginary parts of the frequency domain are extracted, respectively, that is, a feature vector of a dimension [k,f] may be obtained. For example, for a first audio signal with a time length of 4s at a 16k sampling rate, after performing STFT with a window length of W =512 sampling points (that is, a frame shift is 256 sampling points), the number of frames is 249, and the number f of the frequency points per frame is s_n / 2 =512 / 2 =256. Each frequency point is represented by a real part and an imaginary part. Therefore, a feature vector with a dimension of [249, 256] may be obtained.

[0106] STFT is used to perform feature extraction in the above example, but the present application is not limited hereto. Other feature extraction methods may also be used, for example, Convolutional Neural Networks (CNNs) are used to perform feature extraction.

[0107] After the feature vector is obtained by performing the feature extraction, the first feature vector is obtained by performing feature encoding on the extracted feature vector.

[0108] In an exemplary embodiment of the present application, an encoder (e.g. encoder module 530 shown in FIG. 5) may be directly used to encode the extracted feature to obtain the first feature vector. In view of the fact that the present application needs to virtualize different microphone characteristics for different frequencies, in another exemplary embodiment of the present application, a frequency is divided into a plurality of sub-bands, and the frequency points in the same sub-band use the features of the same VMs (that is, the same number and positions of the VMs) at the same time so as to reduce computational complexity. This will be described in detail with reference to FIGS. 15 and 16 below. If the frequency is not divided into a plurality of sub-bands, a full-band processing method is used, and the VM analysis and feature generation module described subsequently will virtualize the same number and positions of VMs for different frequencies (i.e., frequency points).

[0109] FIG. 15 is a block diagram illustrating an encoder module according to an exemplary embodiment of the present application. FIG. 16 is a flowchart illustrating a network of an encoder module according to an exemplary embodiment of the present application.

[0110] As shown in FIG. 15, the encoder module 1500 includes a feature extraction module 1510, a sub-feature division module 1520 and a plurality of sub-encoder 1530, wherein the feature extraction module 1510 is used to perform the operation of extracting the feature vector from the first audio signal as described above. This has been described in detail above, and thus will not be repeated here again.

[0111] As shown in FIG. 16, the feature extraction module may perform short-term Fourier transform in step S1610 to transform the first audio signal into feature vectorF. the sub-feature division module performs frequency band division in step S1620 to divide the extracted feature vectorFinto a plurality of first sub-band feature vectors. For example, a 16k frequency band is divided into N sub-bands. In consideration of performance and model complexity, N may be equal to 4, 5 or 6 in an example. However, the present application is not limited hereto. For example, as shown in FIG. 16, a 16k frequency band may be divided into 4 sub-bands. For the frequency domain data obtained in the previous step, 256 frequency-points data per frame (i.e., the extracted feature vector) may be divided into 4 first sub-band feature vectors , , and . The frequency points included in respective first sub-band feature vectors are respectively {1~32}, {33~64}, {65~128} and {129~256}, and the corresponding frequencies are respectively 0~2k, 2k~4k, 4k-8k and 8k~16k.

[0112] After the extracted feature vectorFis divided into the plurality of first sub-band feature vectors, a plurality of second sub-band feature vectors are obtained in step S1630 by using the corresponding sub encoder to encode each of the plurality of first sub-band feature vectors, wherein the first feature vector includes the plurality of second sub-band feature vectors. As shown in FIGS. 15 and 16, the number N of the divided first sub-band feature vectors is 4, there are N=4 sub-encoders 1530 correspondingly, each first sub-band feature vector is input to a corresponding sub-encoder (e.g., two-dimensional (2D) Convolutional Neural Network (2D-CNN)) for encoding to obtain 4 second sub-band feature vectors , , and , wherein the obtained plurality of second sub-band feature vectors may be collectively referred to as the first feature vector. By dividing the extracted feature vector into the plurality of first sub-band feature vectors, and using different sub-encoders to encode the corresponding first sub-band feature vectors, the present application may realize parallel encoding, thereby reducing complexity, and improving a processing speed of the model.

[0113] Returning to refer to FIG. 14, in step S1420, the second audio signal is obtained based on the first feature vector, the VM position feature vector and the spatial information of the RMs.

[0114] The VM position feature vector (i.e., the number and positions of VMs determined for each frequency unit at each time unit) determined in step S620 above does not consider directivity of RM, wherein the directivity of RM indicates that a microphone picks up sounds from different directions, which decides sensitivity and signal reception ability of the microphone at different angles. The directivity of the microphone may be divided into Omnidirectional, Cardioid, Wide Cardioid and Super Cardioid, wherein a microphone with Omnidirectional directivity may pick up audio signals from all directions indiscriminately; a microphone with Cardioid directivity may mainly pick up audio signals ahead; a microphone with Wide Cardioid directivity may be similar to the microphone with Omnidirectional directivity, and pick up audio signals from all directions, but its ability to pick up audio signals behind is slightly weaker than the microphone with Omnidirectional directivity; and a microphone with Super Cardioid directivity may mainly pick up audio signals ahead and behind, but its ability to pick up audio signals behind is slightly weaker than the ability to pick up audio signals ahead. For an electronic apparatus such as a mobile phone, it generally includes three RMs with different directivity, wherein directivity of top and bottom RMs is generally Omnidirectional or Wide Cardioid, and directivity of back RM is generally Cardioid or Super Cardioid. Thus, in order to enable the microphone array to pick up more signals inside the area of interest and strong interference signals outside the area of interest, the present application may add directivity to the determined VMs according to the relative position between VMs and RMs, thereby enhancing discrimination between the signals inside the area of interest and the interference signals outside the area of interest to better extract the signals inside the area of interest.

[0115] Specifically, the step S1420 may include: determining relative position information between the VMs and the RMs based on the VM position feature vector and the spatial information of the RMs; determining directivity of each VM based on the relative position information between the VMs and the RMs and the directivity of the RM; obtaining the second feature vector based on the relative position information between the VMs and the RMs, the first feature vector and the directivity of each VM; and obtaining the second audio signal based on the first feature vector and the second feature vector.

[0116] In an exemplary embodiment of the present application, as shown in FIG. 7A, relative position information between each VM and each RM may be determined according to the VM position feature vector and the spatial information of RM by a first sub-network (i.e., a directivity processing module 721) in a second neural network (i.e., a VM feature generation module 720). Hereafter, directivity PVMof each VM may be determined according to the relative position information between each VM and each RM and directivity PRMof a RM closest to each VM by the first sub-network. In detail, directivity PVMof each VM may be preset according to the relative position information between each VM and each RM and directivity PRMof the RM closest to each VM by the first sub-network, the directivity PVMof each VM may be adjusted to more point to the signals inside the area of interest or the interference signals according to a signal direction, thereby determining the directivity of each VM, as shown in FIG. 17.

[0117] For example, directivity PVM,1of VM1 may be preset according to the relative position information between VM1 and each RM and directivity PRMof the RM closest to VM1 by the first sub-network, the directivity PVM,1of VM1 may be adjusted to more point to the signals inside the area of interest(e.g. audio signal 1710) or the interference signals according to a signal direction, thereby determining the directivity of VM1.

[0118] For example, directivity PVM,2of VM2 may be preset according to the relative position information between VM2 and each RM and directivity PRMof the RM closest to VM2 by the first sub-network, the directivity PVM,2of VM2 may be adjusted to more point to the signals inside the area of interest or the interference signals(e.g. audio signal 1720) according to a signal direction, thereby determining the directivity of VM2.

[0119] After the directivity of each VM is determined, as shown in FIG. 7A, the second feature vector is obtained based on the relative position information between the VMs and the RMs, the first feature vector and the directivity of each VM, through a second sub-network (i.e., an amplitude and phase processing module 722) of the second neural network. Specifically, the second sub-network may calculate phase difference and amplitude attenuation between RM and VM based on the relative position information between VM and RM and the first feature vector, and then generate the second feature vector based on the directivity of each VM and the phase difference and amplitude attenuation between RM and VM.

[0120] After the second feature vector is obtained, the second audio signal may be obtained based on the first feature vector and the second feature vector. This will be described in detail below.

[0121] At first, a mask related to the second audio signal inside the target image area is obtained based on the first feature vector and the second feature vector, and this operation may be performed by a sound extraction module 520 in FIG. 5.

[0122] Specifically, since the number and positions of VMs usually vary at different time units and frequency units, the second feature vector is not continuous in the time direction, space direction (also referred to as microphone direction) or frequency direction, so that it is difficult for a neural network to transmit hidden statuses while processing these features. For example, as shown in FIG. 18, for RM1, a hidden status is initialized to 0, and after the neural network uses the hidden status to process the feature at the frequency unitf=1, the hidden status is updated; hereinafter, the neural network may use the updated hidden status to process a next feature (i.e., a feature at the frequency unitf=2), thereby completing transmission of feature information from the frequency unitf=1to the frequency unitf=2. However, as shown in FIG. 18, for VM, it is difficult for a neural network to transmit hidden statuses while processing the features. For example, for VM2, a hidden status is initialized to 0, and after the neural network uses the hidden status to process the feature at the frequency unitf=1, the hidden status is updated.However, there is no VM2 at the frequency unitf=2, and further there is no feature data related to VM2.Thus, the hidden status cannot be transmitted smoothly. Subsequently, there is VM2 at the frequency unitf=4, and correspondingly, there is the feature data related to VM2. If the previous hidden status is directly used, the neural network may be made impossible to learn the feature information at the frequency unitsf=2andf=3; similarly, for VM1, since there is no VM1 at the frequency unitf=1, and further there is no feature data related to VM1, no new hidden status is generated. Thus, an initial hidden status =0 will not be updated, so that feature information from the frequency unitf=1will be lost during the processing, thereby affecting the performance of feature processing.

[0123] Since the VM position feature vector is used while generating the second feature vector, the second feature vector includes spatial information. The present application processes correlation between the VM position feature vector of a current step (e.g., a current frequency unit) and the hidden status of the previous step (the previous frequency unit) through following two operations to realize normal transfer of the hidden status during the network feature processing: (1) after processing the feature corresponding to the previous step in the feature vector corresponding to each VM, obtaining a global hidden status by performing feature processing on hidden statuses of all VMs, wherein a fused weight of each hidden status may be learned through a network or a weighted sum; for example, when a certain VM existing in the previous step is closer to the VM existing in the current step, a weight of the certain VM existing in the previous step is greater; however, the present application is not limited hereto; and (2) while processing the feature corresponding to the current step in the feature vector corresponding to each VM, for a certain VM, updating a hidden status of the certain VM in the current step by calculating the correlation between a position feature of the certain VM in the VM position feature vector and a spatial feature included in the global hidden status.

[0124] Based on this principle, by processing the feature vector in the time direction, the space direction (also referred to as microphone direction) and the frequency direction, the present applicant realizes the normal transfer of the hidden status during the network feature processing to thereby improve accuracy of a final prediction result.

[0125] In an exemplary embodiment of the present application, the obtaining the mask related to the second audio signal inside the target image area based on the first feature vector and the second feature vector may include: obtaining a merged feature vector by merging a first feature vectorXand a second feature vectorX_vmalong the space direction; obtaining the mask by performing feature processing on the merged feature vector in a first direction, a second direction and a third direction in sequence, wherein the first direction, the second direction and the third direction are one of the space direction, the time direction and the frequency direction, respectively, and in addition, the first direction, the second direction and the third direction are different from each other, that is to say, when the feature processing is performed in the first direction, an input feature vector is the merged feature vector; when the feature processing is performed in the second direction, the input feature vector is a feature vector obtained after performing the feature processing in the first direction; when the feature processing is performed in the third direction, the input feature vector is a feature vector obtained after performing the feature processing in the second direction, and after the feature processing is performed in the third direction, a finally obtained feature vector may be represented as the mask. That is to say, the feature processing may be performed in any one of following orders: (1) an order of the space direction, the time direction and the frequency direction; (2) an order of the space direction, the frequency direction and the time direction; (3) an order of the time direction, the space direction, and the frequency direction; (4) an order of the time direction, the frequency direction and the space direction; (5) an order of the frequency direction, the space direction and the time direction; and (6) an order of the frequency direction, the time direction and the space direction. A process of performing the feature processing in the order of the frequency direction, the time direction and the space direction will be described in detail with reference to FIGS. 19 to 22 below.

[0126] FIG. 19 is a system block diagram illustrating a sound extraction module according to an exemplary embodiment of the present application. FIG. 20 is a schematic diagram illustrating a process of performing feature processing for each time unit in a frequency direction according to an exemplary embodiment of the present application.

[0127] In the present application, the frequency unit may be a frequency point or a sub-band. When the frequency unit is a sub-band, the sub-band may include a plurality of frequency points. Hereafter, the description is made by taking the frequency unit being a sub-band as an example. In the present application, a sound extraction module 1900 may also be referred to as a third neural network, including a first feature processing module 1910, a second feature processing module 1920, a third feature processing module 1930, a first fusion module 1940, a second fusion module 1950, a third fusion module 1960, a first update module 1970, a second update module 1980 and a third update module 1990, wherein the first feature processing module to the third feature processing module 1910-1930 may use recurrent neural networks (RNNs), but the present application is not limited hereto, and other networks having time sequence processing ability (such as CNNs, attention networks, long short-term memory (LSTM) networks, and the like) may also be used. Similarly, the first fusion module to the third fusion module 1940-1960 and the first update module to the third update module 1970-1990 may use CNNs, RNNs or the like.

[0128] In FIGS. 19 and 20, represents a hidden status set of a network after the first feature processing module 1910 or the second feature processing module 1920 of FIG. 19 processes a feature corresponding to a kthsub-band in a feature vector corresponding to each VM, wherein N represents the number of VMs; represents an initialized hidden status corresponding to an ithVM before processing a feature corresponding to a kthfrequency unit in a feature vector corresponding to the ithVM; represents a hidden status set of a network after the third feature processing module 1930 of FIG. 19 processes a feature corresponding to each sub-band in the feature vector corresponding to the ithVM; and represents a position feature vector of the ithVM corresponding to a sub-bandf=k(also referred to as ) at time t.

[0129] As shown in FIGS. 19 and 20, when processing a feature vector corresponding to RM, for each sub-band, a feature corresponding to a current sub-band in an input feature vector corresponding to each RM is processed based on an initial hidden status of each RM at the current sub-band, to obtain a processed feature and a processed hidden status corresponding to each RM, wherein the processed hidden status corresponding to each RM is used as an initial hidden status of each RM at a next sub-band. This processing operation may be performed by the first feature processing module 1910 of the third neural network in FIG. 19. Since the feature processing is performed in the frequency direction at first, the input feature vector corresponding to each RM is a first feature vector in the merged feature vector, wherein as described above, the merged feature vector is obtained by merging the first feature vector and the second feature vector along the space direction, wherein the merged feature vector in a 3D form is shown in FIGS. 19 and 20 (e.g. feature vector in 3D 50), but the present application is not limited hereto, and the merged feature vector may further be a feature vector in higher dimension, such as 4D, 5D or the like. The process of processing the feature vector of RM is the same as the process described with reference to FIG. 18, and thus will not be repeated here again.

[0130] When processing the feature vector corresponding to VM, an input feature vector corresponding to each VM is the second feature vector in the merged feature vector. As shown in FIGS. 19 and 20, for each sub-band, following operations are performed: processing a feature corresponding to a current sub-band in an input feature vector corresponding to each VM based on an initial hidden status of each VM at the current sub-band, to obtain a processed feature and a processed hidden status corresponding to each VM; obtaining a global hidden status by performing feature processing on the processed hidden status corresponding to each VM; and obtaining an initial hidden status of each VM at a next sub-band by performing the feature processing on a feature vector corresponding to the next sub-band in the VM position feature vector and the global hidden status.

[0131] As shown in FIG. 20, for a current sub-bandf=1(i.e.,f1), the first feature processing module of the third neural network processes a feature corresponding to the current sub-bandf=1in an input feature vector corresponding to VM2 based on an initial hidden status of VM2 at the current sub-bandf=1, to obtain a processed feature corresponding to the current sub-bandf=1and a processed hidden status corresponding to VM2. In a similar way, the first feature processing module may process a feature corresponding to the current sub-bandf=1in an input feature vector corresponding to each VM in a similar way (if there is no feature corresponding to the current sub-bandf=1in the feature vector corresponding to a certain VM, the processing is not performed), so that a processed feature and a processed hidden status (wherein i represent a serial number of VM)corresponding to each VM may be obtained. Then, the feature processing is performed on the processed hidden status corresponding to each VM by using the first fusion module 1940 (e.g. fusion module 2010) of the third neural network to obtain the global hidden status . For example, the feature processing performed by the first fusion module may be a convolutional operation performed through CNN, but the present application is not limited hereto, and it may also be a feature processing performed by a network with an attention mechanism. Hereafter, a feature processing may be performed on a feature vector corresponding to a next sub-bandf=2(i.e.,f2) in the VM position feature vector and the global hidden status by using the first update module 1970 (e.g. update module 2020) of the third neural network to obtain an initial hidden status of each VM at the next sub-bandf=2. In FIG. 20, there are only VM1 and VM4 at the next sub-bandf=2, thus, initial hidden status of VM1 and initial hidden status of VM4 at the next sub-bandf=2may be obtained. Then, the feature processing may be performed for the sub-bandf=2in a similar way after the feature processing is performed for the sub-bandf=1. For example, the first feature processing module 1910 may process features corresponding to the current sub-bandf=2in the input feature vector corresponding to VM1 and the input feature vector corresponding to VM4 based on the initial hidden statuses and , so that processed features corresponding to the current sub-bandf=2, a processed hidden status corresponding to VM1 and a processed hidden status corresponding to VM4 may be obtained. Processing the feature vector corresponding to each VM in the above way may realize normal transfer of the hidden statuses during the network feature processing.

[0132] The feature vector corresponding to each VM processed in the frequency direction may be obtained through the above process as described above with reference to FIG. 20. Then, as shown in FIG. 19, the obtained feature vector corresponding to each VM processed in the frequency direction is input to the second feature processing module 1920, and the feature processing is performed in the time direction then. This will be described in detail with reference to FIGS. 21 and 19 below.

[0133] FIG. 21 is a schematic diagram illustrating a process of performing feature processing for each frequency unit in a time direction according to an exemplary embodiment of the present application.

[0134] As shown in FIGS. 19 and 21, when processing a feature vector corresponding to RM, for each frame (such as a frame is 16ms), a feature corresponding to a current frame in an input feature vector corresponding to each RM based on an initial hidden status of each RM at the current frame, to obtain a processed feature and a processed hidden status corresponding to each RM, wherein the processed hidden status corresponding to each RM is used as an initial hidden status of each RM at a next frame.

[0135] For example, as shown in FIG. 21, for RM1, the second feature processing module 1920 of the third neural network processes a feature corresponding to a first frame in an input feature vector corresponding to RM1 based on an initial hidden status at the first frame, to obtain a processed feature and an updated hidden status . The updated hidden status is used as an initial hidden status of RM1 at a next frame.

[0136] While processing a feature vector corresponding to VM, as shown in FIGS. 19 and 21, for each frame, following operations are performed: processing a feature corresponding to the current frame in an input feature vector corresponding to each VM based on an initial hidden status of each VM at the current frame, to obtain a processed feature and a processed hidden status corresponding to each VM; obtaining a global hidden status by performing feature processing on the processed hidden status corresponding to each VM; and obtaining an initial hidden status of each VM at the next frame by performing a feature processing on a feature vector corresponding to the next frame in the VM position feature vector and the global hidden status.

[0137] As shown in FIG. 21, for a first frame, the second feature processing module 1920 processes a feature corresponding to the first frame in an input feature vector corresponding to VM2 based on an initial hidden status of VM2 at the first frame, to obtain a processed feature and a processed hidden status corresponding to VM2. In a similar way, the second feature processing module 1920 process a feature corresponding to the first frame in an input feature vector corresponding to each VM (if there is no feature corresponding to the first frame in the feature vector corresponding to a certain VM, the processing is not performed), so that a processed feature and a processed hidden status (wherein i represents a serial number of VM) corresponding to each VM may be obtained. For example, as shown in FIG. 21, since there are only the features corresponding to the first frame in the feature vector corresponding to VM2, a feature vector corresponding to VM5 and a feature vector corresponding to VM6, the hidden statuses obtained in the current step are only a hidden status corresponding to VM2, a hidden status corresponding to VM5 and a hidden status corresponding to VM6. Then, the feature processing is performed on the processed hidden statuses , and by using the second fusion module 1950 (e.g. fusion module 2110) of the third neural network to obtain a global hidden status . For example, the feature processing performed by the second fusion module 1950 may be a convolutional operation performed through CNN, but the present application is not limited hereto, and it may also be a feature processing performed by a network with an attention mechanism. Hereafter, the feature processing is performed on a feature vector corresponding to a next frame (i.e., a second frame) in the VM position feature vector and the global hidden status by using the second update module 1980 (e.g. update module 2120) of the third neural network to obtain an initial hidden status of each VM at the second frame. In FIG. 21, there are only VM1 and VM4 at the second frame, thus, initial hidden statuses and of VM1 and VM4 at the second frame may be obtained through the above update operation performed by the second update module 1980. Then, the feature processing may be performed for the second frame in a similar way after the feature processing is performed for the first frame. For example, the second feature processing module 1920 may process features corresponding to the second frame in the input feature vector corresponding to VM1 and the input feature vector corresponding to VM4 based on the initial hidden statuses and , so as to obtain processed features corresponding to the second frame, a processed hidden status corresponding to VM1 and a processed hidden status corresponding to VM4. Processing the feature vector corresponding to each VM in the above way may realize normal transfer of the hidden statuses during the network feature processing.

[0138] The feature vector corresponding to each VM processed in the time direction may be obtained through the above process as described with reference to FIG. 21. Then, as shown in FIG. 19, the obtained feature vector corresponding to each VM processed in the time direction is input to the third feature processing module 1930, and the feature processing is performed in the space direction then. This will be described in detail with reference to FIGS. 22 and 19 below.

[0139] FIG. 22 is a schematic diagram illustrating a process of performing feature processing for each time unit in a space direction according to an exemplary embodiment of the present application.

[0140] As shown in FIGS. 19 and 22, while processing a feature vector corresponding to RM, for each RM, following operations are performed: processing a feature corresponding to each sub-band in an input feature vector corresponding to a current RM based on an initial hidden status of the current RM at each sub-band, to obtain a processed feature and a processed hidden status corresponding to each sub-band, wherein the processed hidden status corresponding to each sub-band is used as an initial hidden status of each sub-band at a next RM.

[0141] For example, as shown in FIG. 22, for RM1, the third feature processing module 1930 of the third neural network processes a feature corresponding to each sub-band in an input feature vector corresponding to RM1 based on an initial hidden status of RM1 at each sub-band, to obtain a processed feature and a processed hidden status corresponding to each sub-band, wherein the processed hidden status corresponding to each sub-band is used as an initial hidden status of each sub-band at RM1. The feature processing is performed for RM2 and RM3 in a similarly way.

[0142] While processing a feature vector corresponding to VM, as shown in FIGS. 19 and 22, for each VM, following operations are performed: processing a feature corresponding to each sub-band in an input feature vector corresponding to a current VM based on an initial hidden status of the current VM at each sub-band, to obtain a processed feature and a processed hidden status corresponding to each sub-band; obtaining a global hidden status by performing feature processing on the processed hidden status corresponding to each sub-band; and obtaining an initial hidden status of a next VM at each sub-band by performing feature processing on a feature vector corresponding to a next VM in the VM position feature vector and the global hidden status.

[0143] As shown in FIG. 22, for VM1, the third feature processing module 1930 processes features corresponding to a sub-bandf=2and a sub-bandf=3in an input feature vector corresponding to VM1 based on initial hidden statuses and of VM1 at the sub-bandf=2and the sub-bandf=3, to obtain processed features and processed hidden statuses and corresponding to the sub-bandf=2and the sub-bandf=3(if there is no feature corresponding to a certain sub-band in the feature vector corresponding to a certain VM1, the processing is not performed for the certain sub-band). Then, the feature processing is performed on the processed hidden statuses and by using the third fusion module 1960 (e.g. fusion module 2210) of the third neural network to obtain a global hidden status . For example, the feature processing performed by the third fusion module 1960 may be a convolutional operation performed through CNN, but the present application is not limited hereto, and it may also be a feature processing performed by a network with an attention mechanism. Hereafter, feature processing is performed on a feature vector corresponding to a next VM (i.e., VM2) in the VM position feature vector and the global hidden status by using the third update module 1990 (e.g. update module 2220) of the third neural network to obtain an initial hidden status of VM2 at each sub-band. In FIG. 22, there are only a feature corresponding to the sub-bandf=1, a feature corresponding to a sub-bandf=4and a feature corresponding to a sub-bandf=5in a feature vector corresponding to VM2 at time t, thus, initial hidden statuses , and of VM2 at the sub-bandf=1, the sub-bandf=4and the sub-bandf=5may be obtained through the above update operation performed by the third update module 1990. Then, the feature processing may be performed for VM2 in a similar way after the feature processing is performed for VM1. For example, the third feature processing module 1930 may process a feature corresponding to the sub-bandf=1, a feature corresponding to the sub-bandf=4and a feature corresponding to the sub-bandf=5in the input feature vector corresponding to VM2 based on the initial hidden statuses , and , so that processed features corresponding to VM2 and a processed hidden status corresponding to the sub-bandf=1, a processed hidden status corresponding to the sub-bandf=4and a processed hidden status corresponding to the sub-bandf=5may be obtained. Processing the feature vector corresponding to each VM in the above way may realize normal transfer of the hidden statuses during the network feature processing.

[0144] Finally, the mask related to the second audio signal inside the target image area may be obtained by performing the feature processing in the frequency direction, the time direction and the space direction as described above, wherein the mask may represent an information percentage of second audio signal (i.e., the target audio signal) occupied in the first audio signal.

[0145] The second audio signal may be obtained based on the mask and the first feature vector after the mask is obtained. This operation may be performed by a decoder module 540 in FIG. 5. Specifically, the obtaining the second audio signal based on the mask and the first feature vector includes: obtaining a fifth feature vector based on the mask and the first feature vector; and obtaining the second audio signal by performing an audio signal restoration operation on the fifth feature vector.

[0146] In an exemplary embodiment of the present application, if the encoder module 530 in FIG. 5 directly encodes the extracted feature vector to obtain the first feature vector, the decoder module 540 may directly perform a dot product operation on the mask and the first feature vector to obtain the fifth feature vector, then decode the fifth feature vector, so as to restore the target audio signal, i.e., the second audio signal.

[0147] In another exemplary embodiment of the present application, if the encoder module 530 in FIG. 5 uses the structure shown in FIG. 15, it divides the first feature vector into a plurality of first sub-band feature vectors, encodes the divided plurality of first sub-band feature vectors using a plurality of sub encoders, so as to obtain the first feature vector. Then, the decoding may be realized by correspondingly using a plurality of sub-decoders at the decoder module 540. This will be described in detail with reference to FIGS. 23 and 24 below.

[0148] FIG. 23 is a block diagram illustrating a decoder module according to an exemplary embodiment of the present application. FIG. 24 is a flowchart illustrating a network of a decoder module according to an exemplary embodiment of the present application.

[0149] As shown in FIG. 23, the decoder module 2300 includes a plurality of sub-decoders 2310, a feature merge module 2320 and a time domain signal restoration module 2330.

[0150] In an example illustrated in FIG. 24, the obtained first feature vector includes four second sub-band feature vectors , , and . In this case, an ithsub-band mask in the mask and an ith second sub-band feature vector in the first feature vector are processed (for example, a dot product operation of corresponding elements is performed) through an ith sub-decoder to obtain an iththird sub-band feature vector, and then feature conversion is performed on the four obtained third sub-band feature vectors to obtain four prediction features , , and in operation S2410, wherein the feature conversion operation may be realized through a linear fully connected layer, but the present application is not limited hereto, and it may also be realized through networks (such as CNNs or the like) capable of performing feature conversion operation. Hereafter, in operation S2420, feature merge (which may also be understood as frequency band merge) is performed on the four prediction features , , and through the feature merge module to obtain a merged feature , i.e., the fifth feature vector, for facilitating subsequent feature conversion operation.

[0151] That is to say, the obtaining the fifth feature vector based on the mask and the first feature vector includes: obtaining a plurality of third sub-band feature vectors by processing each sub-band mask in the mask and a corresponding second sub-band feature vector in the first feature vector; obtaining a plurality of prediction features by performing feature conversion on the plurality of third sub-band feature vectors; and obtaining the fifth feature vector by performing feature merge on the plurality of prediction features.

[0152] After the fifth feature vector is obtained, in operation S2430, an audio signal restoration operation may be performed on the fifth feature vector through the time domain signal restoration module 2330 in FIG. 23 to obtain the target audio signal, i.e., the second audio signal. For example, the time domain signal restoration module 2330 may use short-term inverse Fourier transform to perform the audio signal restoration operation, but the present application is not limited hereto, and the time domain signal restoration module 2330 may also use other feature transformation methods, for example, CNN networks may be used to perform the audio signal restoration operation. In the present application, if the encoder module uses the short-term Fourier transform to perform feature extraction, then in the decoder module, the time domain signal restoration module may use the short-term inverse Fourier transform to realize the audio signal restoration operation on the fifth feature vector. Accordingly, if the encoder module uses the CNN network to perform feature extraction, then in the decoder module, the time domain signal restoration module may use the CNN network to realize the audio signal restoration operation on the fifth feature vector.

[0153] A plurality of neural networks (i.e., the first neural network, the second neural network and the third neural network) are used in the method of extracting the target audio signal (second audio signal) that the user wants from the input mixed audio signal (i.e., the first audio signal) as described above with reference to the accompanying drawings, and they may be obtained through offline training. This will be described in detail below.

[0154] FIG. 25 is a schematic diagram illustrating a process of training a sound extraction module (the third neural network) according to an exemplary embodiment of the present application. In brief, the sound extraction module is trained using simulation features generated by a simulator.

[0155] At first, positions of VMs are obtained by using Expert Guidance according to the positions of the target audio signal and the interference signal. Then, feature vectors simulate_vm of VMs are obtained through simulation of the simulator 2510 based on these positions of VMs. Besides, signals of RMs (such as a signal of RM1, a signal of RM2 and a signal of RM3) are encoded using an encoder module 2520 to obtain feature vectors of RMs. Hereafter, the obtained feature vectors simulate_vm of VMs and the obtained feature vectors of RMs are input to the third neural network (i.e., the sound extraction module 2530) for feature processing so as to obtain a mask. After the dot product operation is performed on the mask and the feature vectors of RMs, a result of the dot product operation is input to the decoder module 2540 for feature decoding so as to extract an audio signal. A maximum signal to noise ration (SNR) 2550 is calculated as a target function based on the extracted audio signal and the target audio signal, and the third neural network is iteratively trained by adjusting model parameters of the third neural network until the target function SNR 2550 obtained through the above process converges.

[0156] After the trained third neural network is obtained through the above training process, then the trained third neural network and the VM analysis and feature generation module (i.e., the first neural network and the second neural network) are jointly trained. This will be described in detail with reference to FIG. 26 below.

[0157] FIG. 26 is a schematic diagram illustrating a process of jointly training the third neural network and the VM analysis and feature generation module (wherein the VM analysis and feature generation module includes the first neural network and the second neural network) according to an exemplary embodiment of the present application.

[0158] Specifically, signals of RMs (such as a signal of RM1, a signal of RM2 and a signal of RM3) are encoded using an encoder module 2630 to obtain a feature vector (i.e., a first feature vector) of RMs. In addition, the obtained feature vector of RMs, DOA information, target direction information (i.e., DOI information), and spatial information of RMs are input to the VM analysis and feature generation module 2610 to obtain the feature vector of VMs (i.e., the second feature vector). In addition, the information input to the VM analysis and feature generation module 2610 further includes at least one of status information of an electronic apparatus, historical statistical information, and directivity of RM. Specifically, for each frequency unit at each time unit, following operations are performed: obtaining a first signal spatial distribution feature based on the DOA information and the target direction information; and determining the number and positions of the VMs based on the first signal spatial distribution feature and the spatial information of the RMs, wherein the number and positions of the VMs obtained for each frequency unit at each time unit constitute a VM position feature vector. Hereafter, relative position information between VMs and RMs is determined based on the VM position feature vector and the spatial information of the RMs, directivity of each VM is determined based on the relative position information between VMs and RMs and the directivity of RM, and the feature vector of VM is obtained based on the relative position information between VMs and RMs, the feature vector of RM and the directivity of each VM. Then, the feature vector of VM and the feature vector of RM are input to the third neural network (i.e., the sound extraction module 2640) to obtain the mask. Then, after the dot product operation is performed on the mask and the feature vector of RM, a result of the dot product operation is input to the decoder module 2650 for feature decoding so as to extract an audio signal. During training, a SNR 2660 is calculated based on the extracted audio signal and the target audio signal as a first target function, a minimum mean square error (MSE) 2620 is calculated based on the feature vector of VM obtained through the VM analysis and feature generation module 2610 and the feature vector simulate_vm of VM obtained through simulation of the simulator as a second target function, and these neural networks are iteratively trained by adjusting model parameters of the first neural network and the second neural network and by slightly adjusting model parameters of the third neural network simultaneously until the first target function and the second target function obtained through the above process converge.

[0159] The method of extracting the target audio signal (second audio signal) that the user wants from the input mixed audio signal (i.e., the first audio signal) is described above with reference to the accompanying drawings. This method may be applied to a variety of scenarios that require voice focus, so that video or audio recorders may clearly pick up sounds the user is interested in (such as distant sounds or information in a specific direction, or the like). Three application scenarios of the method performed by the electronic apparatus as described above in the application are described through following examples, but the actual use scenarios are not limited to these three scenarios.

[0160] FIG. 27 is a schematic diagram illustrating application of a method performed by an electronic apparatus according to the present application to a scenario of recording a video. In the example, users can choose to record sound objects of interest and control a range of recording the sound objects.

[0161] At first, a user turns on an audio zoom function in a video mode. Then, the user performs a touch operation (such as a swipe operation) on a screen 2710 to select an area of interest AOI (which may also be referred to as a target image area) on the screen 2710. Hereafter, the method performed by the electronic apparatus may obtain target direction information (i.e., a direction and width of the area of interest AOI), which is also referred to as DOI information. Specifically, coordinate values of the area of interest AOI selected by the user are acquired through a video processing module, the width of the area of interest AOI is calculated according to the coordinate values and a zoom state, and then the direction of the area of interest AOI is calculated according to the current zoom state and an angle offset relative to a front (e.g. 45°). On this basis, the target sound (i.e., the second audio signal) inside the area of interest AOI may be extracted through the method performed by the electronic apparatus proposed by the present application. In the example, the user may accurately select the area of interest AOI, and may accurately extract the target sound in the area for enhancement and other processing. For example, the user may accurately extract the target sound of a first person 2720 or a second person 2730 in the area of interest AOI, not the interfering sound of a third person 2730 outside the area of interest AOI. That is to say, the above-described method performed by the electronic apparatus may be not only applicable to sound extraction, but also applicable to sound enhancement and sound separation.

[0162] In addition, the method performed by the electronic apparatus according to the present application may also be applied to a scenario of recording videos. For example, users may record voices of singers at concerts by zooming a screen, or record music of street performers in noisy environments by zooming the screen.

[0163] At first, a user turns on an audio zoom function in a video mode. Hereafter, the user zooms in or out the screen (that is, zooming in or out an audio pickup range). Then, the method performed by the electronic apparatus may obtain target direction information (i.e., a direction and width of the area of interest). In this case, the direction of the area of interest may default to straight ahead (that is, the front). On this basis, the target sound (i.e., the second audio signal) inside the area of interest may be extracted through the method performed by the electronic apparatus proposed by the present application so as to be used as audio signals when the user records the video. In the example, by zooming in or out the screen, the method performed by the electronic device proposed by the present application may obtain the target direction information, so that the sound inside the screen may be accurately extracted while eliminating interference outside the screen.

[0164] In addition, the method performed by the electronic apparatus according to the present application may also be applied to a scenario of recording a voice memo or calling. For example, when a user makes a phone call voice memo, a posture of holding a mobile phone may vary, and a relative position of the user's mouth and the mobile phone may be different. By applying the method performed by the electronic apparatus proposed by the present application, the user's voice may be accurately extracted following changes in the relative position between the user's mouth and the mobile phone, so that the users may also record their own voice while blocking background noises even if in a noisy environment.

[0165] Specifically, at first, a user turns on an audio zoom function in a video mode. Then, target direction information is obtained, and in this case, a target direction is a direction of the user's voice relative to the mobile phone. Hereinafter, through the method performed by the electronic apparatus proposed by the present application, the user's voice (i.e., the second audio signal) may be extracted and surrounding sounds may be suppressed.

[0166] In embodiments of the present disclosure, there is also provided an electronic apparatus that includes at least one processor, and alternatively, further includes at least one transceiver and / or at least one memory coupled to the at least one processor, wherein, the at least one processor is configured to perform the steps of the method provided in any alternative embodiment of the present disclosure.

[0167] FIG. 28 illustrates a schematic diagram of a structure of an electronic apparatus applicable to an exemplary embodiment of the present application. As shown in FIG. 28, the electronic apparatus 4000 shown in FIG. 28 includes: a processor 4001 and a memory 4003, wherein the processor 4001 and the memory 4003 are coupled, e.g., through a bus 4002. Alternatively, the electronic apparatus 4000 may further include a transceiver 4004 which may be used for data interaction between the electronic apparatus and other electronic apparatuses, such as transmitting of data and / or receiving of data. It should be noted that, each of the processor 4001, the memory 4003, and the transceiver 4004 is not limited to one in a practice application, and the structure of the electronic apparatus 4000 does not constitute a limitation of the embodiments of the present disclosure. Alternatively, the electronic apparatus may be the first network node, the second network node, or the third network node.

[0168] The processor 4001 may be a Central Processing Unit (CPU), general purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, transistor logic device, hardware part, or any combination thereof. It may implement or perform various exemplary logic boxes, modules, and circuits described in conjunction with the disclosed contents of the present disclosure. The processor 4001 may also be a combination that implements computing functions, such as a combination containing one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0169] The bus 4002 may include a pathway to transfer information between the above components. The bus 4002 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, and the like. The bus 4002 may be classed as an address bus, a data bus, a control bus, and the like. For ease of representation, only one bold line is shown in FIG. 28, but it does not mean that there is only one bus or one type of bus.

[0170] The memory 4003 may be a Read Only Memory (ROM) or other types of static storage apparatuses that can store static information and instructions, a Random Access Memory (RAM) or other types of dynamic storage apparatuses that can store information and instructions, may be an Electrically Erasable Programmable Read Only Memory (EEPROM), Compact Disc Read Only Memory (CD-ROM) or other optical disc storages, an optical disc storage (including a compressed disc, laser disc, optical disc, digital universal disc, Blu-ray disc, etc.), a disk storage medium, other magnetic storage apparatuses, or any other medium that can be used to carry or store computer programs and can be read by a computer, it is not limited herein.

[0171] The memory 4003 is used to store computer programs or executable instructions for performing the embodiments of the present disclosure, and is controlled for execution by the processor 4001. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the steps shown in the preceding method of the embodiments.

[0172] An embodiment of the present disclosure provides a computer readable storage medium storing computer programs or instructions, and the computer programs or instructions, when being executed by at least one processor may perform or implement the steps in the preceding method of the embodiments and corresponding contents.

[0173] An embodiment of the present disclosure provides a computer program product including computer programs, and the computer programs, when being executed by a processor, may implement the steps shown in the preceding method of the embodiments and corresponding contents.

[0174] The terms "first", "second", "third", "fourth", "1", "2" and the like (if exists) in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequence. It should be understood that, data used as such may be interchanged in appropriate situations, so that the embodiments of the present disclosure described here may be implemented in an order other than the illustration or text description.

[0175] It should be understood that, although each operation step is indicated by an arrow in the flowcharts of the embodiments of the present disclosure, an implementation order of these steps is not limited to an order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in the flowcharts may be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include a plurality of sub steps or stages based on an actual implementation scenario. Some or all of these sub steps or stages may be executed at the same time, and each sub step or stage in these sub steps or stages may also be executed at different times. In scenarios with different execution times, an execution order of these sub steps or stages may be flexibly configured according to a requirement, which is not limited by the embodiment of the present disclosure.

[0176] The text and drawings are provided as examples only to assist readers in understanding the present disclosure. They are not intended and should not be interpreted as limiting the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, based on the content disclosed herein, it is apparent to those skilled in the art that, changes can be made to the illustrated embodiments and examples without departing from the scope of the present disclosure, and other similar implementation methods based on the technical concepts of the present disclosure also belong to a protection scope of the embodiments of the present disclosure.

Claims

1.A method performed by an electronic apparatus, comprising:obtaining an input first audio signal;obtaining a number and positions of virtual microphones (VMs) based on direction of arrival (DOA) information of the first audio signal, target direction information, and spatial information of real microphones (RMs) of the electronic apparatus; andobtaining a second audio signal based on the first audio signal and the number and positions of the VMs.2.The method of claim 1, wherein the obtaining the number and positions of the VMs based on the DOA information of the first audio signal, the target direction information, and the spatial information of the RMs of the electronic apparatus comprises: for each frequency unit at each time unit, performing following operations:obtaining a first signal spatial distribution feature based on the DOA information and the target direction information; anddetermining the number and positions of the VMs based on the first signal spatial distribution feature and the spatial information of the RMs.3.The method of claim 2, wherein the obtaining the first signal spatial distribution feature based on the DOA information and the target direction information comprises:obtaining a third feature vector by feature mapping the DOA information;obtaining a fourth feature vector by encoding the target direction information;obtaining a second signal spatial distribution feature based on the third feature vector and the fourth feature vector; andobtaining the first signal spatial distribution feature based on historical statistical information of the first audio signal and the second signal spatial distribution feature.4.The method of claim 3, wherein the historical statistical information comprises voice signal feature information and noise signal feature information of a last frame in a previous time unit,and wherein the obtaining the first signal spatial distribution feature based on the historical statistical information of the first audio signal and the second signal spatial distribution feature comprises:obtaining the first signal spatial distribution feature, by calculating correlation of audio signals inside a target image area for each RM based on the voice signal feature information of the last frame in the previous time unit and the second signal spatial distribution feature, and calculating correlation of audio signals outside the target image area for each RM based on the noise signal feature information of the last frame in the previous time unit and the second signal spatial distribution feature.5.The method of claim 2, wherein the determining the number and positions of the VMs based on the first signal spatial distribution feature and the spatial information of the RMs comprises:determining the number and positions of the VMs based on the first signal spatial distribution feature, the spatial information of the RMs and status information of the electronic apparatus.6.The method of claim 5, wherein the determining the number and positions of the VMs based on the first signal spatial distribution feature, the spatial information of the RMs and the status information of the electronic apparatus comprises:encoding the status information of the electronic apparatus into a status vector;determining a microphone aperture according to the spatial information of the RMs; anddetermining the number and positions of the VMs based on the first signal spatial distribution feature, the microphone aperture and the status vector.7.The method of claim 2, wherein the number and positions of the VMs determined for each frequency unit at each time unit constitute a three-dimensional (3D) VM position feature vector in a time direction, a frequency direction and a space direction, andwherein the obtaining the second audio signal based on the first audio signal and the number and positions of the VMs comprises:obtaining a first feature vector corresponding to the first audio signal; andobtaining the second audio signal based on the first feature vector, the VM position feature vector and the spatial information of the RMs.8.The method of claim 7, wherein the obtaining the second audio signal based on the first feature vector, the VM position feature vector and the spatial information of the RMs comprises:determining a relative position information between the VMs and the RMs based on the VM position feature vector and the spatial information of the RMs;determining a directivity of each VM based on the relative position information between the VMs and the RMs and directivity of the RMs;obtaining the second feature vector based on the relative position information between the VMs and the RMs, the first feature vector and the directivity of each VM; andobtaining the second audio signal based on the first feature vector and the second feature vector.9.The method of claim 8, wherein the obtaining the second audio signal based on the first feature vector and the second feature vector comprises:obtaining a mask related to the second audio signal inside a target image area based on the first feature vector and the second feature vector; andobtaining the second audio signal based on the mask and the first feature vector.10.The method of claim 9, wherein the obtaining the mask related to the second audio signal inside the target image area based on the first feature vector and the second feature vector comprises:obtaining a merged feature vector by merging the first feature vector and the second feature vector along a space direction; andobtaining the mask by performing feature processing on the merged feature vector in a first direction, a second direction and a third direction in sequence,wherein the first direction, the second direction and the third direction are one of the space direction, a time direction and a frequency direction, respectively.11.The method of any one of claims 2 to 5 and 7 to 10, wherein the frequency unit is a frequency point or a sub-band, wherein the sub-band comprises a plurality of frequency points.12.The method of any one of claims 2 to 5 and 7 to 10, wherein the time unit is one or more frames.13.The method of claim 9, wherein the obtaining the second audio signal based on the mask and the first feature vector comprises:obtaining a fifth feature vector based on the mask and the first feature vector; andobtaining the second audio signal by performing an audio signal restoration operation on the fifth feature vector.14.An electronic apparatus, comprising:at least one processor; andat least one memory storing computer executable instructions,wherein the computer executable instructions, when executed by the at least one processor, cause the at least one processor to perform the method of any one of claims 1 to 13.15.A computer readable storage medium storing instructions, wherein the instructions, when executed by at least one processor, cause the at least one processor to perform the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Apparatus and Method To Track Position For Multiple Sound Source

    KR101612704B1

  • Sound acquisition via the extraction of geometrical information from direction of arrival estimates

    KR1020140045910A

  • Method and apparatus for providing augmented contents through augmented reality view based on preset unit space

    KR1020240083090A

  • Additive composition for semiconductor polishing process, polishing slurry composition and method for manufacturing semiconductor devices

    KR1020260024804A

  • Systems and methods for audio adjustment

    US20210321213A1

Cited By

  • Sound detection method

    CN122116940A