Audio acquisition method and device, electronic equipment and computer readable medium

By setting a center point and multiple audio acquisition components in an electronic device, and combining rotation angle and vertical distance, the audio acquisition model is used to identify audio signals in the pickup area, solving the problem of high cost of audio acquisition in the microphone pickup area and realizing convenient and efficient audio acquisition.

CN121985251APending Publication Date: 2026-05-05VOICEAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VOICEAI TECH CO LTD
Filing Date
2025-12-22
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies for acquiring audio signals from the microphone pickup area are costly and inconvenient.

Method used

By setting a center point in an electronic device and multiple audio acquisition components surrounding the center point, sampled audio signals, rotation angles, and vertical distances are acquired. Audio recognition is then performed using an audio acquisition model to identify the target audio signal in the pickup area.

Benefits of technology

It reduces the cost of audio acquisition, improves convenience, and allows audio acquisition in the pickup area to be completed without the need for additional hardware equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985251A_ABST
    Figure CN121985251A_ABST
Patent Text Reader

Abstract

The invention discloses an audio acquisition method and device, electronic equipment and a computer readable medium, the audio acquisition method and device are applied to the electronic equipment, the electronic equipment comprises an audio acquisition module, the audio acquisition module comprises a central point and a plurality of audio acquisition assemblies arranged around the central point, and the distances between the audio acquisition assemblies and the central point are the same. The reference lines of the acquisition areas corresponding to the audio acquisition assemblies are parallel to each other. The method comprises the following steps: acquiring a sampling audio signal acquired by each audio acquisition assembly; obtaining a rotation angle and a vertical distance of each audio acquisition assembly; and carrying out audio identification on the sampled audio signal acquired by each audio acquisition assembly, the rotation angle of the audio acquisition assembly and the vertical distance through the audio acquisition model to obtain a target audio signal of the pickup area corresponding to the audio acquisition module. Acquisition of audios in the pickup area can be completed without the help of additional hardware equipment. The use cost is reduced, and the convenience of audio acquisition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio technology, and more specifically, to an audio acquisition method, apparatus, electronic device, and computer-readable medium. Background Technology

[0002] Currently, with the development of electronic information technology, it is possible to collect audio signals from the microphone's pickup area. However, the current method of collecting audio signals from the pickup area is costly and inconvenient. Summary of the Invention

[0003] This application proposes an audio acquisition method, apparatus, electronic device, and computer-readable medium to improve the above-mentioned deficiencies.

[0004] In a first aspect, embodiments of this application provide an audio acquisition method applied to an electronic device. The electronic device includes an audio acquisition module, which includes a center point and a plurality of audio acquisition components surrounding the center point. Each audio acquisition component is equidistant from the center point, and reference lines for the acquisition areas corresponding to each audio acquisition component are parallel to each other. The method includes: acquiring sampled audio signals acquired by each audio acquisition component; acquiring the rotation angle and vertical distance of each audio acquisition component, wherein the rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component to the center point and the reference direction, the reference direction includes the direction from the center point to a reference point on the circumference formed by the plurality of audio acquisition components, and the vertical distance is used to characterize the vertical length of the audio acquisition component from the center point; and performing audio recognition on the sampled audio signals acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance using an audio acquisition model to obtain the target audio signal of the pickup area corresponding to the audio acquisition module.

[0005] Optionally, in some embodiments, before performing audio recognition on the sampled audio signals acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance using an audio acquisition model to obtain the target audio signal of the pickup area corresponding to the audio acquisition module, the method further includes: acquiring a training dataset and a reference dataset, wherein the training dataset comprises multiple sets of training data, each set of training data including the training audio signal received by each audio acquisition component, the training rotation angle, and the training vertical distance; the reference dataset includes reference data corresponding to each set of training data, and each set of reference data includes a reference audio signal corresponding to each audio acquisition component; performing audio recognition on each set of training data using an initial model to obtain an initial audio signal corresponding to each set of training data; training the initial model based on a first difference between the initial audio signal and the reference audio signal corresponding to the set of training data to reduce the first difference; and using the trained initial model as the audio acquisition model.

[0006] Optionally, in some embodiments, when the sound source is within the pickup area, the reference audio signal is the audio signal emitted by the sound source; when the sound source is not within the pickup area, the reference audio signal is a silence audio signal.

[0007] Optionally, in some embodiments, the reference data further includes a reference label, which is used to characterize whether the sound source is located in the pickup area. The step of performing audio recognition on each set of training data through an initial model to obtain an initial audio signal corresponding to each set of training data includes: performing audio recognition on each set of training data through an initial model to obtain an initial audio signal and an initial label corresponding to each set of training data; training the initial model based on a first difference between the initial audio signal and the reference audio signal corresponding to the set of training data to reduce the first difference includes: training the initial model based on the first difference and a second difference to reduce the first difference and the second difference, wherein the second difference is used to characterize the difference between the initial label and the reference label corresponding to the set of training data.

[0008] Optionally, in some embodiments, the step of performing audio recognition on each set of training data using an initial model to obtain the initial audio signal corresponding to each set of training data includes: acquiring the real spectrum and imaginary spectrum corresponding to the training audio signal received by each of the audio acquisition components; determining the amplitude spectrum feature vector of the training data based on the real spectrum and imaginary spectrum corresponding to each set of training audio signals; calculating the phase difference feature vector between the channels corresponding to each audio acquisition component in the training data based on the real spectrum and imaginary spectrum corresponding to each set of training audio signals; constructing the angular distance feature vector corresponding to the training data based on the training rotation angle and training vertical distance corresponding to each of the audio acquisition components; and calculating the initial audio signal corresponding to the training data based on the real spectrum, imaginary spectrum, amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector.

[0009] Optionally, in some embodiments, determining the amplitude spectrum feature vector of the training data based on the real and imaginary spectra corresponding to each group of training audio signals includes: constructing an initial amplitude spectrum feature vector of the training data based on the square root of the sum of the squares of the real and imaginary spectra corresponding to each group of training audio signals; and performing amplitude spectrum compression on the initial amplitude spectrum feature vector to obtain the amplitude spectrum feature vector of the training data.

[0010] Optionally, in some embodiments, the step of calculating the phase difference feature vector between channels corresponding to each audio acquisition component in the training data based on the real and imaginary spectra corresponding to each group of training audio signals includes: calculating the initial phase spectrum of the training data based on the real and imaginary spectra corresponding to each group of training audio signals; and determining the phase difference feature vector between channels corresponding to each audio acquisition component based on the initial phase spectrum.

[0011] Optionally, in some embodiments, determining the phase difference feature vector between the channels corresponding to each audio acquisition component based on the initial phase spectrum includes: using the initial phase spectrum to represent the initial phase difference feature vector between the channels corresponding to each audio acquisition component in a triangular manner; and performing phase encoding processing on the initial phase difference feature vector to obtain the phase difference feature vector between the channels corresponding to each audio acquisition component.

[0012] Optionally, in some embodiments, constructing the angular distance feature vector corresponding to the set of training data based on the training rotation angle and training vertical distance corresponding to each of the audio acquisition components includes: concatenating the training rotation angles corresponding to each of the audio acquisition components to obtain a rotation angle feature vector; concatenating the training vertical distances corresponding to each of the audio acquisition components to obtain a vertical distance feature vector; and concatenating the rotation angle feature vector and the vertical distance feature vector to obtain the angular distance feature vector corresponding to the set of training data.

[0013] Optionally, in some embodiments, calculating the initial audio signal corresponding to the set of training data based on the real spectrum, imaginary spectrum, amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector includes: concatenating the amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector to obtain a first mixed feature vector; performing encoding and decoding processing on the first mixed feature vector to obtain mask information corresponding to each of the audio acquisition components; performing noise reduction processing on the corresponding real spectrum and imaginary spectrum based on the mask information to obtain a denoised real spectrum and a denoised imaginary spectrum, respectively; and calculating the initial audio signal corresponding to the set of training data based on the denoised real spectrum and the denoised imaginary spectrum.

[0014] Optionally, in some embodiments, the step of calculating the initial audio signal corresponding to the set of training data based on the denoised real spectrum and the denoised imaginary spectrum includes: summing and averaging the denoised real spectra of multiple channels in the set of training data through a pooling layer to obtain a single-channel real spectrum; summing and averaging the denoised imaginary spectra of multiple channels in the set of training data through a pooling layer to obtain a single-channel imaginary spectrum; and performing an inverse Fourier transform based on the single-channel real spectrum and the single-channel imaginary spectrum to obtain the initial audio signal corresponding to the set of training data.

[0015] Optionally, in some embodiments, obtaining the real spectrum and imaginary spectrum corresponding to the training audio signal received by each of the audio acquisition components includes: performing a Fourier transform on each group of training audio signals to obtain a first formula; transforming the first formula using Euler's formula to obtain a second formula; using the real part of the second formula as the real spectrum corresponding to the training audio signal, and using the imaginary part of the second formula as the imaginary part corresponding to the training audio signal.

[0016] Secondly, embodiments of this application also provide an audio acquisition device applied to an electronic device. The electronic device includes an audio acquisition module, which includes a center point and a plurality of audio acquisition components surrounding the center point. Each audio acquisition component is equidistant from the center point, and the reference lines of the acquisition areas corresponding to each audio acquisition component are parallel to each other. The device includes: a first acquisition unit, used to acquire sampled audio signals acquired by each audio acquisition component; a second acquisition unit, used to acquire the rotation angle and vertical distance of each audio acquisition component, wherein the rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component and the center point and the reference direction, the reference direction including the direction from the center point to a reference point on the circumference formed by the plurality of audio acquisition components, and the vertical distance is used to characterize the vertical length of the audio acquisition component from the center point; and an identification unit, used to perform audio identification on the sampled audio signals acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance through an audio acquisition model to obtain the target audio signal of the pickup area corresponding to the audio acquisition module.

[0017] Thirdly, embodiments of this application also provide an electronic device, the electronic device including one or more processors; a memory; an audio acquisition module connected to the processor; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to perform the method described in the first aspect.

[0018] Fourthly, embodiments of this application also provide a computer-readable medium storing processor-executable program code, which, when executed by the processor, causes the processor to perform the method described in the first aspect.

[0019] The audio acquisition method, apparatus, electronic device, and computer-readable medium provided in this application include: acquiring sampled audio signals acquired by each audio acquisition component; acquiring the rotation angle and vertical distance of each audio acquisition component, wherein the rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component to the center point and a reference direction, the reference direction including the direction from the center point to a reference point on the circumference formed by the plurality of audio acquisition components, and the vertical distance is used to characterize the vertical length of the audio acquisition component from the center point; and performing audio recognition on the sampled audio signals acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance using an audio acquisition model to obtain the target audio signal of the pickup area corresponding to the audio acquisition module. Therefore, in this application, sampled audio signals are acquired jointly by the center point and multiple audio acquisition components surrounding the center point, and then the target audio signal of the pickup area corresponding to the audio acquisition module is jointly identified by combining the rotation angle and vertical distance of each audio acquisition component. This eliminates the need for additional hardware equipment to acquire audio from the pickup area, reducing usage costs and improving the convenience of audio acquisition.

[0020] Other features and advantages of the embodiments of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the embodiments of this application. The objects and other advantages of the embodiments of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart of the audio acquisition method provided in an embodiment of this application is shown; Figure 2 A schematic diagram of the audio acquisition module provided in an embodiment of this application is shown; Figure 3 A flowchart of an audio acquisition method according to another embodiment of this application is shown; Figure 4 A schematic diagram of the bottleneck structure in an embodiment of this application is shown; Figure 5 A flowchart of an audio acquisition method according to another embodiment of this application is shown; Figure 6 A structural block diagram of the audio acquisition device provided in an embodiment of this application is shown; Figure 7 This application also provides a structural block diagram of an electronic device according to another embodiment; Figure 8 This paper shows a structural block diagram of a computer-readable storage medium provided in an embodiment of this application; Figure 9 A structural block diagram of a computer program product provided in an embodiment of this application is shown. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.

[0024] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0025] Currently, with the development of electronic information technology, it is possible to collect audio signals from the microphone's pickup area. However, the current method of collecting audio signals from the pickup area is costly and inconvenient. How to collect audio signals from the pickup area conveniently and at a lower cost is an urgent problem to be solved.

[0026] Currently, audio signals from the pickup area can be acquired by obtaining distance cues and combining them with specific algorithms. For example, the specific algorithm could be a target speech extraction (TSE) algorithm, and distance cues can be obtained using a camera or other external devices, thereby enabling the acquisition of audio signals from the pickup area.

[0027] However, the inventors discovered during their research that current methods for collecting audio signals from the pickup area rely on other devices, such as cameras to obtain distance cues. This results in high costs and low convenience in collecting audio signals from the pickup area.

[0028] Therefore, this application provides an audio acquisition method, apparatus, electronic device, and computer-readable medium to solve or partially solve the above-mentioned problems.

[0029] Please see Figure 1 , Figure 1 A flowchart of an audio acquisition method provided in an embodiment of this application is shown. This method can be applied to an electronic device, and the specific execution entity can be a processor in the electronic device. The method includes steps S110 to S130.

[0030] The electronic device may be configured with an audio acquisition module, which includes a center point and a plurality of audio acquisition components surrounding the center point. Each audio acquisition component is equidistant from the center point, and the reference lines of the acquisition area corresponding to each audio acquisition component are parallel to each other.

[0031] Each audio acquisition component is equidistant from the center point; this can be achieved by equidistant distances between the geometric center of each audio acquisition component and the center point of the audio acquisition module. In essence, the center point and the multiple audio acquisition components form a circle with the center point as its center and the distance from any audio acquisition component to the center point as its radius. Each audio acquisition component is located on the circumference of this circle.

[0032] In some implementations, the structure formed by the multiple audio acquisition components can also be referred to as a circular array structure, in which the multiple audio acquisition components can rotate around a central point. For example, a slow, uniform rotation of 1 degree is completed in t seconds. For example, if t is 1, then a uniform rotation of 1 degree is completed in 1 second. Thus, one complete rotation can be completed in 360 seconds.

[0033] For example, the audio acquisition module may include eight audio acquisition components.

[0034] For example, the audio acquisition component can be a microphone, and thus the audio acquisition module can be a microphone array.

[0035] In addition, each of the audio acquisition components may have a corresponding acquisition area, which may have a reference line to characterize the direction of the acquisition area, that is, to characterize the direction in which the audio acquisition component is facing.

[0036] Step S110: Acquire the sampled audio signal acquired by each of the audio acquisition components.

[0037] As described above, the audio acquisition module includes multiple audio acquisition components, which can control each audio acquisition component to acquire audio signals, thus obtaining the sampled audio signals acquired by each audio acquisition component.

[0038] For example, taking an audio acquisition module comprising eight audio acquisition components as an example, the corresponding sampled audio signals can be acquired by each of the eight audio acquisition components. It should be noted that each sampled audio signal is acquired by a corresponding audio acquisition component.

[0039] Step S120: Obtain the rotation angle and vertical distance of each audio acquisition component, wherein the rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component and the center point and the reference direction, the reference direction includes the direction from the center point to a reference point on the circumference formed by the plurality of audio acquisition components, and the vertical distance is used to characterize the vertical length of the audio acquisition component and the center point.

[0040] Furthermore, it is also necessary to obtain the rotation angle and vertical distance of each audio acquisition component. It should be noted that the obtained rotation angle and vertical distance correspond to the same moment when the audio acquisition component acquires and samples the audio signal.

[0041] The rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component and the center point and the reference direction. The reference direction includes the direction from the center point to a reference point on the circumference formed by the plurality of audio acquisition components. The vertical distance is used to characterize the vertical length between the audio acquisition component and the center point.

[0042] For example, please refer to Figure 2 , Figure 2 A schematic diagram of the audio acquisition module provided in an embodiment of this application is shown. Figure 2 The audio acquisition module 200 shown includes a center point 210 and eight audio acquisition components 220. A reference direction can be set, which is the direction from the center point 210 to a reference point C on the circumference 230 formed by the multiple audio acquisition components. It should be noted that the reference point C can also be located at other positions on the circumference 230. Figure 2 The location of reference point C is merely one example and does not constitute a limitation on the embodiments of this application. Therefore, in Figure 2 In the diagram, the reference direction 240 can be represented by arrow M. The angle θ formed by the line 211 connecting the audio acquisition component 220 and the center point 210, and the reference direction 240, is the rotation angle.

[0043] Optionally, the rotation angle can be further defined as the angle formed by the line connecting the audio acquisition component and the center point from the reference direction to the rotation angle of the audio acquisition component.

[0044] Additionally, the vertical distance can be used to characterize the vertical length h between the audio acquisition component and the center point 210. For example, Figure 2 If the horizontal line 250 passes through the center point 210, then the length of the perpendicular line from the audio acquisition component 220 to the horizontal line 250 is the vertical distance h of the audio acquisition component 220.

[0045] Optionally, when the audio acquisition component 220 is above the horizontal line 250, the vertical distance can be set to a positive number, i.e., h; when the audio acquisition component 220 is below the horizontal line 250, the vertical distance can be set to a negative number, i.e., -h.

[0046] In some implementations, the electronic device can know in advance parameters such as the distance between the audio acquisition component and the center point, the rotation speed, the rotation angle and vertical distance of each audio acquisition component acquired last time, and the time length since the last acquisition of the rotation angle and vertical distance. Thus, it can calculate the current rotation angle and vertical distance of each audio acquisition component based on the rotation speed of multiple audio acquisition components around the center point.

[0047] It should be noted that when the audio acquisition module is not rotated, it can be in its initial state. At this time, the initial factory position of the audio acquisition module can be set, which will facilitate the subsequent calculation of the rotation angle and vertical distance of each audio acquisition component.

[0048] Please continue reading Figure 2 ,exist Figure 2 The image also shows a front region 260 of the audio acquisition module 200, which can be considered as the pickup area corresponding to the audio acquisition module. Optionally, this pickup area can be characterized by α.

[0049] Step S130: Using the audio acquisition model, perform audio recognition on the sampled audio signal acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance to obtain the target audio signal of the pickup area corresponding to the audio acquisition module.

[0050] Furthermore, the sampling audio signal acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance can be identified by the audio acquisition model to obtain the target audio signal of the pickup area corresponding to the audio acquisition module.

[0051] For example, the sampled audio signal acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance can be input into the audio acquisition model as input parameters. The parameters output by the audio acquisition model are the target audio signal of the pickup area corresponding to the audio acquisition module.

[0052] It should be noted that the audio acquisition model can be a pre-trained model. This model can be used to identify the target audio signal based on the input audio signal. The target audio signal is the audio signal located in the pickup area corresponding to the audio acquisition module. The method for training this audio acquisition model can be found in the following embodiments.

[0053] In addition, since the target audio signal of the pickup area is acquired, this audio acquisition method can also be called fixed-distance pickup.

[0054] The audio acquisition method, apparatus, electronic device, and computer-readable medium provided in this application include: acquiring sampled audio signals acquired by each audio acquisition component; acquiring the rotation angle and vertical distance of each audio acquisition component, wherein the rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component to the center point and a reference direction, the reference direction including the direction from the center point to a reference point on the circumference formed by the plurality of audio acquisition components, and the vertical distance is used to characterize the vertical length of the audio acquisition component from the center point; and performing audio recognition on the sampled audio signals acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance using an audio acquisition model to obtain the target audio signal of the pickup area corresponding to the audio acquisition module. Therefore, in this application, sampled audio signals are acquired jointly by the center point and multiple audio acquisition components surrounding the center point, and then the target audio signal of the pickup area corresponding to the audio acquisition module is jointly identified by combining the rotation angle and vertical distance of each audio acquisition component. This eliminates the need for additional hardware equipment to acquire audio from the pickup area, reducing usage costs and improving the convenience of audio acquisition.

[0055] Please see Figure 3 , Figure 3 A flowchart of an audio acquisition method provided in an embodiment of this application is shown. This method can be applied to an electronic device, and the specific execution entity can be a processor in the electronic device. The method includes steps S210 to S270.

[0056] Similar to the aforementioned embodiments, the electronic device includes an audio acquisition module, which includes a center point and a plurality of audio acquisition components surrounding the center point. Each audio acquisition component is equidistant from the center point, and the reference lines of the acquisition areas corresponding to each audio acquisition component are parallel to each other. Detailed descriptions will not be repeated here.

[0057] Step S210: Acquire the sampled audio signal acquired by each of the audio acquisition components.

[0058] Step S220: Obtain the rotation angle and vertical distance of each audio acquisition component, wherein the rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component and the center point and the reference direction, the reference direction includes the direction from the center point to a reference point on the circumference formed by the plurality of audio acquisition components, and the vertical distance is used to characterize the vertical length of the audio acquisition component and the center point.

[0059] Steps S210 and S220 have been described in detail in the foregoing embodiments and will not be repeated here.

[0060] Step S230: Obtain a training dataset and a reference dataset, wherein the training dataset consists of multiple sets of training data, each set of training data includes the training audio signal received by each of the audio acquisition components, the training rotation angle, and the training vertical distance, and the reference dataset includes reference data corresponding to each set of training data, each set of reference data includes the reference audio signal corresponding to each of the audio acquisition components.

[0061] To obtain an audio acquisition model, an initial model can be trained using a training dataset and a reference dataset. Therefore, the first step is to obtain a training dataset and a reference dataset. The training dataset can include multiple sets of training data, while the reference dataset can include multiple sets of reference data.

[0062] Specifically, the training data may include the training audio signal received by each of the audio acquisition components, the training rotation angle, and the training vertical distance, while the reference data has a one-to-one correspondence with the training data, and each of the reference data includes the reference audio signal corresponding to each of the audio acquisition components.

[0063] In some implementations, a dataset may first be acquired, which includes multiple mono audio streams. k This allows us to subsequently simulate the reverberation effect of sound sources in specific regions to obtain both the training dataset and the label dataset. For example, if S represents the dataset, then S can be represented as... In the case where the audio acquisition module includes 8 audio acquisition components, n can be 8. When the specific area is a room, the reverberation effect of the sound source within the room can be simulated using the room impulse response method, and the planar distance between the sound source and the center point of the audio acquisition module can be calculated. This facilitates subsequent simulations to obtain the training dataset and the label dataset.

[0064] Specifically, a room impulse response method can be used to construct a sound source s k The corresponding training audio signal for each audio acquisition component, wherein the training audio signal is also the representation of the sound source s k In this case, the audio acquisition component receives the sound signal. For example, it can be achieved through X. k To represent a set of training audio signals corresponding to the k-th sound source. ,in, These can be referred to as training audio.

[0065] In addition, when constructing the training dataset, it is also necessary to record the rotation angle and vertical distance of each audio acquisition component, so as to obtain the training rotation angle and training vertical distance corresponding to each set of training audio signals.

[0066] Furthermore, the reference data corresponding to each set of training data may include the reference audio signal corresponding to each audio acquisition component. The reference audio signal can be considered as the standard audio signal corresponding to that audio acquisition component.

[0067] As described above, the audio acquisition module has a corresponding pickup area. Therefore, when the sound source is within the pickup area, the reference audio signal can be confirmed as the audio signal emitted by the sound source; while when the sound source is not within the pickup area, the reference audio signal is a silence audio signal.

[0068] For specific examples, please continue reading. Figure 2 If Figure 2 The area α in front of the audio acquisition module 200 shown is the pickup area corresponding to the audio acquisition module. β is the fan-shaped angle of area α, R is the radius of area α, and D is the distance from the sound source to the center point 210. Therefore, if the distance D is greater than the fan-shaped radius R, the sample can be determined as a negative sample, and the reference audio signal is set as a silent audio signal. If the planar distance D is less than or equal to the fan-shaped radius R and is not within the fan-shaped angle β, the sample can be determined as a negative sample, and the reference audio signal is set as a silent audio signal. If the planar distance D is greater than the fan-shaped radius R and is within the fan-shaped angle β, the sample can be determined as a positive sample, and the reference audio signal can be confirmed as the audio signal emitted by the sound source.

[0069] Negative samples are used to characterize the sound source being located in the pickup area. In addition, positive samples are used to characterize that the sound source is within the pickup zone α.

[0070] For example, a silent audio signal can be an audio signal that does not include sound.

[0071] Step S240: Perform audio recognition on each set of training data using the initial model to obtain the initial audio signal corresponding to each set of training data.

[0072] Furthermore, the initial model can be trained using the training dataset and the reference dataset to obtain the audio acquisition model. The initial model can then be used to perform audio recognition on each set of training data to obtain the initial audio signal corresponding to each set of training data.

[0073] The initial model can be a neural network model.

[0074] In some implementations, the initial audio signal corresponding to the training data can be obtained by calculating the training audio signal in the training data and combining it with the rotation angle and vertical distance corresponding to each group of training audio signals. For a detailed description, please refer to the following embodiments.

[0075] Step S250: Train the initial model based on the first difference between the initial audio signal and the reference audio signal corresponding to the set of training data, so as to reduce the first difference.

[0076] After obtaining the initial audio signal, since the initial audio signal corresponds to a set of training data, a reference audio signal corresponding to that set of training data can be determined. The reference audio signal can be considered as a standard audio signal emitted by the sound source. Therefore, a first difference between the initial audio signal and the reference audio signal corresponding to the set of training data can be obtained, and the initial model is then trained based on this first difference. The purpose of training the initial model is to reduce this first difference.

[0077] For example, the first difference can be quantified by a loss function, such as the mean squared error or the squared absolute error loss function. This application does not specifically limit the specific differences.

[0078] In some implementations, the initial model can be updated through backpropagation learning, such as updating the neural network and multi-channel pooling layer in the initial model.

[0079] Step S260: Use the trained initial model as the audio acquisition model.

[0080] Furthermore, the trained initial model can be used as the audio acquisition model.

[0081] Optionally, the number of batches for training the initial model can be set. If the number of batches that have been trained is greater than or equal to the number of batches, the training of the initial model can be considered complete. In this case, the trained initial model can be used as the audio acquisition model.

[0082] Optionally, a first threshold corresponding to the first difference can be set. After obtaining the first difference each time, the relationship between the first difference and the first threshold is determined. If the first difference is greater than or equal to the first threshold, the training of the initial model continues. If the first difference is less than the first threshold, the training of the initial model is considered complete. At this time, the trained initial model can be used as the audio acquisition model.

[0083] Step S270: Using the audio acquisition model, perform audio recognition on the sampled audio signal acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance to obtain the target audio signal of the pickup area corresponding to the audio acquisition module.

[0084] Step S270 has been described in detail in the foregoing embodiments and will not be repeated here.

[0085] Optionally, the reference data in the above steps may further include reference labels, wherein the reference labels are used to characterize whether the sound source is within the pickup area. For example, the reference labels may include positive samples and negative samples, wherein negative samples are used to characterize that the sound source is outside the pickup area, and positive samples are used to characterize that the sound source is within the pickup area. Thus, in some embodiments, step S240 may further include step S241.

[0086] Step S241: Perform audio recognition on each set of training data using the initial model to obtain the initial audio signal and initial label corresponding to each set of training data.

[0087] In other words, after inputting the training data as input parameters into the initial model, the initial model can output not only the initial audio signal, but also the initial label. The initial label is the result obtained by the initial model analysis to determine whether the sound source corresponding to the training data is within the pickup area.

[0088] Furthermore, step S250 may also include step S251.

[0089] Step S251: Train the initial model based on the first difference and the second difference to reduce the first difference and the second difference, wherein the second difference is used to characterize the difference between the initial label and the reference label corresponding to the set of training data.

[0090] After obtaining the initial audio signal and the initial label, the first difference and the second difference can be determined respectively. The first difference is used to characterize the difference between the initial audio signal and the reference audio signal corresponding to the training data; while the second difference is used to characterize the difference between the initial label and the reference label corresponding to the training data.

[0091] Therefore, the initial model can be trained based on the first difference and the second difference respectively to reduce the first difference and the second difference.

[0092] Similar to the first difference, the second difference can also be quantified using a loss function, which will not be elaborated here.

[0093] Step S250 will be described in detail below with reference to specific embodiments. Specifically, step S250 may include steps S252 to S256.

[0094] Step S252: Obtain the real spectrum and imaginary spectrum corresponding to the training audio signal received by each of the audio acquisition components.

[0095] First, the acquired training audio signal can be processed to obtain the real and imaginary spectra of the training audio signal received by the audio acquisition component.

[0096] In some implementations, the real spectrum and the imaginary spectrum can be determined using Fourier transform and Euler's formula. Specifically, step S252 may include steps S2521 to S2523.

[0097] Step S2521: Perform Fourier transform on the training audio in each group of training audio signals to obtain the first formula.

[0098] Step S2522: Transform the first formula using Euler's formula to obtain the second formula.

[0099] Step S2523: Take the real part of the second formula as the real spectrum corresponding to the training audio, and take the imaginary part of the second formula as the imaginary spectrum corresponding to the training audio.

[0100] First, a Fourier transform can be performed on each group of training audio signals to obtain the first formula. As mentioned above, each group of training audio signals includes multiple training audio files. Therefore, performing a Fourier transform on each group of training audio signals is essentially performing a Fourier transform on each individual training audio file within that group.

[0101] For example, using training audio as... Taking this as an example, then... The Fourier transform can be performed using the following formula (1).

[0102] (1) in, The window length is used to characterize the Fourier transform. For example, the window length is equal to the number of Fourier transform points, nfft = 1024 points; n is the nth sampling point within the window. It is a window function, such as the Hanning window. It is the point where the window function slides along the time axis; It is the frequency. The above formula (1) is the first formula.

[0103] Furthermore, the exp in the above equation can be expanded into a complex form, that is, by transforming the first equation using Euler's formula, the corresponding real and imaginary spectra can be obtained. Specifically, the second equation can be represented by the following equation (2).

[0104] (2) The second formula can also be represented as In the form of , Re represents the real part of the second formula, and Im represents the imaginary part of the second formula. Therefore, the real part of the second formula can be used as the real spectrum corresponding to the training audio, and the imaginary part of the second formula can be used as the imaginary spectrum corresponding to the training audio.

[0105] Furthermore, taking an audio acquisition module comprising eight audio acquisition components as an example, each group of training audio signals... The training consists of eight audio samples. By sequentially following the steps described above, the real and imaginary spectra of each training audio sample in each training audio signal are obtained.

[0106] For example, if through Characterize the training audio signal, through To characterize each training audio, one can... Characterization The real spectrum corresponding to each training audio in each training audio; through The imaginary spectrum is represented by each training audio in each training audio.

[0107] In some implementations, the real and imaginary spectra can be represented by vectors, for example, by a three-dimensional vector of size [ch, F, T]. Here, ch is the number of channels, F is the frequency axis, and T is the number of frames, for example, it can be set to a size of [ch=8, F=nfft / 2=512, T=8].

[0108] Step S253: Based on the real spectrum and imaginary spectrum corresponding to each group of training audio signals, determine the amplitude spectrum feature vector of the training data.

[0109] Then, based on the real and imaginary spectra corresponding to each set of training audio signals, the amplitude spectrum feature vector of that set of training data can be determined. For example, the amplitude spectrum feature vector can be obtained by calculating the vector magnitude. Alternatively, the amplitude spectrum feature vector can be obtained by calculating the vector magnitude and then performing amplitude compression.

[0110] Specifically, step S253 may include steps S2531 and S2532.

[0111] Step S2531: Based on the square root of the sum of the squares of the real and imaginary spectra corresponding to each group of training audio signals, construct the initial amplitude spectrum feature vector of the training data.

[0112] Step S2532: Perform amplitude spectrum compression on the initial amplitude spectrum feature vector to obtain the amplitude spectrum feature vector of the training data.

[0113] It should be noted that the initial amplitude spectrum feature vector of the training data is constructed by taking the square root of the sum of the squares of the real and imaginary spectra of each training audio signal. In essence, the initial amplitude spectrum feature vector of the training audio signal is constructed by taking the square root of the sum of the squares of the real and imaginary spectra of each training audio signal in each training audio signal. Then, based on the initial amplitude spectrum feature vector of each training audio signal, the initial amplitude spectrum feature vector of the training audio signal is obtained, which is the initial amplitude spectrum feature vector of the training data.

[0114] For example, if Re represents the real spectrum and Im represents the imaginary spectrum, then the initial amplitude spectrum eigenvector can be represented as... .

[0115] Furthermore, amplitude spectrum compression is performed on the initial amplitude spectrum feature vector to obtain the amplitude spectrum feature vector of this set of training data. For example, this can be achieved through... If the amplitude spectrum eigenvectors are characterized, then amplitude spectrum compression can be characterized as follows: Where c is the compression factor, for example, c can be 0.5.

[0116] In this embodiment of the application, by performing amplitude spectrum compression on the initial amplitude spectrum feature vector, the difference between human voice and noise can be reduced, thereby enhancing the model's noise suppression effect.

[0117] Step S254: Based on the real spectrum and imaginary spectrum corresponding to each group of training audio signals, calculate the phase difference feature vector between the channels corresponding to each audio acquisition component in the training data.

[0118] Furthermore, after obtaining the real spectrum and the imaginary spectrum, the phase difference feature vector between the channels corresponding to each audio acquisition component in the training data can be calculated based on the real spectrum and the imaginary spectrum corresponding to each group of training audio signals.

[0119] In some implementations, the initial phase spectrum can be calculated first based on the real and imaginary spectra, and then the phase difference feature vector between the channels corresponding to each audio acquisition component can be determined based on the initial phase spectrum.

[0120] As described above, the audio acquisition module includes multiple audio acquisition components, allowing the determination of corresponding phase difference feature vectors between any two audio acquisition components. Taking eight audio acquisition components as an example, this results in the determination of 28 sets of phase difference feature vectors.

[0121] Specifically, step S254 may include steps S2541 and S2542.

[0122] Step S2541: Calculate the initial phase spectrum of the training data based on the real and imaginary spectra corresponding to each group of training audio signals.

[0123] Step S2542: Determine the phase difference feature vector between the channels corresponding to each audio acquisition component based on the initial phase spectrum.

[0124] First, the initial phase spectrum can be determined using the arctangent function. Specifically, the initial phase spectrum of each set of training audio signals can be determined using... To characterize, thus the initial phase spectrum can be obtained through Calculated.

[0125] Then, based on the initial phase spectrum, the phase difference feature vectors between the channels corresponding to each audio acquisition component are determined.

[0126] Step S2542 may also include steps S2543 and S2544.

[0127] Step S2543: Using the initial phase spectrum, the initial phase difference feature vector between the channels corresponding to each audio acquisition component is represented in a triangular form.

[0128] Step S2544: Perform phase encoding processing on the initial phase difference feature vector to obtain the phase difference feature vector between the channels corresponding to each audio acquisition component.

[0129] The initial phase difference feature vectors between channels corresponding to each audio acquisition component can be characterized in triangular form (Inter-Phase Difference, IPD) using the initial phase spectrum. For example, by representing the initial phase difference feature vectors with IPD, the initial phase difference feature vectors can be... The phase difference feature vector between any two audio acquisition components can be determined by calculation.

[0130] Then, the initial phase difference feature vector is further processed by phase encoding to obtain the phase difference feature vector between the channels corresponding to each audio acquisition component.

[0131] For example, a bottleneck structure can be used as a phase encoder to perform phase encoding processing on the initial phase difference feature vector. See also [example available]. Figure 4 , Figure 4 A schematic diagram of the bottleneck structure in an embodiment of this application is shown. Figure 4 The diagram shows the input layer 401, the intermediate layer 402, and the output layer 403.

[0132] If we take 8 sets of initial phase difference feature vectors as an example, the input layer 401 uses the 8 sets of initial phase difference feature vectors as the input of the bottleneck structure. The intermediate layer 402 reduces the dimensionality of the input, compressing the 8 sets into 2 sets. Then, the output layer 403 increases the dimensionality of the intermediate layer 402, from 2 sets to 8 sets, and then outputs the data through the output layer 403.

[0133] In this embodiment of the application, the initial phase difference feature vector can be compressed and then restored by the phase encoder with the bottleneck structure, and key components can be extracted.

[0134] Optionally, as described above, taking 8 audio acquisition components as an example, 28 sets of phase difference feature vectors can be determined, while the input of the subsequent bottleneck structure may only require a smaller number of sets of phase difference feature vectors, such as only 8 sets. Therefore, in some implementations, several sets of phase difference feature vectors can be randomly or redundantly selected as the input of the subsequent bottleneck structure. Taking 8 sets of phase difference feature vectors as an example, the size of the 8 sets of phase difference feature vectors can be characterized by [ch=8, F=512, T=8].

[0135] Step S255: Based on the training rotation angle and training vertical distance corresponding to each of the audio acquisition components, construct the angular distance feature vector corresponding to the set of training data.

[0136] As described above, for each set of training audio signals, each audio component can correspond to a training rotation angle and a training vertical distance. Since both the training rotation angle and the training vertical distance are geometric information and have a certain correlation, joint encoding of the two can be considered. For example, an angular distance feature vector corresponding to the training data can be constructed based on the training rotation angle and training vertical distance corresponding to each audio acquisition component. The angular distance feature vector is the vector obtained by jointly encoding the training rotation angle and the training vertical distance.

[0137] In some implementations, step S255 may include steps S2551 to S2553.

[0138] Step S2551: Concatenate the training rotation angles corresponding to each audio acquisition component to obtain a rotation angle feature vector.

[0139] Step S2552: Concatenate the training vertical distances corresponding to each audio acquisition component to obtain a vertical distance feature vector.

[0140] Step S2553: Concatenate the rotation angle feature vector and the vertical distance feature vector to obtain the angle distance feature vector corresponding to the training data.

[0141] It is understandable that for a set of training audio signals, multiple audio acquisition components in the audio acquisition module will each have a corresponding training rotation angle. Therefore, the training rotation angles corresponding to each audio acquisition component can be concatenated to obtain a rotation angle feature vector. If T=8 frames are used as a window, the size of the concatenated rotation angle feature vector can be [ch=8, F=1, T=8], where the rotation angle feature vector can be denoted as... .

[0142] Furthermore, for a set of training audio signals, each of the multiple audio acquisition components in the audio acquisition module will have a corresponding training vertical distance. Therefore, the training vertical distances corresponding to each audio acquisition component can be concatenated to obtain a vertical distance feature vector. Similarly, if a window of T=8 frames is used, the size of the concatenated distance feature vector can be [ch=8, F=1, T=8], where the distance feature vector can be denoted as H.

[0143] Furthermore, the rotation angle feature vector and the vertical distance feature vector can be concatenated to obtain the angle distance feature vector corresponding to the training data.

[0144] For example, the rotation angle feature vector and the second dimension of the distance feature vector can be concatenated to obtain an angle-distance feature vector of size [8,2,T]. For example, the angle-distance feature vector can be denoted as... .

[0145] Optionally, a bottleneck structure can be used to encode and decode the vector concatenated based on the second dimension to obtain an angle distance feature vector with more prominent features. Similar to the aforementioned use of a bottleneck structure as a phase encoder to perform phase encoding on the initial phase difference feature vector, it will not be elaborated here.

[0146] Step S256: Based on the real spectrum, imaginary spectrum, amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector, calculate the initial audio signal corresponding to the set of training data.

[0147] After obtaining the real spectrum, imaginary spectrum, amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector through the aforementioned steps, the initial audio signal corresponding to the training data set can be calculated. For example, the initial audio signal can be obtained through vector concatenation, noise reduction, or other methods.

[0148] Specifically, step S256 may include steps S2561 to S2564.

[0149] Step S2561: Concatenate the amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector to obtain the first mixed feature vector.

[0150] First, the phase difference feature vector and the angular distance feature vector can be concatenated to obtain a concatenated vector. Then, the amplitude spectrum feature vector is concatenated with this concatenated vector again to obtain the first mixed feature vector.

[0151] Taking a phase difference feature vector with dimensions [8, nfft / 2, T] and an angular distance feature vector with dimensions [8, 2, T] as an example, a concatenated vector can be obtained by concatenating along the second dimension. The dimensions of the concatenated vector are then [8, nfft / 2 + 2, T]. In some implementations, the concatenated vector can be denoted as... .

[0152] Furthermore, the amplitude spectrum feature vector can be concatenated with the concatenated vector again. For example, the main neural network of the initial model can be of the U-Net type, which includes an encoding part and a decoding part. Thus, the amplitude spectrum feature vector obtained through amplitude spectrum compression can be input into the initial model and concatenated with the concatenated vector in a certain layer of the encoding part. For example, concatenation can be performed along the frequency axis to obtain a first mixed feature vector. This first mixed feature vector then continues to participate in the subsequent encoding part.

[0153] For example, the encoding function can be implemented using an encoding network, and the decoding function can be implemented using a decoding network.

[0154] Step S2562: Encode and decode the first mixed feature vector to obtain mask information corresponding to each of the audio acquisition components.

[0155] It should be noted that the first mixed feature vector obtained above is actually corresponding to each audio acquisition component. Therefore, in some embodiments, further encoding and decoding of the first mixed feature vector can yield mask information corresponding to each audio acquisition component. The mask information is used to characterize the degree of sound masking. For example, the mask information can be a probability matrix. If the probability is closer to 1, it indicates a lower degree of sound masking, suggesting a potentially larger proportion of effective human voice in the sound; conversely, if the probability is closer to 0, it indicates a higher degree of sound masking, suggesting a potentially larger proportion of noise in the sound.

[0156] Step S2563: Based on the mask information, perform noise reduction processing on the corresponding real spectrum and imaginary spectrum respectively to obtain the noise-reduced real spectrum and noise-reduced imaginary spectrum respectively.

[0157] Furthermore, noise reduction processing can be performed on the corresponding real and imaginary spectra based on the mask information to obtain the noise-reduced real and imaginary spectra, respectively. For example, if the audio acquisition module includes 8 audio acquisition components, the mask information can be multiplied by the real and imaginary spectra obtained through the aforementioned steps to obtain the noise-reduced real and imaginary spectra, respectively.

[0158] For example, if we take Characterize mask information, in order to Characterizing the denoised real spectrum, with Characterizing the denoised imaginary spectrum, with Characterizing the spectrum of real numbers, with The imaginary spectrum can be characterized by... Obtain the denoised real spectrum; through The denoised imaginary spectrum is obtained.

[0159] Step S2564: Calculate the initial audio signal corresponding to the set of training data based on the denoised real spectrum and the denoised imaginary spectrum.

[0160] Therefore, after obtaining the denoised real spectrum and the denoised imaginary spectrum, the initial audio signal corresponding to the training data set can be calculated based on the denoised real spectrum and the denoised imaginary spectrum. In some implementations, a multi-channel pooling layer and inverse Fourier transform can be combined to calculate the initial audio signal. Specifically, step S2564 may include steps S2565 to S2567.

[0161] Step S2565: Summing and averaging the denoised real spectra of multiple channels in the training data using a pooling layer to obtain the single-channel real spectrum.

[0162] Step S2566: Summing and averaging the denoised imaginary spectra of the multiple channels in the training data using a pooling layer to obtain the single-channel imaginary spectrum.

[0163] Step S2567: Perform inverse Fourier transform based on the single-channel real spectrum and the single-channel imaginary spectrum to obtain the initial audio signal corresponding to the training data.

[0164] As described above, the denoised real and imaginary spectra obtained correspond to a set of training audio signals and each audio acquisition component. Therefore, the multi-channel denoised real and imaginary spectra corresponding to each set of training audio signals essentially correspond one-to-one with the audio acquisition components.

[0165] Therefore, in some implementations, a single-channel real spectrum can be obtained by summing and averaging the denoised real spectra of multiple channels in the training data using a pooling layer.

[0166] The pooling layer is used to calculate the pooling of the channel axes; for example, it can be calculated using mean pooling. To characterize a single-channel real spectrum, one can use... The single-channel real spectrum was calculated.

[0167] Alternatively, a single-channel imaginary spectrum can be obtained by summing and averaging the denoised imaginary spectra of multiple channels in the training data using a pooling layer. Similarly, if... To characterize the single-channel imaginary spectrum, one can use... The single-channel imaginary spectrum was calculated.

[0168] Here, ch is the number of channels, thus enabling multiple channels to be merged into a single channel.

[0169] For some implementations, single-channel real spectrum and single-channel imaginary spectrum The dimensions can all be [1, nfft / 2, T].

[0170] Furthermore, an inverse Fourier transform can be performed based on the single-channel real spectrum and the single-channel imaginary spectrum to obtain the initial audio signal corresponding to the training data.

[0171] Specifically, the single-channel real spectrum and the single-channel imaginary spectrum can be represented as the Fourier transform of exp using Euler's formula, and then an inverse Fourier transform can be performed to obtain the initial audio signal corresponding to the training data. In some implementations, the initial audio signal can be represented as... .

[0172] The audio acquisition method provided in this application, through an audio acquisition module comprising a center point and multiple audio acquisition components surrounding the center point, acquires an array with a larger synthetic aperture. It can combine the phase difference features of each audio acquisition component to achieve phase difference compensation, resulting in a more accurate target audio signal. Furthermore, in this application embodiment, by constructing joint encoding of angular distance feature vectors and applying it within the neural network, the acquired target audio signal becomes more stable and robust. Additionally, in this application embodiment, by performing amplitude spectrum compression on the initial amplitude spectrum feature vector, the difference between human voice and noise can be reduced, thereby enhancing the model's noise suppression effect. Moreover, through a bottleneck-structured phase encoder, the initial phase difference feature vector can be compressed and then restored, enabling the extraction of key components.

[0173] Please see Figure 5 , Figure 5 A flowchart of an audio acquisition method provided in an embodiment of this application is shown. This method can be applied to an electronic device, and the specific execution entity can be a processor in the electronic device. The method includes steps S310 to S3140.

[0174] Similar to the aforementioned embodiments, the electronic device includes an audio acquisition module, which includes a center point and a plurality of audio acquisition components surrounding the center point. Each audio acquisition component is equidistant from the center point, and the reference lines of the acquisition areas corresponding to each audio acquisition component are parallel to each other. Detailed descriptions will not be repeated here.

[0175] Step S310: Sample the audio signal.

[0176] Step S320: Real number spectrum and imaginary number spectrum.

[0177] Step S330: Amplitude spectrum eigenvectors.

[0178] After obtaining the real and imaginary spectra for each set of training audio signals, the amplitude spectrum feature vector of that set of training data can be further determined. Additionally, the amplitude spectrum feature vector can also be obtained from the real and imaginary spectra.

[0179] Step S340: Initial phase spectrum.

[0180] Step S350: Phase difference eigenvector.

[0181] Specifically, the initial phase spectrum of the training data can be calculated based on the real and imaginary spectra corresponding to each group of training audio signals; then, the phase difference feature vector between the channels corresponding to each audio acquisition component can be determined based on the initial phase spectrum.

[0182] Step S360: Rotation angle and vertical distance.

[0183] Step S370: Angular distance feature vector.

[0184] Additionally, the rotation angle and vertical distance of each audio acquisition component can be obtained, and then an angle-distance feature vector can be constructed.

[0185] After steps S350 and S370, the process can proceed to step S380.

[0186] Step S380: Encode the network.

[0187] Step S390: Decode the network.

[0188] Step S3100: Mask information.

[0189] Therefore, by combining the encoding and decoding networks, the mask information can be obtained. After step S390, the mask information needs to be determined by combining the real and imaginary spectra from step S320.

[0190] Step S3110: Noise reduction of real spectrum and noise reduction of imaginary spectrum.

[0191] Furthermore, the corresponding real and imaginary spectra can be denoised based on the mask information to obtain denoised real and imaginary spectra, respectively.

[0192] Step S3120: Pooling layer.

[0193] Step S3130: Single-channel real spectrum and single-channel imaginary spectrum.

[0194] Then, the pooling layer is used to sum and average the denoised real spectra of multiple channels in the training data to obtain the single-channel real spectrum; the pooling layer is used to sum and average the denoised imaginary spectra of multiple channels in the training data to obtain the single-channel imaginary spectrum.

[0195] Step S3140: Target audio signal.

[0196] This allows us to obtain the target audio signal.

[0197] For a detailed description of each of the above steps, please refer to the aforementioned embodiments; they will not be repeated here.

[0198] Please see Figure 6This document illustrates a structural block diagram of an audio acquisition device 600 provided in an embodiment of this application, which is applied to an electronic device. The electronic device includes an audio acquisition module, which includes a center point and a plurality of audio acquisition components surrounding the center point. Each audio acquisition component is equidistant from the center point, and the reference lines of the acquisition areas corresponding to each audio acquisition component are parallel to each other. The device 600 includes a first acquisition unit 610, a second acquisition unit 620, and a recognition unit 630.

[0199] The first acquisition unit 610 is used to acquire the sampled audio signal acquired by each of the audio acquisition components.

[0200] The second acquisition unit 620 is used to acquire the rotation angle and vertical distance of each of the audio acquisition components, wherein the rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component and the center point and the reference direction, the reference direction includes the direction from the center point to a reference point on the circumference formed by the plurality of audio acquisition components, and the vertical distance is used to characterize the vertical length of the audio acquisition component and the center point.

[0201] The recognition unit 630 is used to perform audio recognition on the sampled audio signal acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance through the audio acquisition model, so as to obtain the target audio signal of the pickup area corresponding to the audio acquisition module.

[0202] Optionally, the audio acquisition device 600 may also include a training unit ( Figure 6 (Not shown in the image), this training unit can be used to acquire a training dataset and a reference dataset. The training dataset contains multiple sets of training data, each set including the training audio signal received by each audio acquisition component, the training rotation angle, and the training vertical distance. The reference dataset includes reference data corresponding to each set of training data, and each set of reference data includes a reference audio signal corresponding to each audio acquisition component. An initial model is used to perform audio recognition on each set of training data to obtain an initial audio signal corresponding to each set of training data. The initial model is trained based on a first difference between the initial audio signal and the reference audio signal corresponding to that set of training data to reduce the first difference. The trained initial model is then used as the audio acquisition model.

[0203] Wherein, when the sound source is in the pickup area, the reference audio signal is the audio signal emitted by the sound source; when the sound source is not in the pickup area, the reference audio signal is a silence audio signal.

[0204] Optionally, the training unit can also be used to perform audio recognition on each set of training data using an initial model to obtain an initial audio signal and an initial label corresponding to each set of training data; and train the initial model based on a first difference and a second difference to reduce the first difference and the second difference, wherein the second difference is used to characterize the difference between the initial label and the reference label corresponding to the set of training data.

[0205] Optionally, the training unit can also be used to acquire the real spectrum and imaginary spectrum corresponding to the training audio signal received by each of the audio acquisition components; determine the amplitude spectrum feature vector of the training data based on the real spectrum and imaginary spectrum corresponding to each group of training audio signals; calculate the phase difference feature vector between the channels corresponding to each audio acquisition component in the training data based on the real spectrum and imaginary spectrum corresponding to each group of training audio signals; construct the angular distance feature vector corresponding to the training data based on the training rotation angle and training vertical distance corresponding to each of the audio acquisition components; and calculate the initial audio signal corresponding to the training data based on the real spectrum, imaginary spectrum, amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector.

[0206] Optionally, the training unit can also be used to construct an initial amplitude spectrum feature vector for the training data based on the square root of the sum of the squares of the real and imaginary spectra corresponding to each training audio signal; and to perform amplitude spectrum compression on the initial amplitude spectrum feature vector to obtain the amplitude spectrum feature vector for the training data.

[0207] Optionally, the training unit can also be used to calculate the initial phase spectrum of the training data based on the real and imaginary spectra corresponding to each group of training audio signals; and to determine the phase difference feature vector between the channels corresponding to each audio acquisition component based on the initial phase spectrum.

[0208] Optionally, the training unit can also be used to represent the initial phase difference feature vector between the channels corresponding to each audio acquisition component in a triangular manner using the initial phase spectrum; and to perform phase encoding processing on the initial phase difference feature vector to obtain the phase difference feature vector between the channels corresponding to each audio acquisition component.

[0209] Optionally, the training unit can also be used to concatenate the training rotation angles corresponding to each of the audio acquisition components to obtain a rotation angle feature vector; concatenate the training vertical distances corresponding to each of the audio acquisition components to obtain a vertical distance feature vector; and concatenate the rotation angle feature vectors and the vertical distance feature vectors to obtain the angle distance feature vectors corresponding to the set of training data.

[0210] Optionally, the training unit can also be used to concatenate the amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector to obtain a first mixed feature vector; perform encoding and decoding processing on the first mixed feature vector to obtain mask information corresponding to each of the audio acquisition components; perform noise reduction processing on the corresponding real spectrum and imaginary spectrum based on the mask information to obtain a denoised real spectrum and a denoised imaginary spectrum respectively; and calculate the initial audio signal corresponding to the set of training data based on the denoised real spectrum and the denoised imaginary spectrum.

[0211] Optionally, the training unit can also be used to sum and average the denoised real spectra of multiple channels in the training data set through a pooling layer to obtain a single-channel real spectrum; to sum and average the denoised imaginary spectra of multiple channels in the training data set through a pooling layer to obtain a single-channel imaginary spectrum; and to perform an inverse Fourier transform based on the single-channel real spectrum and the single-channel imaginary spectrum to obtain the initial audio signal corresponding to the training data set.

[0212] The training unit can also be used to perform Fourier transform on the training audio in each group of training audio signals to obtain the first formula; transform the first formula using Euler's formula to obtain the second formula; use the real part of the second formula as the real spectrum corresponding to the training audio, and use the imaginary part of the second formula as the imaginary spectrum corresponding to the training audio.

[0213] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0214] In the several embodiments provided in this application, the coupling between the units can be electrical, mechanical or other forms of coupling.

[0215] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0216] Please see Figure 7 This illustration shows a structural block diagram of an electronic device 700 according to another embodiment of this application. The electronic device 700 can be a smartphone, tablet computer, or other electronic device capable of running computer programs. The electronic device 100 in this application may include one or more components: a processor 710 and a memory 720. The one or more processors are used to execute the methods described in the foregoing embodiments.

[0217] The processor 710 may include one or more processing cores. The processor 710 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 720, and by calling data stored in the memory 720. Optionally, the processor 710 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 710 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and computer programs; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 710 and may be implemented separately using a communication chip.

[0218] The memory 720 may include Double Data Rate Synchronous Dynamic Random Access Memory (DDR) and Static Random Access Memory (SRAM). Optionally, the memory 720 may also include external storage, such as a hard disk drive (HDD), solid-state drive (SSD), USB flash drive, or flash memory card. The memory 720 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 720 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created by the electronic device 700 during use.

[0219] Please refer to Figure 8 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 800 stores program code that can be called by a processor to execute the methods described in the above method embodiments.

[0220] The computer-readable storage medium 800 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 800 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 800 has storage space for program code 810 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 810 may, for example, be compressed in a suitable form.

[0221] Please refer to Figure 9 The diagram illustrates a structural block diagram 900 of a computer program product provided in an embodiment of this application. The computer program product 900 includes a computer program / instructions 910, which, when executed by a processor, implements the steps of the aforementioned method.

[0222] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An audio acquisition method, characterized in that, An electronic device is used, the electronic device including an audio acquisition module, the audio acquisition module including a center point and a plurality of audio acquisition components surrounding the center point, wherein each audio acquisition component is equidistant from the center point, and reference lines of the acquisition areas corresponding to each audio acquisition component are parallel to each other, the method including: Acquire the sampled audio signal acquired by each of the audio acquisition components; Obtain the rotation angle and vertical distance of each audio acquisition component, wherein the rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component to the center point and the reference direction, the reference direction includes the direction from the center point to a reference point on the circumference formed by the plurality of audio acquisition components, and the vertical distance is used to characterize the vertical length of the audio acquisition component to the center point; The audio acquisition model performs audio recognition on the sampled audio signals acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance, to obtain the target audio signal of the pickup area corresponding to the audio acquisition module.

2. The method according to claim 1, characterized in that, Before performing audio recognition on the sampled audio signals acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance through the audio acquisition model to obtain the target audio signal of the pickup area corresponding to the audio acquisition module, the method further includes: Obtain a training dataset and a reference dataset, wherein the training dataset comprises multiple sets of training data, each set of training data includes the training audio signal received by each of the audio acquisition components, the training rotation angle, and the training vertical distance, and the reference dataset includes reference data corresponding to each set of training data, each set of reference data includes the reference audio signal corresponding to each of the audio acquisition components; The initial model is used to perform audio recognition on each set of training data to obtain the initial audio signal corresponding to each set of training data. The initial model is trained based on the first difference between the initial audio signal and the reference audio signal corresponding to the set of training data, so as to reduce the first difference; The trained initial model is used as the audio acquisition model.

3. The method according to claim 2, characterized in that, When the sound source is located in the pickup area, the reference audio signal is the audio signal emitted by the sound source; When the sound source is not located in the pickup area, the reference audio signal is a silent audio signal.

4. The method according to claim 3, characterized in that, The reference data also includes reference labels, which are used to characterize whether the sound source is located in the pickup area. The step of performing audio recognition on each set of training data using an initial model to obtain the initial audio signal corresponding to each set of training data includes: The initial model is used to perform audio recognition on each set of training data to obtain the initial audio signal and initial label corresponding to each set of training data. The initial model is trained based on a first difference between the initial audio signal and the reference audio signal corresponding to the set of training data, in order to reduce the first difference, including: The initial model is trained based on the first difference and the second difference to reduce the first difference and the second difference, wherein the second difference is used to characterize the difference between the initial label and the reference label corresponding to the set of training data.

5. The method according to claim 2, characterized in that, The step of performing audio recognition on each set of training data using an initial model to obtain the initial audio signal corresponding to each set of training data includes: Obtain the real spectrum and imaginary spectrum corresponding to the training audio signal received by each of the audio acquisition components; Based on the real and imaginary spectra corresponding to each group of training audio signals, determine the amplitude spectrum feature vector of the training data for that group; Based on the real and imaginary spectra of each training audio signal, calculate the phase difference feature vector between the channels of each audio acquisition component in the training data. Based on the training rotation angle and training vertical distance corresponding to each of the audio acquisition components, an angular distance feature vector corresponding to this set of training data is constructed; Based on the real spectrum, imaginary spectrum, amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector, the initial audio signal corresponding to the training data is calculated.

6. The method according to claim 5, characterized in that, The determination of the amplitude spectrum feature vector of the training data based on the real and imaginary spectra corresponding to each group of training audio signals includes: The initial amplitude spectrum feature vector of the training data is constructed by taking the square root of the sum of the squares of the real and imaginary spectra of each training audio signal. The initial amplitude spectrum feature vector is compressed to obtain the amplitude spectrum feature vector of the training data.

7. The method according to claim 5, characterized in that, The step of calculating the phase difference feature vector between channels corresponding to each audio acquisition component in the training data based on the real and imaginary spectra of each group of training audio signals includes: Calculate the initial phase spectrum of the training data based on the real and imaginary spectra corresponding to each group of training audio signals; Based on the initial phase spectrum, the phase difference feature vectors between the channels corresponding to each audio acquisition component are determined.

8. The method according to claim 7, characterized in that, The step of determining the phase difference feature vector between channels corresponding to each audio acquisition component based on the initial phase spectrum includes: The initial phase spectrum is used to represent the initial phase difference feature vector between the channels corresponding to each audio acquisition component in a triangular form. Phase encoding processing is performed on the initial phase difference feature vector to obtain the phase difference feature vector between the channels corresponding to each audio acquisition component.

9. The method according to claim 5, characterized in that, The step of constructing an angular distance feature vector corresponding to the training data based on the training rotation angle and training vertical distance corresponding to each of the audio acquisition components includes: The training rotation angles corresponding to each of the audio acquisition components are concatenated to obtain a rotation angle feature vector; The vertical distances corresponding to the training vertical distances of each audio acquisition component are concatenated to obtain a vertical distance feature vector; The rotation angle feature vector and the vertical distance feature vector are concatenated to obtain the angle distance feature vector corresponding to the training data.

10. The method according to claim 5, characterized in that, The calculation of the initial audio signal corresponding to the training data based on the real spectrum, imaginary spectrum, amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector includes: The amplitude spectrum feature vector, phase difference feature vector, and angular distance feature vector are concatenated to obtain the first mixed feature vector; The first mixed feature vector is encoded and decoded to obtain mask information corresponding to each audio acquisition component; Based on the mask information, the corresponding real and imaginary spectra are denoised to obtain the denoised real and imaginary spectra, respectively. The initial audio signal corresponding to the set of training data is calculated based on the denoised real spectrum and the denoised imaginary spectrum.

11. The method according to claim 10, characterized in that, The calculation of the initial audio signal corresponding to the training data based on the denoised real spectrum and the denoised imaginary spectrum includes: The single-channel real spectrum is obtained by summing and averaging the denoised real spectra of the multi-channel training data through the pooling layer. The single-channel imaginary spectrum is obtained by summing and averaging the denoised imaginary spectra of the multi-channel training data through the pooling layer. The initial audio signal corresponding to the training data is obtained by performing an inverse Fourier transform based on the single-channel real spectrum and the single-channel imaginary spectrum.

12. The method according to claim 5, characterized in that, The step of obtaining the real spectrum and imaginary spectrum corresponding to the training audio signal received by each of the audio acquisition components includes: Perform a Fourier transform on the training audio in each group of training audio signals to obtain the first formula; The first equation is transformed using Euler's formula to obtain the second equation; The real part of the second formula is taken as the real spectrum corresponding to the training audio, and the imaginary part of the second formula is taken as the imaginary spectrum corresponding to the training.

13. An audio acquisition device, characterized in that, An electronic device is used in which an audio acquisition module is included. The audio acquisition module includes a center point and a plurality of audio acquisition components arranged around the center point. Each audio acquisition component is equidistant from the center point, and reference lines for the acquisition areas corresponding to each audio acquisition component are parallel to each other. The device includes: The first acquisition unit is used to acquire the sampled audio signal acquired by each of the audio acquisition components; The second acquisition unit is used to acquire the rotation angle and vertical distance of each audio acquisition component, wherein the rotation angle is used to characterize the angle formed by the line connecting the audio acquisition component and the center point and the reference direction, the reference direction includes the direction from the center point to the reference point on the circumference formed by the plurality of audio acquisition components, and the vertical distance is used to characterize the vertical length of the audio acquisition component and the center point. The recognition unit is used to perform audio recognition on the sampled audio signal acquired by each audio acquisition component, the rotation angle of the audio acquisition component, and the vertical distance through the audio acquisition model, so as to obtain the target audio signal of the pickup area corresponding to the audio acquisition module.

14. An electronic device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the method as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The readable storage medium stores program code that can be invoked by a processor to execute the method as described in any one of claims 1-12.