Audio acquisition method and device, storage medium and equipment

By using the pre-generated reference voiceprint vector and cross attention mechanism in the neural network model, user audio features are extracted from target audio data, which solves the problem of inaccurate user audio data extraction in noise scenarios in the prior art, and improves the noise reduction effect and device response speed.

CN120496549APending Publication Date: 2025-08-15MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510602189.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing personalized voice noise reduction technology is difficult to accurately extract user audio data in noise scenarios, especially when ambient noise and human voice noise are complex, the noise reduction effect is high uncertainty and depends on high-quality registered audio data.

Method used

Based on the pre-generated user reference voiceprint vector, user audio features are extracted from the target audio data through a neural network model, and user audio is separated using the cross attention mechanism and mask matrix to improve the representation ability of voiceprint vectors.

Benefits of technology

Improve the noise reduction effect in complex noise environments, reduce dependence on high-quality registered audio data, and enhance the response speed and accuracy of voice control devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496549A_ABST
    Figure CN120496549A_ABST
Patent Text Reader

Abstract

The invention provides an audio obtaining method and device, a storage medium and equipment, the method is applied to the field of electronic equipment, the method obtains target audio features from target audio data of a user in an application scene, and the target audio features comprise noise features and user audio features of the user; the pre-generated reference voiceprint vector of the user is taken as a reference, the user audio feature is acquired from the target audio, and then the user audio data is acquired based on the user audio feature. Wherein the reference voiceprint vector is a fusion voiceprint vector of common representation extracted from a pre-collected registered audio data set of the user, the representation capability of the voiceprint vector is improved, and then the noise reduction effect in the audio acquisition process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio technology, and more specifically, to an audio acquisition method, apparatus, storage medium, and device in the field of audio technology. Background Art

[0002] Personalized voice noise reduction technology refers to the ability to accurately extract user audio data in noisy environments. Since noisy environments can include both ambient and vocal noise, accurately extracting user audio data requires not only addressing the interference of ambient noise but also distinguishing between similar-sounding human voices. To improve the accuracy of voice noise reduction, users are required to pre-register audio data as a reference. However, this method relies heavily on the quality of the registered audio data, resulting in significant uncertainty in the voice noise reduction and impacting its effectiveness. Summary of the Invention

[0003] The present application provides an audio acquisition method, apparatus, storage medium, and device, which can improve the characterization capability of voiceprint vectors, thereby enhancing the noise reduction effect during the audio acquisition process.

[0004] In a first aspect, an audio acquisition method is provided, which includes: collecting target audio data of a user in an application scenario, and obtaining target audio features of the target audio data; obtaining user audio features of the user in the target audio features based on a pre-generated reference voiceprint vector of the user; and generating user audio data of the user based on the user audio features; wherein the reference voiceprint vector is a fused voiceprint vector of common representations extracted from a pre-collected set of registered audio data of the user.

[0005] The above technical solution obtains target audio features from the target audio data of the user in the application scenario. The target audio features include noise features and the user's audio features. Using a pre-generated reference voiceprint vector of the user as a reference, the user audio features are obtained from the target audio, and then the user audio data is obtained based on the user audio features. The reference voiceprint vector is a fused voiceprint vector of common features extracted from a pre-collected set of registered audio data for the user. This improves the representational capabilities of the voiceprint vector and, in turn, enhances the noise reduction effect during the audio acquisition process.

[0006] In a second aspect, an audio acquisition device is provided, the device comprising:

[0007] A target audio feature acquisition unit is used to collect target audio data of the user in the application scenario and obtain target audio features of the target audio data;

[0008] A user audio feature acquisition unit, configured to acquire the user audio feature of the user from the target audio feature based on a pre-generated reference voiceprint vector of the user;

[0009] A user audio data generating unit, configured to generate user audio data of a user based on the user audio features;

[0010] The reference voiceprint vector is a fused voiceprint vector of common features extracted from a pre-collected set of registered audio data of the user.

[0011] In a third aspect, an electronic device is provided, the electronic device comprising: a memory for storing executable program code;

[0012] A processor is used to call and run executable program code from a memory to execute the method in the above-mentioned first aspect or any possible implementation of the first aspect.

[0013] In a fourth aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the method in the first aspect or any possible implementation of the first aspect.

[0014] In a fifth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method in the above-mentioned first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a schematic diagram of a scenario of an audio acquisition method provided in an embodiment of the present application;

[0016] Figure 2 This is a flow chart of an audio acquisition method provided in an embodiment of the present application;

[0017] Figure 3 This is a flow chart of an audio acquisition method provided in an embodiment of the present application;

[0018] Figure 4 This is a schematic diagram of an example of a simulation data set provided in an embodiment of the present application;

[0019] Figure 5 This is a flow chart of an audio acquisition method provided in an embodiment of the present application;

[0020] Figure 6 1 is a flow chart of a method for extracting a reference voiceprint vector provided in an embodiment of the present application;

[0021] Figure 7 1 is a flow chart of a method for extracting a reference voiceprint vector provided in an embodiment of the present application;

[0022] Figure 8 This is a flowchart of a method for determining a mask matrix provided in an embodiment of the present application;

[0023] Figure 9 This is a flowchart of a method for determining a mask matrix provided in an embodiment of the present application;

[0024] Figure 10 This is a flow chart of a method for collecting target audio data provided by an embodiment of the present application;

[0025] Figure 11 This is a structural diagram of an audio acquisition device provided in an embodiment of the present application;

[0026] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will clearly and thoroughly describe the technical solutions in this application in conjunction with the accompanying drawings. In the description of the embodiments of this application, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more than two.

[0028] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.

[0029] The audio acquisition method provided in the embodiment of the present application can be applied to noise reduction scenarios when voice controls electronic devices or when voice calls are made through electronic devices. Electronic devices include but are not limited to mobile phones, personal computers, laptops, smart home devices, vehicle-mounted devices, wearable devices and other terminal devices that can be controlled by voice. The electronic device collects the target audio data in the application scenario through a microphone (Microphone, MIC). In order to ensure that the electronic device can correctly obtain the user audio data of the user who issued the instruction to accurately execute the instruction, it is internally configured with a noise reduction model that can extract the user audio data of the user from the target audio data in application scenarios with severe noise. Application scenarios include scenarios in which users make voice calls, online meetings, online education, medical calls, etc. through electronic devices such as mobile phones, personal computers, laptops, etc., and also include scenarios in which users control smart home devices or vehicle-mounted devices through voice. The noise in the above application scenarios may include ambient noise and human voice noise. When the noise is more complex, the noise reduction model not only needs to deal with the interference of ambient noise, but also needs to distinguish human voice noise with similar timbre.

[0030] See Figure 1 , Figure 1 This is a scene diagram of an audio acquisition method provided by an embodiment of the present application. Figure 1 As shown in the figure, the electronic device is a robot vacuum cleaner in a smart home device, and the application scenario is a scenario where a user controls the robot vacuum cleaner to execute cleaning commands through voice. When user A sends user audio data containing cleaning commands to robot vacuum cleaner B, while user C is speaking, robot vacuum cleaner B uses its built-in MIC to collect the target audio data for the current application scenario. This target audio data contains user A's user audio data and the human voice noise emitted by user C. Robot vacuum cleaner B needs to obtain user A's user audio data from the target audio data to execute the cleaning command corresponding to the user audio data.

[0031] Based on the above problems, an embodiment of the present application provides an audio acquisition method for acquiring target audio features from target audio data of a user in an application scenario. The target audio features include noise features and the user's audio features. Using a pre-generated reference voiceprint vector of the user as a reference, the user audio features are acquired from the target audio, and then the user audio data is acquired based on the user audio features. The reference voiceprint vector is a fused voiceprint vector of common features extracted from a pre-collected set of registered audio data of the user, which improves the representational capability of the voiceprint vector and thereby enhances the noise reduction effect during the audio acquisition process.

[0032] based on Figure 1 The following scene diagram will be combined with Figure 2-Figure 10 , the audio acquisition method provided in the embodiment of the present application is introduced in detail.

[0033] See Figure 2 , Figure 2 This is a flow chart of an audio acquisition method provided by an embodiment of the present application. Figure 2 As shown, the method of the embodiment of the present application may include the following steps S101 to S103.

[0034] S101, collecting target audio data of the user in the application scenario and obtaining target audio features of the target audio data;

[0035] Specifically, the electronic device collects the target audio data of the user in the application scenario, and obtains the target audio features from the target audio data by audio encoding the target audio data. The user has registered the audio information in the electronic device in advance, and the application scenario is the scenario in which the user controls the electronic device by voice or the scenario in which the user makes a voice call through the electronic device. The target audio data includes at least user audio data, and the target audio data may also include noise data, and the noise data includes environmental noise and human voice noise. Audio encoding is a technical means for obtaining audio features from audio data, including time-frequency transformation and time-frequency mapping of audio data. The audio features are used to match the voice information pre-registered by the user to obtain the user audio data from the target audio data, so that the electronic device can correctly identify the user audio data and then execute the user's voice commands.

[0036] For example, if the electronic device is a robot vacuum cleaner in a smart home, the application scenario could be a user controlling the robot vacuum cleaner to start cleaning or switch cleaning modes via voice commands. Assuming the user issues a command such as "automatically start cleaning the entire house at 3:00 PM every day," and there is ambient noise and other people talking around when the user issues the command, the target audio data collected by the robot vacuum cleaner would be "automatically start cleaning the entire house at 3:00 PM every day + ambient noise + other people talking except the user."

[0037] S102, obtaining a user audio feature of the user from the target audio feature based on a pre-generated reference voiceprint vector of the user;

[0038] Specifically, based on a user's reference voiceprint vector pre-generated in the electronic device, the user audio features of the user that match the reference voiceprint vector are obtained from the target audio features. A registered audio data set is a collection of a user's registered audio data collected by the electronic device, including at least one piece of registered audio data, used to determine whether the user has control authority over the electronic device. A reference voiceprint vector is a fused voiceprint vector of common features extracted from the pre-collected user's registered audio data set. A voiceprint vector is a digitized vector of voice features that uniquely identifies a user, extracted through analysis and processing of the audio data.

[0039] S103: Generate user audio data of the user based on the user audio features.

[0040] Specifically, the user's audio features are decoded to obtain the user's audio data, so that the electronic device can perform voice recognition on the user's audio data and execute the corresponding voice command. Feature decoding is the inverse operation process corresponding to audio encoding, which is used to obtain audio data based on audio features, including inverse mapping and inverse time-frequency transformation. When a user controls a smart home device through voice, the audio acquisition method provided in the embodiment of the present application is used to obtain the user's audio data, which can significantly reduce the impact of noise, thereby improving the response speed of the smart home device and the accuracy of voice control.

[0041] In an embodiment of the present application, target audio features are obtained from the target audio data of the user in the application scenario. The target audio features include noise features and the user's audio features. A pre-generated reference voiceprint vector of the user is used as a reference to obtain the user's audio features from the target audio, and then the user's audio data is obtained based on the user's audio features. The reference voiceprint vector is a fused voiceprint vector of common features extracted from a pre-collected set of registered audio data for the user, which improves the representational capabilities of the voiceprint vector and thereby enhances the noise reduction effect during the audio acquisition process.

[0042] See Figure 3 , Figure 3 This is a flow chart of an audio acquisition method provided by an embodiment of the present application. Figure 3 As shown, the method of the embodiment of the present application may include the following steps S201 to S206.

[0043] S201, obtaining a simulation data set, and training a neural network model based on the simulation data set;

[0044] Specifically, the speech noise reduction model adopted in the embodiment of the present application is a neural network model, a simulation data set is obtained, and the neural network model is trained based on the simulation data set. The neural network model is a computational model designed to imitate the structure and function of human brain neurons, and is composed of multiple layers of interconnected "neurons". Using the neural network model as a speech noise reduction model can adapt to complex noise scenes, and at the same time support multimodal fusion (audio, image data) to enhance the noise reduction effect, thereby improving the accuracy and applicability of speech noise reduction. When obtaining the simulation data set, noise data and human voice audio are obtained, and the noise data includes environmental noise and human voice noise. The simulation data set is determined based on the noise data and human voice audio. The simulation data set includes first simulation data, second simulation data, and third simulation data, wherein the first simulation data is a composite audio data of environmental noise and human voice audio, the second simulation data is a composite audio data of human voice noise and human voice audio, and the third simulation data is a composite audio data of environmental noise, human voice noise, and human voice audio. The human voice audio is the audio data of the target speaker that the neural network model ultimately needs to identify, and the human voice noise in the noise data is the audio data of other people other than the target speaker. The noise reduction effect of the neural network model is further improved by using a simulation data set composed of multiple simulation data.

[0045] The simulation data set consists of the first simulation data, the second simulation data, and the third simulation data in the same proportion or different proportions. For example, the simulation data set may include 30% of the first simulation data, 30% of the second simulation data, and 40% of the third simulation data. Figure 4 , Figure 4 This is an example diagram of a simulation data set provided in an embodiment of the present application.

[0046] S202, collecting a user's registered audio data set, and extracting the user's reference voiceprint vector from the registered audio data set;

[0047] Specifically, when a user uses an electronic device for the first time, the user's registered audio data set is collected, and the user's reference voiceprint vector is extracted from the registered audio data set. The registered audio data set includes at least one registered audio data, which is used to determine whether the user has control authority over the electronic device. The reference voiceprint vector is a fused voiceprint vector of common characteristics extracted from the user's registered audio data set collected in advance. The voiceprint vector is a numerical vector of the voice characteristics that can uniquely identify the user, which is extracted by analyzing and processing the audio data. Optionally, when the user enters the registered audio data, a prompt message prompts the user to choose to enter multiple audio data in different contexts and different ways of speaking, so as to further enhance the representation ability of the reference voiceprint vector.

[0048] When extracting a reference voiceprint vector, the voiceprint vector of each registered audio data set is extracted from the registered audio data set. The reference voiceprint vector is then fused based on the features of the voiceprint vectors of each registered audio data set using a cross-attention mechanism. The cross-attention mechanism is used to process the correlation between two different voiceprint vectors. Its core idea is to dynamically fuse the information of one voiceprint vector with another by calculating the attention weight of the other.

[0049] S203, collecting target audio data of the user in the application scenario, performing time-frequency transformation on the target audio data to obtain a first time-frequency graph of the target audio data, and mapping the first time-frequency graph into target audio features based on a neural network model;

[0050] Specifically, the electronic device collects target audio data from the user in the application scenario, inputs the target audio data into the neural network model, and preprocesses the target audio data in the audio encoder of the neural network model. After the amplitude of the target audio data is normalized, the target audio data is framed and time-frequency transformed to generate a first time-frequency graph of the target audio data. The first time-frequency graph can be a spectrogram or a mel-spectrogram. Based on a deep neural network, the first time-frequency graph is mapped to target audio features. The target audio features are high-dimensional latent features in the target audio data, namely, the user audio features present in the target audio data, as well as possible environmental noise features and human voice noise features.

[0051] The application scenario is a scenario where a user voice controls an electronic device or makes a voice call through an electronic device. For example, if the electronic device is a robot vacuum cleaner in a smart home device, the application scenario can be a scenario where the user voice controls the robot vacuum cleaner to start cleaning or switch cleaning modes. Assume that the user issues the command "Automatically start cleaning the whole house at 3 pm every day." When the user issues the command, there is ambient noise and other people talking around. The target audio data collected by the robot vacuum cleaner is "Automatically start cleaning the whole house at 3 pm every day + ambient noise + other people talking except the user."

[0052] S204, determining, from the target audio features, a user audio feature that matches a pre-generated reference voiceprint vector of the user through a cross-attention mechanism in the neural network model;

[0053] Specifically, after the audio encoder obtains the target audio features, it inputs them into a feature extractor, which pre-stores the user's reference voiceprint vector as prior information. In the feature extractor, the target audio features are combined with the reference voiceprint vector using a cross-attention mechanism to determine the user audio features in the target audio features that match the reference voiceprint vector. The reference voiceprint vector contains the audio pattern of the user's speech, such as fundamental frequency and formant distribution. The cross-attention mechanism is used to calculate the similarity between the target audio features and the reference voiceprint vector, and identify the time-frequency units in the target audio features that match the audio pattern of the reference voiceprint vector.

[0054] S205, enhancing the feature weight of the user audio feature to separate the user audio feature from the target audio feature through a mask matrix;

[0055] Specifically, after determining the user audio features, the feature weights of the user audio features are enhanced to separate the user audio features from the target audio features through a mask matrix. The mask matrix is a method for distinguishing audio and noise components in the time-frequency domain, and enhancing audio features and suppressing noise features through weighted operations. When training the neural network model, the feature extractor in the neural network model is controlled to learn the mask matrix so that when obtaining the user audio features, a mask matrix with the same dimension as the first time-frequency graph is provided, and the value of each element of the mask matrix is set between 0 and 1, indicating the retention ratio of the user audio features in the time-frequency unit corresponding to the first time-frequency graph, so as to separate the user audio features from the target audio features. If the value of an element in the mask matrix is close to 1, the time-frequency unit corresponding to the element is the user audio feature, which is retained; if the value of an element in the mask matrix is close to 0, the time-frequency unit corresponding to the element is the noise feature, which is suppressed.

[0056] S206: Perform inverse mapping on the user audio features based on the neural network model to obtain a second time-frequency graph of the user audio features, and perform inverse time-frequency transformation on the second time-frequency graph to obtain user audio data of the user.

[0057] Specifically, the user audio features are input into a feature decoder, which has a symmetrical network structure with the audio encoder. In the feature decoder, the user audio features are inversely mapped to generate a second time-frequency map of the user audio features. This second time-frequency map is then inversely transformed to convert it into a time-domain signal. The time-domain signals are then overlapped and added to obtain the user audio data. It is understood that the user audio data may contain residual noise or large volume fluctuations. High-frequency residual noise can be removed using a filter, and volume fluctuations can be balanced using a limiter or dynamic range controller.

[0058] See Figure 5 , Figure 5This is a flow chart of an audio acquisition method provided by an embodiment of the present application. Figure 5 As shown, after the electronic device collects the user's target audio data, it inputs the target audio data into an audio encoder to obtain target audio features. The target audio features are then input into a feature extractor that pre-stores the user's reference voiceprint vector. The feature extractor matches the reference voiceprint vector with the target audio features to separate the user's audio features. The user audio features are then input into a feature decoder, which decodes the user audio features and outputs the user audio data.

[0059] Optionally, the embodiment of the present application may adopt a distributed computing method to assign some computing tasks to a cloud server for processing. Distributed computing is a computing method that distributes large computing tasks into multiple subtasks. Distributed computing can effectively improve resource utilization efficiency and cope with various complex computing scenarios. At the same time, it has a high fault tolerance rate. When a node fails, other nodes can take over the processing task, ensuring the stability of the system. The cloud server and the electronic device communicate through a wireless connection, and can receive data sent by the electronic device, and store, process and analyze the received data. In the embodiment of the present application, audio acquisition can be used as the overall task, and the task assigned to the electronic device is to collect target audio data and send the target audio data to the cloud server. The task of the cloud server is to store the reference voiceprint vector, process the target audio data after receiving the target audio data, and obtain the user audio data from the target audio data based on the reference voiceprint vector. Through the collaborative work of the cloud server and the local electronic device, both real-time performance and computing efficiency are guaranteed.

[0060] Optionally, the embodiment of the present application supports multi-user scenarios. When the target audio data includes user audio data of multiple registered users, the electronic device can distinguish the voiceprint vectors of different users, realize personalized voice noise reduction for multiple users, and identify the voice control instructions of different users, thereby ensuring the use effect of the electronic device. Exemplarily, the current application scenario includes registered users D and E, as well as unregistered user F. The target audio data collected by the electronic device includes user D audio data d of user D, user E audio data e of user E, and audio noise data f and environmental noise data g of F. The electronic device extracts user D audio data d from the target audio data based on the pre-stored reference voiceprint vector m of user D, and extracts user E audio data e from the target audio data based on the pre-stored reference voiceprint vector n of user E, and then determines whether user D audio data d and user E audio data e are control instructions for them. If user D's audio data d and user E's audio data e are both voice control instructions for electronic devices, the electronic device will execute the voice control instructions of user D and user E in sequence according to the time sequence of user D's audio data d and user E's audio data e in the target audio data; if user D's audio data d or user E's audio data e is a voice control instruction for the electronic device, the instruction will be executed directly.

[0061] In a feasible embodiment, the electronic device obtains the user's user audio data set under multiple control scenarios of the user. In order to ensure the real-time performance of the reference voiceprint vector and improve the representation ability of the reference voiceprint vector, the user's reference voiceprint vector is updated based on the user's user audio data set at each preset time interval. Optionally, the electronic device can save the user's user audio data set obtained within the preset time interval in a local memory, or upload it to a cloud server and save the user's user audio data set to the cloud server to reduce the storage pressure of the electronic device. The electronic device obtains the user's audio data set within the preset time interval at each preset time interval, determines the user audio data set as the registered audio data set, extracts the user's reference voiceprint vector from the registered audio data set according to the method of step S202, deletes the user's reference voiceprint vector in the feature extractor, and inputs the new reference voiceprint vector into the feature extractor.

[0062] In the embodiment of the present application, the neural network model is trained by using simulation data including complex noise, thereby improving the generalization ability and robustness of the neural network model and enabling the neural network model to be used in complex and varied application scenarios. A fused voiceprint vector with common characteristics is extracted from the pre-collected user registration audio data set, and used as the user's reference voiceprint vector. The vector is matched with the target audio features in the target audio data to obtain the user's audio features, and then the user's audio data is obtained. This improves the representation ability of the voiceprint vector and more accurately guides the neural network model to distinguish between user audio data and noise, thereby improving the noise reduction effect in the audio acquisition process, reducing the dependence on a single high-quality registration audio data, and ensuring the user's experience.

[0063] See Figure 6 , Figure 6 This is a flow chart of a method for extracting a reference voiceprint vector provided in an embodiment of the present application. Figure 6 As shown, the method of the embodiment of the present application may include the following steps S301-S305, and steps S301-S304 can be used as Figure 2 The detailed steps of step S202 of the embodiment are shown.

[0064] S301, extracting the voiceprint features of each registered audio data in the registered audio data set;

[0065] Specifically, when extracting voiceprint vectors, the voiceprint features of each registered audio data set are first extracted. Voiceprint features can be used in multiple fields such as speech recognition, audio classification, and voiceprint classification. The voiceprint features extracted in the embodiments of the present application include, but are not limited to, Mel-Frequency Cepstral Coefficients (MFCC), Mel-scale Filter Bank Coefficients (FBank), and Linear Predictive Coding (LPC).

[0066] Among them, MFCC simulates the sensitivity of the human auditory system to different frequencies, represents the audio data on the Mel frequency scale, and calculates the cepstral coefficients of the spectrum to capture the spectral envelope of the audio data. The Mel frequency scale is a nonlinear frequency scale used to simulate the human ear's ability to perceive sounds of different frequencies, that is, the human ear is more sensitive to low-frequency sounds than high-frequency sounds. FBank is the output value of a set of filter banks designed based on the Mel frequency scale. By designing a series of filters uniformly distributed on the Mel frequency scale, the energy distribution of audio data in different frequency bands can be captured, thereby extracting useful voiceprint features. LPC estimates the spectral envelope and resonance peak frequency of audio data by performing linear prediction modeling on audio data, thereby achieving effective compression and representation of audio data.

[0067] S302, mapping the voiceprint features of each registered audio data to the same linear space by sharing parameters to obtain a voiceprint vector of each registered audio data;

[0068] Specifically, after the voiceprint features are extracted, the voiceprint features are mapped into a linear space to obtain a voiceprint vector. The voiceprint vectors extracted in the embodiment of the present application are used for fusion to obtain a reference voiceprint vector. Therefore, when mapping the voiceprint features, the voiceprint features of each registered audio data are mapped into the same linear space through shared parameters to obtain a voiceprint vector for each registered audio data. Shared parameters are learnable or preset mapping parameters in the mathematical transformation that converts the original voiceprint features into low-dimensional voiceprint vectors. Mapping through shared parameters can not only reduce the number of parameters when extracting the voiceprint vector, but also avoid the situation where the training cannot converge during the subsequent voiceprint vector fusion due to excessive distribution differences between different registered audio data.

[0069] S303, concatenating the voiceprint vectors of each registered audio data in the registered audio data set according to the feature dimension to obtain a first feature matrix;

[0070] Specifically, the characteristic dimension of a voiceprint vector refers to a set of parameters that mathematically represent the physical acoustic features contained in the audio data, such as 512 or 1024 dimensions. When fusing voiceprint vectors, the voiceprint vectors of each registered audio data set are first concatenated based on the characteristic dimension to produce a first characteristic matrix. A characteristic matrix is a fundamental data structure for storing data features. In the embodiments of this application, the first characteristic matrix is a two-dimensional array, with rows representing the voiceprint vectors of each registered audio data point, and columns representing the characteristic dimensions of the voiceprint vectors.

[0071] S304, a cross attention module based on the cross attention mechanism performs attention calculation on the first feature matrix to obtain a second feature matrix;

[0072] Specifically, the reference voiceprint vector is a fused voiceprint vector that fuses the relevant information of the voiceprint vector of each registered audio data. The cross-attention mechanism in the voiceprint vector fusion module provided in the embodiment of the present application includes two cross-attention modules. The first cross-attention module is used to capture the global correlation between voiceprint vectors (such as the commonality of timbre and pronunciation habits), and the second cross-attention module is used to refine local feature associations (such as the stability of specific phonemes) to enhance the robustness of the cross-attention mechanism. The first attention calculation is performed on the first feature matrix by the first cross-attention module to obtain the third feature matrix, and then the second attention calculation is performed on the third feature matrix by the second cross-attention module to obtain the second feature matrix. The cross-attention mechanism is used to process the correlation between two different voiceprint vectors. The core idea is to dynamically fuse the information of the two by calculating the attention weight of one voiceprint vector to another voiceprint vector.

[0073] S305: Convolve the second feature matrix to obtain a comprehensive voiceprint vector.

[0074] Specifically, the voiceprint vector fusion module also includes a downsampling layer, which performs convolution on the second feature matrix, compressing the high-dimensional features of the second feature matrix into a single-dimensional reference voiceprint vector, reducing redundant information. The downsampling layer provides effective input for the convolution operation by reducing feature dimensions and retaining key information. The convolution operation, through local perception and parameter sharing, extracts and fuses key features from the voiceprint vector, ultimately generating a robust reference voiceprint vector.

[0075] See Figure 7 , Figure 7 This is a flow chart of a method for extracting a reference voiceprint vector provided in an embodiment of the present application. Figure 7 As shown, multiple voiceprint vectors of multiple registered audio data are extracted through the voiceprint vector extraction module, and the multiple voiceprint vectors are fused through the first cross-attention module, the second cross-attention module and the downsampling layer to obtain a reference voiceprint vector.

[0076] In an embodiment of the present application, the electronic device extracts the voiceprint features of each registered audio in the user's registered audio data set. When obtaining the voiceprint vector based on the voiceprint features, different voiceprint features are mapped to the same linear space by sharing parameters, which reduces the number of parameters when extracting the voiceprint vector. At the same time, it avoids the situation where the training cannot converge during the subsequent voiceprint vector fusion due to excessive distribution differences between different registered audio data. Based on the attention calculation of the two cross-attention modules and the convolution of the downsampling layer, a reference voiceprint vector is obtained, which focuses on and retains the common features in the voiceprint vector, making the reference voiceprint more representative and improving the accuracy of the neural network model in extracting user audio features from the target audio features. By optimizing the voiceprint vector extraction and voiceprint vector fusion process, the computational complexity is reduced, so that the embodiment of the present application can be implemented in electronic devices with limited resources, expanding its scope of application.

[0077] Regarding the above-mentioned method of separating target audio features through a mask matrix, the embodiments of the present application also disclose the following two implementation methods.

[0078] In the first possible implementation, see Figure 8 , Figure 8 FIG. 1 is a flow chart of a method for determining a mask matrix provided in an embodiment of the present application. Figure 8 As shown, the method of the embodiment of the present application may include the following steps S401-S403.

[0079] S401, obtaining a target volume and / or a target noise type of target audio data, and determining a weight parameter corresponding to the target audio data based on the target volume and / or the target noise type;

[0080] Specifically, due to the complex and variable nature of the noise environment, an adaptive learning mechanism is introduced into the neural network model to automatically adjust the model parameters based on noise changes. A volume detection module and / or a noise detection module are added after the audio encoder of the neural network model. The volume detection module obtains the target volume of the target audio data, and the noise detection module obtains the target noise type of the target audio data. Based on the target volume and / or target noise type, the weight parameters corresponding to the target audio data are determined and assigned to the neural network model.

[0081] Among them, the adaptive learning mechanism is a mechanism that automatically adjusts the model parameters according to new data or environmental changes. The audio detection module realizes volume detection by quantifying the energy intensity of the target audio data. After the audio encoder divides the target audio data into frames, it calculates the root mean square value of the audio signal of each frame, converts it into a decibel value, and normalizes the decibel value to obtain the target volume of the target audio data. The target volume can be a specific numerical value, or it can be a volume level determined according to a pre-set volume interval, such as high, medium, and low. After the audio encoder generates the first time-frequency graph of the target audio data, the noise detection module extracts the energy distribution of each frequency band in the first time-frequency graph, and uses the trained noise classification model to classify the noise type of the target audio data based on the reference noise vector. Noise types include but are not limited to white noise, traffic noise, human voice noise, and transient noise.

[0082] S402, generating a first mask matrix and a second mask matrix in the neural network model, where the first mask matrix is used to comprehensively suppress noise, and the second mask matrix is used to suppress noise overlapping with the user audio features;

[0083] Specifically, two types of mask matrices are pre-generated in the neural network model. The first mask matrix is used to comprehensively suppress noise and is suitable for high-volume scenes or high-noise scenes to enhance the noise reduction intensity; the second mask matrix is used to suppress noise overlapping with the user's audio features and is suitable for low-volume scenes or low-noise scenes to reduce the noise reduction intensity and thus retain the user's voice details.

[0084] S403 : Generate a third mask matrix based on the weight parameter, the first mask matrix, and the second mask matrix, and separate the user audio features from the target audio features by using the third mask matrix.

[0085] Exemplarily, if the weight parameter is k, the first mask matrix is X, the second mask matrix is Y, and the third mask matrix is Z, then Z=k·X+(1-k)·Y, and the user audio features are separated from the target audio features according to the generated third mask matrix.

[0086] In an embodiment of the present application, weight parameters are generated according to the target volume and / or target noise type, and the first mask matrix for comprehensive noise suppression and the second mask matrix for partial noise suppression are dynamically fused based on the weight parameters, thereby balancing the noise reduction intensity and audio quality in the target audio data and avoiding audio loss caused by noise reduction.

[0087] In the second possible implementation, see Figure 9 , Figure 9 FIG. 1 is a flow chart of a method for determining a mask matrix provided in an embodiment of the present application. Figure 9 As shown, the method of the embodiment of the present application may include the following steps S501-S502.

[0088] S501, obtaining a target volume and / or a target noise type of target audio data, and determining a target audio type of the target audio data based on the target volume and / or the target noise type;

[0089] Due to the complexity and variability of the noise environment, an adaptive learning mechanism is introduced into the neural network model to automatically adjust the parameters of the neural network model according to noise changes. A volume detection module and / or a noise detection module are added after the audio encoder of the neural network model. The target volume of the target audio data is obtained through the volume detection module, and the target noise type of the target audio data is obtained through the noise detection module. Based on a pre-set first mapping relationship between volume and audio type, and / or a second mapping relationship between noise type and audio type, the target audio type of the target audio data corresponding to the target volume and / or target noise type is determined.

[0090] Among them, the adaptive learning mechanism is a mechanism that automatically adjusts the model parameters according to new data or environmental changes. The audio detection module realizes volume detection by quantifying the energy intensity of the target audio data. After the audio encoder divides the target audio data into frames, it calculates the root mean square value of the audio signal of each frame, converts it into a decibel value, and normalizes the decibel value to obtain the target volume of the target audio data. The target volume can be a specific numerical value, or it can be a volume level determined according to a pre-set volume interval, such as high, medium, and low. After the audio encoder generates the first time-frequency graph of the target audio data, the noise detection module extracts the energy distribution of each frequency band in the first time-frequency graph, and uses the trained noise classification model to classify the noise type of the target audio data based on the reference noise vector. Noise types include but are not limited to white noise, traffic noise, human voice noise, and transient noise.

[0091] S502: Determine a target mask matrix corresponding to the target audio data based on the target audio type, and separate the user audio features from the target audio features using the target mask matrix.

[0092] Specifically, multiple noise reduction pathways are pre-set in the neural network model, each corresponding to a different mask matrix to process different types of audio data. Based on the target audio type, the target noise reduction pathway corresponding to the target audio data is determined. The mask matrix in this target noise reduction pathway is then used as the target mask matrix, which is used to separate the user audio features from the target audio features.

[0093] For example, assuming that audio data is classified according to noise type, three noise reduction pathways are pre-set in the neural network model. The first noise reduction pathway is a noise reduction pathway for the first audio type corresponding to white noise, and the mask matrix in the first noise reduction pathway is the fourth mask matrix. The second noise reduction pathway is a noise reduction pathway for the second audio type corresponding to human voice noise, and the mask matrix in the second noise reduction pathway is the fifth mask matrix. The third noise reduction pathway is a noise reduction pathway for the third audio type corresponding to transient noise, and the mask matrix in the third noise reduction pathway is the sixth mask matrix. If the electronic device is a vacuum cleaner in a smart home device, and the noise data in the target audio data is the vocal noise of the user's family, then the target audio type of the target audio data is the second audio type corresponding to human voice noise, and the fifth mask matrix in the second noise reduction pathway is selected to separate the target audio features from the target audio data.

[0094] In an embodiment of the present application, the target audio type of the target audio data is determined based on the target volume and / or target noise type, and then the corresponding target mask matrix is determined based on the target audio type, thereby balancing the noise reduction intensity and audio quality in the target audio data and avoiding audio loss caused by noise reduction.

[0095] See Figure 10 , Figure 10 FIG. 1 is a flow chart of a method for collecting target audio data provided by an embodiment of the present application. Figure 10 As shown, the method of the embodiment of the present application may include the following steps S601 to S604.

[0096] S601, collecting registration audio data input by the user and facial image data of the user when inputting the registration audio data;

[0097] Specifically, the electronic device includes a sound receiving module and an image acquisition module. When a user uses the electronic device for the first time, the sound receiving module collects the registration audio data input by the user, and the image acquisition module collects the facial image data of the user when entering the registration audio data. The facial image data includes image data such as lip movement information and facial expressions when the user speaks. Optionally, when the user enters the registration audio data, a prompt message prompts the user to choose to enter multiple audio data in different contexts and different voice modes, so as to improve the representation ability of the reference voiceprint vector extracted subsequently. At the same time, when the user enters audio data in different contexts and different voice modes, the facial image data of the user collected is also more representative, thereby improving the representation ability of the reference facial image extracted subsequently. Among them, the reference voiceprint vector is a fused voiceprint vector of common characteristics extracted from a registration audio data set composed of the user's registration audio data collected in advance, and the reference facial image is a fused facial image of common characteristics extracted from a facial image data set composed of the user's facial image data collected in advance.

[0098] S602, determining a registered audio data set based on the registered audio data, and determining a facial image data set based on the facial image data;

[0099] S603, extracting a user's reference voiceprint vector from the registered audio data set, and extracting a user's reference facial image from the facial image data set;

[0100] Specifically, the voiceprint features of each registered audio data set are extracted from the registered audio data set, and the voiceprint features of each registered audio data set are mapped to the same linear space through shared parameters to obtain the voiceprint vector of each registered audio data set. When fusing the voiceprint vectors, the voiceprint vectors of each registered audio data set in the registered audio data set are first concatenated according to the feature dimension to obtain a first feature matrix. The cross-attention module based on the cross-attention mechanism performs attention calculation on the first feature matrix to obtain a second feature matrix. The second feature matrix is convolved based on the downsampling layer to compress the high-dimensional features of the second feature matrix into a single-dimensional reference voiceprint vector to reduce redundant information.

[0101] Each facial image in the facial image dataset is preprocessed to locate the facial regions, including the eyes, mouth, and nose, and align multiple facial images. Facial features are extracted from each facial image to obtain multiple facial feature vectors. These multiple facial feature vectors are weighted averaged to generate a reference facial feature vector, which is then reconstructed to obtain a reference facial image.

[0102] Among them, voiceprint features can be used in multiple fields such as speech recognition, audio classification and voiceprint classification. The voiceprint features extracted by the embodiment of the present application include but are not limited to features such as MFCC, FBank and LPC. MFCC simulates the sensitivity of the human ear's auditory system to different frequencies, represents the audio data on the Mel frequency scale, and calculates the cepstral coefficients of the spectrum, thereby capturing the spectral envelope of the audio data. The Mel frequency scale is a nonlinear frequency scale used to simulate the human ear's perception of sounds of different frequencies, that is, the human ear is more sensitive to low-frequency sounds than high-frequency sounds. FBank is the output value of a group of filter groups designed based on the Mel frequency scale. By designing a series of filters uniformly distributed on the Mel frequency scale, the energy distribution of audio data in different frequency bands can be captured, thereby extracting useful voiceprint features. LPC estimates the spectral envelope and resonance peak frequency of audio data by performing linear prediction modeling on audio data, thereby achieving effective compression and representation of audio data. Shared parameters are learnable or preset mapping parameters in the mathematical transformation of converting the original voiceprint features into low-dimensional voiceprint vectors. Mapping through shared parameters can not only reduce the number of parameters when extracting voiceprint vectors, but also avoid the situation where different registered audio data cannot converge during subsequent voiceprint vector fusion due to excessive distribution differences. The feature matrix is a basic data structure for storing data features. The first feature matrix in the embodiment of the present application is a two-dimensional array. The rows of the first feature matrix are used to represent the voiceprint vectors of each registered audio data, and the columns are used to represent the feature dimensions of the voiceprint vectors. The cross-attention mechanism in the voiceprint vector fusion module provided in the embodiment of the present application includes two cross-attention modules. The first cross-attention module is used to capture the global correlation between voiceprint vectors (such as the commonality of timbre and pronunciation habits), and the second cross-attention module is used to refine local feature associations (such as the stability of specific phonemes) to enhance the robustness of the cross-attention mechanism. The first attention calculation is performed on the first feature matrix by the first cross-attention module to obtain a third feature matrix, and then the second attention calculation is performed on the third feature matrix by the second cross-attention module to obtain a second feature matrix. The cross-attention mechanism is used to handle the correlation between two different voiceprint vectors. Its core idea is to dynamically fuse the information of one voiceprint vector with another by calculating the attention weight of one voiceprint vector. The downsampling layer provides effective input for the convolution operation by reducing feature dimensions and retaining key information. The convolution operation, through local perception and parameter sharing, extracts and fuses key features from the voiceprint vectors, ultimately generating a robust reference voiceprint vector.

[0103] S604: Determine the target position of the user in the application scenario based on the user's reference facial image, enhance the sound reception weight at the target position, and collect target audio data of the user in the application scenario.

[0104] Specifically, when the electronic device detects that the user has issued a voice command, the user's reference facial image is matched with the face of the person appearing in the application scene to determine the user's target position in the application scene, thereby enhancing the sound reception weight at the target position, collecting the user's target audio data in the application scene, enhancing the volume and clarity of the user's audio data in the target audio data, and reducing noise interference. Among them, the application scenario is a scenario in which the user voice controls the electronic device or a scenario in which the user makes a voice call through the electronic device. For example, if the electronic device is a vacuum cleaner in a smart home device, the application scenario can be a scenario in which the user controls the vacuum cleaner by voice to start cleaning or switch the cleaning mode. Assuming that the command issued by the user is "automatically start cleaning the whole house at 3 pm every day", when the user issues the command, there is ambient noise and other people's voices around, then the target audio data collected by the vacuum cleaner is "automatically start cleaning the whole house at 3 pm every day + ambient noise + other people's voices except the user".

[0105] In this embodiment of the present application, when a user inputs registration audio data, the electronic device collects the user's facial image data, extracts the user's reference facial image from the user's facial image data set, and then, based on the reference facial image, enhances the sound reception weight at the user's target location in the application scenario. Through multimodal fusion, the volume and clarity of the user's audio data within the target audio data are enhanced, noise interference in the user's audio data is reduced, and the performance of the electronic device is guaranteed.

[0106] based on Figure 1 The following scene diagram will be combined with Figure 11 , the audio acquisition device provided in the embodiment of the present application is introduced in detail. It should be noted that, Figure 11 The audio acquisition device in this application is used to execute Figure 2-Figure 10 For the convenience of explanation, only the part related to the embodiment of the present application is shown. For the specific technical details not disclosed, please refer to the present application. Figure 2-Figure 10 The embodiment shown.

[0107] See Figure 11 , Figure 11 This is a structural diagram of an audio acquisition device provided in an embodiment of the present application. Figure 11 As shown, the audio acquisition device 1 of the embodiment of the present application may include: a target audio feature acquisition unit 11, a user audio feature acquisition unit 12 and a user audio data generation unit 13.

[0108] The target audio feature acquisition unit 11 is used to collect the target audio data of the user in the application scenario and obtain the target audio features of the target audio data;

[0109] A user audio feature acquisition unit 12 is configured to acquire the user audio feature of the user from the target audio feature based on a pre-generated reference voiceprint vector of the user;

[0110] A user audio data generating unit 13 is configured to generate user audio data of a user based on the user audio features;

[0111] The reference voiceprint vector is a fused voiceprint vector of common features extracted from a pre-collected set of registered audio data of the user.

[0112] Optionally, the audio acquisition device 1 is specifically used to collect the user's registered audio data set;

[0113] Extract the user's reference voiceprint vector from the registered audio data set.

[0114] Optionally, the audio acquisition device 1 is specifically configured to extract a voiceprint vector of each registered audio data in the registered audio data set;

[0115] Based on the cross-attention mechanism, the voiceprint vector of each registered audio data in the registered audio data set is feature fused to obtain a reference voiceprint vector.

[0116] Optionally, the audio acquisition device 1 is specifically used to extract the voiceprint features of each registered audio data in the registered audio data set;

[0117] The voiceprint features of each registered audio data are mapped to the same linear space by sharing parameters to obtain the voiceprint vector of each registered audio data.

[0118] Optionally, the audio acquisition device 1 is specifically configured to concatenate the voiceprint vectors of each registered audio data in the registered audio data set according to the feature dimension to obtain a first feature matrix;

[0119] The cross attention module based on the cross attention mechanism performs attention calculation on the first feature matrix to obtain the second feature matrix;

[0120] Convolve the second feature matrix to obtain the comprehensive voiceprint vector.

[0121] Optionally, the target audio feature acquisition unit 11 is specifically configured to perform time-frequency transformation on the target audio data to obtain a first time-frequency graph of the target audio data;

[0122] The first time-frequency image is mapped into target audio features based on a neural network model.

[0123] Optionally, the user audio feature acquisition unit 12 is specifically configured to determine, in the target audio feature, a user audio feature that matches a pre-generated reference voiceprint vector of the user through a cross-attention mechanism in the neural network model;

[0124] The feature weights of the user audio features are enhanced to separate the user audio features from the target audio features through a mask matrix.

[0125] Optionally, the user audio data generating unit 13 is specifically configured to perform inverse mapping on the user audio features based on a neural network model to obtain a second time-frequency graph of the user audio features;

[0126] An inverse time-frequency transform is performed on the second time-frequency graph to obtain user audio data of the user.

[0127] Optionally, the audio acquisition device 1 is specifically used to acquire a simulation data set and train a neural network model based on the simulation data set.

[0128] Optionally, the audio acquisition device 1 is specifically used to acquire noise data and human voice audio, where the noise data includes environmental noise and human voice noise;

[0129] Determine a simulation data set based on the noise data and the human voice audio, the simulation data set including first simulation data, second simulation data, and third simulation data;

[0130] Among them, the first simulation data is the synthesized audio data of environmental noise and human voice audio, the second simulation data is the synthesized audio data of human voice noise and human voice audio, and the third simulation data is the synthesized audio data of environmental noise, human voice noise and human voice audio.

[0131] Optionally, the target audio feature acquisition unit 11 is specifically configured to collect registration audio data input by the user, and facial image data of the user when inputting the registration audio data;

[0132] A registration audio data set is determined based on the registration audio data, and a facial image data set is determined based on the facial image data.

[0133] Optionally, the audio acquisition device 1 is specifically configured to extract a reference facial image of a user from a facial image data set.

[0134] Optionally, the audio acquisition device 1 is specifically configured to determine a target position of the user in the application scenario based on a reference facial image of the user;

[0135] Increase the reception weight at the target direction.

[0136] Optionally, the audio acquisition device 1 is specifically configured to update the user's reference voiceprint vector based on the user's audio data at intervals of a preset duration.

[0137] Optionally, the audio acquisition device 1 is specifically used to acquire a target volume and / or a target noise type of target audio data;

[0138] A weight parameter corresponding to the target audio data is determined based on the target volume and / or the target noise type.

[0139] Optionally, the user audio feature acquisition unit 12 is specifically configured to generate a first mask matrix and a second mask matrix in the neural network model, wherein the first mask matrix is used to comprehensively suppress noise, and the second mask matrix is used to suppress noise overlapping with the user audio feature;

[0140] generating a third mask matrix based on the weight parameters, the first mask matrix, and the second mask matrix;

[0141] The user audio features are separated from the target audio features by a third mask matrix.

[0142] Optionally, the audio acquisition device 1 is specifically configured to determine a target audio type of the target audio data based on a target volume and / or a target noise type.

[0143] Optionally, the user audio feature acquisition unit 12 is specifically configured to determine a target mask matrix corresponding to the target audio data based on the target audio type;

[0144] The user audio features are separated from the target audio features through the target mask matrix.

[0145] In the embodiments of the present application, the neural network model is trained using simulated data including complex noise, thereby improving the generalization and robustness of the neural network model and enabling the neural network model to be used in complex and diverse application scenarios. A fused voiceprint vector with common characteristics is extracted from a pre-collected set of registered audio data of the user and used as the user's reference voiceprint vector. This is then matched with the target audio features in the target audio data to obtain the user's audio features, and then the user's audio data is obtained. This improves the representational capability of the voiceprint vector and, in turn, enhances the noise reduction effect during the audio acquisition process.

[0146] See Figure 12 , Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0147] For example, Figure 12 As shown, the electronic device 1200 includes: a processor 1201 and a memory 1202, wherein the processor 1201 is electrically connected to the memory 1202.

[0148] The processor 1201 is the control center of the electronic device 1200 and may include one or more processing cores. The processor 1201 utilizes various interfaces and lines to connect the various parts of the entire electronic device. By running or calling computer programs stored in the memory 1202, and calling data stored in the memory 1202, the processor 1201 executes various functions of the electronic device and processes data, thereby performing overall control over the electronic device 1200. Optionally, the processor 1201 may be implemented in at least one hardware form of a digital signal processing (DSP), a field programmable gate array (FPGA), or a programmable logic array (PLA). The processor 1201 may integrate one or a combination of a CPU, a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user pages, and applications; the GPU is responsible for rendering and drawing display content; and the modem is used to handle wireless communications. It is understandable that the above-mentioned modem may not be integrated into the processor 1201 and may be implemented separately through a communication chip.

[0149] The memory 1202 can be used to store software programs and modules. The processor 1201 executes various functional applications and data processing by running the computer programs and modules stored in the memory 1202. The memory 1202 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, computer programs required for at least one function, etc.; the data storage area may store data generated based on the use of the electronic device 1200.

[0150] In addition, the memory 1202 may include a high-speed random access memory and a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 1202 may also include a memory controller to provide the processor 1201 with access to the memory 1202.

[0151] In this embodiment, the processor 1201 in the electronic device 1200 loads instructions corresponding to one or more computer program processes into the memory 1202 according to the following steps, and the processor 1201 runs the computer program stored in the memory 1202 to implement various functions as follows:

[0152] Collect target audio data of the user in the application scenario and obtain target audio features of the target audio data;

[0153] Based on a pre-generated reference voiceprint vector of the user, obtaining a user audio feature of the user from the target audio feature;

[0154] generating user audio data of the user based on the user audio features;

[0155] The reference voiceprint vector is a fused voiceprint vector of common features extracted from a pre-collected set of registered audio data of the user.

[0156] Optionally, before collecting target audio data of the user in the application scenario, the processor 1201 further executes:

[0157] Collecting user's registered audio data set;

[0158] Extract the user's reference voiceprint vector from the registered audio data set.

[0159] Optionally, when extracting the user's reference voiceprint vector from the registration audio data set, the processor 1201 specifically performs:

[0160] Extracting a voiceprint vector of each registered audio data from the registered audio data set;

[0161] Based on the cross-attention mechanism, the voiceprint vector of each registered audio data in the registered audio data set is feature fused to obtain a reference voiceprint vector.

[0162] Optionally, when extracting the voiceprint vector of each registered audio data from the registered audio data set, the processor 1201 specifically performs:

[0163] Extracting voiceprint features of each registered audio data in the registered audio data set;

[0164] The voiceprint features of each registered audio data are mapped to the same linear space by sharing parameters to obtain the voiceprint vector of each registered audio data.

[0165] Optionally, when the processor 1201 performs feature fusion on the voiceprint vector of each registered audio data in the registered audio data set based on the cross-attention mechanism to obtain the reference voiceprint vector, the processor 1201 specifically performs:

[0166] splicing the voiceprint vectors of each registered audio data in the registered audio data set according to the feature dimension to obtain a first feature matrix;

[0167] The cross attention module based on the cross attention mechanism performs attention calculation on the first feature matrix to obtain the second feature matrix;

[0168] Convolve the second feature matrix to obtain the comprehensive voiceprint vector.

[0169] Optionally, when executing the step of obtaining the target audio features of the target audio data, the processor 1201 specifically performs:

[0170] Performing a time-frequency transformation on the target audio data to obtain a first time-frequency graph of the target audio data;

[0171] The first time-frequency image is mapped into target audio features based on a neural network model.

[0172] Optionally, when the processor 1201 obtains the user audio feature of the user from the target audio feature based on the pre-generated reference voiceprint vector of the user, it specifically performs:

[0173] Determine the user audio features that match the pre-generated user reference voiceprint vector in the target audio features through the cross-attention mechanism in the neural network model;

[0174] The feature weights of the user audio features are enhanced to separate the user audio features from the target audio features through a mask matrix.

[0175] Optionally, when generating user audio data of a user based on the user audio feature, the processor 1201 specifically performs:

[0176] Perform inverse mapping of the user audio features based on the neural network model to obtain a second time-frequency graph of the user audio features;

[0177] An inverse time-frequency transform is performed on the second time-frequency graph to obtain user audio data of the user.

[0178] Optionally, before collecting target audio data of the user in the application scenario, the processor 1201 further executes:

[0179] Obtain a simulation data set and train a neural network model based on the simulation data set.

[0180] Optionally, when executing the acquisition of the simulation data set, the processor 1201 specifically performs:

[0181] Acquire noise data and human voice audio, where the noise data includes environmental noise and human voice noise;

[0182] Determine a simulation data set based on the noise data and the human voice audio, the simulation data set including first simulation data, second simulation data, and third simulation data;

[0183] Among them, the first simulation data is the synthesized audio data of environmental noise and human voice audio, the second simulation data is the synthesized audio data of human voice noise and human voice audio, and the third simulation data is the synthesized audio data of environmental noise, human voice noise and human voice audio.

[0184] Optionally, when collecting the user's registered audio data set, the processor 1201 specifically performs:

[0185] Collecting the registration audio data input by the user and the facial image data of the user when inputting the registration audio data;

[0186] A registration audio data set is determined based on the registration audio data, and a facial image data set is determined based on the facial image data.

[0187] Optionally, after collecting the user's registered audio data set, the processor 1201 further executes:

[0188] A reference facial image of a user is extracted from a facial image dataset.

[0189] Optionally, before collecting target audio data of the user in the application scenario, the processor 1201 further executes:

[0190] Determining the target orientation of the user in the application scenario based on the user's reference facial image;

[0191] Increase the reception weight at the target direction.

[0192] Optionally, the processor 1201 further updates the user's reference voiceprint vector based on the user's audio data at intervals of a preset duration.

[0193] Optionally, after collecting target audio data of the user in the application scenario and obtaining target audio features of the target audio data, the processor 1201 further executes:

[0194] Obtaining a target volume and / or a target noise type of target audio data;

[0195] A weight parameter corresponding to the target audio data is determined based on the target volume and / or the target noise type.

[0196] Optionally, when separating the user audio feature from the target audio feature using the mask matrix, the processor 1201 specifically performs:

[0197] Generating a first mask matrix and a second mask matrix in the neural network model, wherein the first mask matrix is used to comprehensively suppress noise, and the second mask matrix is used to suppress noise overlapping with the user audio features;

[0198] generating a third mask matrix based on the weight parameters, the first mask matrix, and the second mask matrix;

[0199] The user audio features are separated from the target audio features by a third mask matrix.

[0200] Optionally, after obtaining the target volume and / or target noise type of the target audio data, the processor 1201 further executes:

[0201] A target audio type of the target audio data is determined based on the target volume and / or the target noise type.

[0202] Optionally, when separating the user audio feature from the target audio feature using the mask matrix, the processor 1201 specifically performs:

[0203] Determining a target mask matrix corresponding to the target audio data based on the target audio type;

[0204] The user audio features are separated from the target audio features through the target mask matrix.

[0205] In the embodiments of the present application, the neural network model is trained using simulated data including complex noise, thereby improving the generalization and robustness of the neural network model and enabling the neural network model to be used in complex and diverse application scenarios. A fused voiceprint vector with common characteristics is extracted from a pre-collected set of registered audio data of the user and used as the user's reference voiceprint vector. This is then matched with the target audio features in the target audio data to obtain the user's audio features, and then the user's audio data is obtained. This improves the representational capability of the voiceprint vector and, in turn, enhances the noise reduction effect during the audio acquisition process.

[0206] It should be understood that the device provided in the embodiment of the present application is used to execute the above-mentioned audio acquisition method, and thus can achieve the same effect as the above-mentioned implementation method.

[0207] In the case of an integrated unit, the device may include a processing module and a storage module. When the device is applied to an electronic device, the processing module may be used to control and manage the operation of the electronic device. The storage module may be used to support the electronic device in executing relevant program codes, etc.

[0208] The processing module may be a processor or controller that implements or executes the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processing system (DSP) and a microprocessor, and the storage module may be a memory.

[0209] In addition, the device provided in the embodiment of the present application can specifically be a chip, component or module, and the chip may include a connected processor and memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute an audio acquisition method provided in the above embodiment.

[0210] An embodiment of the present application also provides a computer-readable storage medium, which stores computer program code. When the computer program code is run on a computer, the computer executes the above-mentioned related method steps to implement an audio acquisition method provided in the above embodiment.

[0211] This embodiment further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement an audio acquisition method provided in the above embodiment.

[0212] Among them, the device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0213] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0214] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0215] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An audio acquisition method, characterized in that: The method comprises: Collecting target audio data of the user in the application scenario and obtaining target audio features of the target audio data; Based on a pre-generated reference voiceprint vector of the user, obtaining a user audio feature of the user from the target audio feature; generating user audio data of the user based on the user audio feature; The reference voiceprint vector is a fused voiceprint vector of common features extracted from a pre-collected set of registered audio data of the user.

2. The method according to claim 1, characterized in that Before collecting the target audio data of the user in the application scenario, the method further includes: Collecting a registered audio data set of the user; A reference voiceprint vector of the user is extracted from the registered audio data set.

3. The method according to claim 2, characterized in that Extracting the reference voiceprint vector of the user from the registered audio data set includes: Extracting a voiceprint vector of each registered audio data from the registered audio data set; Based on the cross-attention mechanism, feature fusion is performed on the voiceprint vector of each registered audio data in the registered audio data set to obtain a reference voiceprint vector.

4. The method according to claim 3, characterized in that Extracting the voiceprint vector of each registered audio data from the registered audio data set includes: Extracting a voiceprint feature of each of the registered audio data in the registered audio data set; The voiceprint features of each of the registered audio data are mapped to the same linear space through shared parameters to obtain a voiceprint vector of each of the registered audio data.

5. The method according to claim 3, characterized in that The step of performing feature fusion on the voiceprint vector of each registered audio data in the registered audio data set based on the cross-attention mechanism to obtain a reference voiceprint vector includes: splicing the voiceprint vectors of each registered audio data in the registered audio data set according to the feature dimension to obtain a first feature matrix; A cross-attention module based on the cross-attention mechanism performs attention calculation on the first feature matrix to obtain a second feature matrix; Convolution is performed on the second feature matrix to obtain a comprehensive voiceprint vector.

6. The method according to claim 1, characterized in that The acquiring the target audio feature of the target audio data includes: Performing a time-frequency transformation on the target audio data to obtain a first time-frequency graph of the target audio data; The first time-frequency graph is mapped into target audio features based on a neural network model.

7. The method according to claim 6, characterized in that The acquiring the user audio feature of the user from the target audio feature based on the pre-generated reference voiceprint vector of the user includes: Determining, from the target audio features, a user audio feature that matches a pre-generated reference voiceprint vector of the user through a cross-attention mechanism in the neural network model; The feature weight of the user audio feature is enhanced to separate the user audio feature from the target audio feature through a mask matrix.

8. The method according to claim 6, characterized in that Generating user audio data of the user based on the user audio feature includes: Performing inverse mapping on the user audio feature based on the neural network model to obtain a second time-frequency graph of the user audio feature; Perform an inverse time-frequency transform on the second time-frequency graph to obtain user audio data of the user.

9. The method according to claim 6, characterized in that Before collecting the target audio data of the user in the application scenario, the method further includes: A simulation data set is obtained, and the neural network model is trained based on the simulation data set.

10. The method according to claim 9, characterized in that The obtaining of the simulation data set comprises: Acquire noise data and human voice audio, wherein the noise data includes environmental noise and human voice noise; Determine a simulation data set based on the noise data and the human voice audio, the simulation data set including first simulation data, second simulation data, and third simulation data; Among them, the first simulation data is the synthesized audio data of the environmental noise and the human voice audio, the second simulation data is the synthesized audio data of the human voice noise and the human voice audio, and the third simulation data is the synthesized audio data of the environmental noise, the human voice noise and the human voice audio.

11. The method according to claim 2, characterized in that The collecting of the user's registered audio data set includes: collecting registration audio data input by the user and facial image data of the user when inputting the registration audio data; A registration audio data set is determined based on the registration audio data, and a facial image data set is determined based on the facial image data.

12. The method according to claim 11, characterized in that After collecting the user's registered audio data set, the method further includes: A reference facial image of the user is extracted from the facial image dataset.

13. The method according to claim 12, characterized in that Before collecting the target audio data of the user in the application scenario, the method further includes: Determining a target position of the user in an application scenario based on a reference facial image of the user; The sound reception weight at the target direction is enhanced.

14. The method according to claim 1, wherein The method further comprises: The reference voiceprint vector of the user is updated based on the user audio data of the user at intervals of a preset duration.

15. The method according to claim 7, characterized in that After collecting the target audio data of the user in the application scenario and obtaining the target audio features of the target audio data, the method further includes: Acquire a target volume and / or a target noise type of the target audio data; A weight parameter corresponding to the target audio data is determined based on the target volume and / or target noise type.

16. The method according to claim 15, characterized in that The separating the user audio feature from the target audio feature by using a mask matrix includes: Generating a first mask matrix and a second mask matrix in the neural network model, wherein the first mask matrix is used to comprehensively suppress noise, and the second mask matrix is used to suppress noise overlapping with the user audio feature; generating a third mask matrix based on the weight parameters, the first mask matrix, and the second mask matrix; The user audio feature is separated from the target audio feature by using the third mask matrix.

17. The method according to claim 15, characterized in that After obtaining the target volume and / or target noise type of the target audio data, the method further includes: A target audio type of the target audio data is determined based on the target volume and / or the target noise type.

18. The method according to claim 17, characterized in that The separating the user audio feature from the target audio feature by using a mask matrix includes: Determining a target mask matrix corresponding to the target audio data based on the target audio type; The user audio features are separated from the target audio features by using the target mask matrix.

19. An audio acquisition device, characterized in that: The device comprises: A target audio feature acquisition unit is used to collect target audio data of the user in the application scenario and obtain target audio features of the target audio data; A user audio feature acquisition unit, configured to acquire the user audio feature of the user from the target audio feature based on a pre-generated reference voiceprint vector of the user; A user audio data generating unit, configured to generate user audio data of the user based on the user audio features; The reference voiceprint vector is a fused voiceprint vector of common features extracted from a pre-collected set of registered audio data of the user.

20. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable program code; A processor is configured to call and run the executable program code from the memory, so that the electronic device executes the method according to any one of claims 1 to 18.

21. A computer program product, characterized in that The computer program product comprises: Computer program code, when said computer program code is executed, implements the method according to any one of claims 1 to 18.

22. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program code, and when the computer program code is executed, the method according to any one of claims 1 to 18 is implemented.