Audio data processing method and device, audio data processing equipment and storage medium
Through audio noise reduction model and voiceprint feature optimization technology, the problem of speech recognition accuracy decrease in complex noise environments in the existing technology is solved, and efficient speech recognition and noise reduction effects are achieved.
Patent Information
- Application Number
- CN202510603226.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-15
AI Technical Summary
Existing audio processing technologies are difficult to effectively distinguish target features in complex noise environments, and the voiceprint model cannot be adaptively adjusted in real time, resulting in a decrease in speech recognition accuracy.
By obtaining the audio data input by the user, calling the audio noise reduction model for noise reduction, and updating the audio noise reduction parameters when the instruction recognition probability meets the conditions, combining the voiceprint feature optimization model to achieve closed-loop optimization.
It improves the accuracy of speech recognition and noise reduction in complex noise environments, and enhances the adaptability of audio processing equipment in dynamic scenarios.
Smart Images

Figure CN120496509A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of signal processing, and in particular to an audio data processing method, apparatus, audio data processing device, and storage medium. Background Art
[0002] In related audio processing technologies, users are required to repeatedly provide pure voice samples in a quiet environment. This process is not only cumbersome to operate, but is also inherently limited by the singleness of the initial samples and the limitations of the scenario, which makes it difficult for the system to effectively distinguish target features in complex noise; secondly, the static fixed design of the voiceprint model of existing audio processing technology makes it impossible to dynamically optimize through real-time feedback. When the user produces a timbre deviation, due to the lack of an adaptive adjustment mechanism, the recognition accuracy will continue to decay as the acoustic features drift. Summary of the Invention
[0003] The present application provides an audio data processing method, apparatus, audio data processing device and storage medium, aiming to improve the speech recognition accuracy and noise reduction effect of the audio processing device in practical applications.
[0004] The technical solution is as follows:
[0005] In a first aspect, an embodiment of the present application provides an audio data processing method, comprising:
[0006] Obtaining first audio data input by a user, and invoking a preset audio noise reduction model to perform noise reduction on the first audio data to obtain first noise-reduced audio data, where the audio noise reduction model includes audio noise reduction parameters used for the user when the audio processing device is in an awake state;
[0007] Obtaining instruction data contained in the first noise reduction audio data, and determining an instruction recognition probability of the instruction data based on a preset instruction set;
[0008] If the command recognition probability is greater than or equal to a preset probability threshold, obtaining a voiceprint feature corresponding to the first noise reduction audio data;
[0009] The audio noise reduction parameters for the user in the audio noise reduction model are updated based on the voiceprint features of the first noise reduction audio data.
[0010] In a second aspect, an embodiment of the present application provides an audio data processing device, comprising:
[0011] an audio noise reduction unit, configured to obtain first audio data input by a user, and invoke a preset audio noise reduction model to perform noise reduction on the first audio data to obtain first noise-reduced audio data, wherein the audio noise reduction model includes audio noise reduction parameters used when the audio processing device is in an awake state;
[0012] an instruction parsing unit, configured to obtain instruction data contained in the first noise reduction audio data, and determine an instruction recognition probability of the instruction data based on a preset instruction set;
[0013] a feature extraction unit, configured to obtain a voiceprint feature corresponding to the first noise reduction audio data if the command recognition probability is greater than or equal to a preset probability threshold;
[0014] A parameter updating unit is configured to update audio noise reduction parameters for the user in the audio noise reduction model based on the voiceprint features of the first noise reduction audio data.
[0015] In a third aspect, an embodiment of the present application provides an audio data processing device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the above-mentioned audio data processing method is implemented.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the above-mentioned audio data processing method is implemented.
[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the above audio data processing method when executed by a processor.
[0018] In the above technical solution, audio data is denoised using audio noise reduction parameters to obtain high-fidelity noise-reduced audio data, improving the recognition rate of user audio data in high-noise environments. By updating the audio noise reduction parameters in real time based on the voiceprint characteristics of the noise-reduced audio data when the command recognition probability of the noise-reduced audio data meets preset conditions, the adaptability of the audio processing device in dynamic scenarios is improved, and closed-loop optimization of voiceprint characteristics and audio noise reduction parameters is achieved, thereby significantly improving the noise reduction performance and speech recognition accuracy of the audio processing device. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 This is a system architecture diagram of an audio data processing method provided by an embodiment of the present application;
[0021] Figure 2This is a scenario diagram of an audio data processing method provided in an embodiment of the present application;
[0022] Figure 3 This is a flowchart of an audio data processing method provided by an embodiment of the present application;
[0023] Figure 4 This is a scenario diagram of an audio data processing method provided in an embodiment of the present application;
[0024] Figure 5 This is a flowchart of an audio data processing method provided by an embodiment of the present application;
[0025] Figure 6 This is a scenario diagram of an audio data processing method provided in an embodiment of the present application;
[0026] Figure 7 This is a flowchart of an audio data processing method provided by an embodiment of the present application;
[0027] Figure 8 This is a flowchart of an audio data processing method provided by an embodiment of the present application;
[0028] Figure 9 This is a scenario diagram of an audio data processing method provided in an embodiment of the present application;
[0029] Figure 10 is a structural diagram of an audio data processing device provided in an embodiment of the present application;
[0030] Figure 11 It is a structural diagram of an audio data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] To make the features and advantages of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0032] The following will clearly and thoroughly describe the technical solutions in this application in conjunction with the accompanying drawings. In the description of the embodiments of this application, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more than two.
[0033] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.
[0034] To improve the recognition accuracy and noise reduction effect of audio data, the present invention provides an audio data processing method, which is performed by an audio data processing device. The following is a detailed description. It should be noted that the order in which the following embodiments are described does not limit the preferred order of the embodiments.
[0035] It should be noted that the audio data processing method provided in the embodiments of the present application is applicable to application scenarios that require noise reduction and voice command recognition of user voice, including but not limited to smart home scenarios, car voice assistant scenarios, mobile device voice assistant scenarios, office scenarios, and medical scenarios. In the smart home scenario, the audio processing device can recognize the device control commands contained in the user's voice and then control the smart device.
[0036] An audio data processing method provided in an embodiment of the present application is preferably applicable to a smart home scenario. The following is a specific description of an audio data processing method in conjunction with a smart home scenario.
[0037] See Figure 1 , Figure 1 This is a system architecture diagram of an audio data processing method provided by an embodiment of the present application. Figure 1As shown, the audio processing device mainly includes an audio acquisition module, an audio noise reduction module and a state control module, wherein the audio acquisition module is used to acquire audio data in real scenes, including environmental audio data and human voice audio data, and the audio noise reduction module runs a preset audio noise reduction model, which is used to perform noise reduction processing on the audio data collected by the audio acquisition module. The state control module is used to control the working state of the audio processing device, and the audio noise reduction module can also run a preset voiceprint extraction model, which is used to extract the voiceprint features of the audio. The audio noise reduction model and the voiceprint extraction model can run independently or in combination, which is not specifically limited here. The working state of the audio processing device includes a standby state and a wake-up state. When the audio processing device is in the standby state, the audio noise reduction module of the audio processing device only provides default noise reduction parameters to perform noise reduction processing on the audio data, wherein the default noise reduction parameters refer to fixed audio noise reduction parameters executed by the audio noise reduction module according to the general scene adaptation rules, including a fixed noise suppression threshold and a fixed filter benchmark configuration.
[0038] In a practical scenario, the audio acquisition module collects audio data generated in real-world scenarios in real time. If it collects audio data from any user, it uses the status control module to obtain the operating status of the audio processing device. If it determines that the audio processing device is in standby mode, the audio data is input into the audio noise reduction module, which uses default noise reduction parameters to perform basic noise reduction processing on the audio data to generate noise-reduced audio data.
[0039] Furthermore, the noise reduction audio data is converted into text data, and the text data is analyzed to see whether it contains a preset wake-up text. If it contains any wake-up text, a feedback signal is sent to the state control module of the audio processing device to switch the audio processing device from the standby state to the wake-up state. The number and content of the wake-up texts are set by the user according to the actual application scenario and are not limited here.
[0040] After the audio processing device enters the wake-up state, the audio noise reduction module will call the preset voiceprint feature algorithm to determine the voiceprint features corresponding to the noise reduction audio data, and call the adaptive noise reduction algorithm to generate audio noise reduction parameters including frequency domain suppression threshold and filter gain based on the voiceprint features.
[0041] It should be noted that the audio noise reduction module runs a pre-trained audio noise reduction model, which is used to extract voiceprint features from the noise-reduced audio data. The audio noise reduction model can extract voiceprint features from the noise-reduced audio data by calculating audio features, including Mel-frequency cepstral coefficients or linear prediction cepstral parameters. Voiceprint features from the noise-reduced audio data can also be extracted using the convolutional layers or attention mechanisms within the audio noise reduction model. The specific method used can be determined based on the actual scenario and is not specifically limited here.
[0042] In another actual scenario, the working status of the audio processing device is obtained in real time through the state control module. If the working status of the audio processing device is the awake state, the audio data generated by the user is obtained based on the audio acquisition module, and the audio noise reduction parameters obtained above are applied to the audio noise reduction module of the audio processing device to perform noise reduction on the audio data to obtain the noise reduction audio data corresponding to the audio data.
[0043] Furthermore, the noise-reduced audio data is converted into corresponding instruction data, and the instruction data is further converted into text data to determine the number of instruction words contained in the instruction data. Each instruction word in the text data is then matched with each preset instruction word in a preset instruction set to determine the instruction recognition probability of the instruction data. If the instruction recognition probability is greater than or equal to a preset probability threshold, the noise-reduced audio data is feature-processed based on the audio noise reduction module to obtain the voiceprint features corresponding to the noise-reduced audio data. Based on the voiceprint features of the noise-reduced audio data, the audio noise reduction parameters of the audio noise reduction module are updated to perform noise reduction processing on the user's audio data collected in the next audio acquisition signal.
[0044] Please also see Figure 2 , Figure 2 This is a scene diagram of an audio data processing method provided by an embodiment of the present application. Figure 2 As shown, when the user inputs audio data to the audio processing device for the first time, the audio processing device will detect whether there is a preset wake-up word in the audio processing device. If the preset wake-up word is included, the audio data will be input into the audio noise reduction model. At this time, the audio data will be denoised based on the default audio noise reduction parameters in the audio noise reduction model to obtain audio noise reduction data; the audio processing device will perform voice command recognition on the audio noise reduction data to obtain the instruction recognition probability of the instruction data in the audio noise reduction data. If the instruction recognition probability is greater than the preset probability threshold, the voiceprint features of the audio noise reduction data will be further obtained, and the audio noise reduction parameters of the audio noise reduction model will be updated based on the voiceprint features.
[0045] When the user is not inputting audio data for the first time, that is, the user continuously inputs audio data to the audio processing device, the updated audio noise reduction parameters in the audio noise reduction model will be called to reduce the noise of the audio data, and it will be determined again whether the instruction recognition probability of the instruction data corresponding to the audio noise reduction data is greater than the preset probability threshold; if the instruction recognition probability is greater than the preset probability threshold, the voiceprint features of the audio noise reduction data will be further obtained, and the audio noise reduction parameters of the audio noise reduction model will be updated again based on the voiceprint features; if the instruction recognition probability is less than the preset probability threshold, the audio processing device will be controlled to enter the standby state.
[0046] In the embodiments of the present application, audio data is denoised using audio noise reduction parameters to obtain high-fidelity noise-reduced audio data, thereby improving the recognition rate of user audio data in high-noise environments. By updating the audio noise reduction parameters in real time based on the voiceprint characteristics of the noise-reduced audio data when the command recognition probability of the noise-reduced audio data meets preset conditions, the adaptability of the audio processing device in dynamic scenarios is improved, and closed-loop optimization of voiceprint characteristics and audio noise reduction parameters is achieved, thereby significantly improving the noise reduction performance and recognition accuracy of the audio processing device.
[0047] The following will be combined Figure 3-Figure 9 , an audio data processing method provided in an embodiment of the present application is introduced in detail.
[0048] See Figure 3 , Figure 3 This is a flow chart of an audio data processing method provided by an embodiment of the present application. Figure 3 As shown, the method of the embodiment of the present application may include the following steps S101-S104.
[0049] S101: Acquire first audio data input by a user, and call a preset audio noise reduction model to perform noise reduction on the first audio data to obtain first noise-reduced audio data.
[0050] Specifically, when the audio processing device is in the awake state, the audio acquisition module based on the audio processing device obtains the first audio data input by the user from the real scene, and inputs the first audio data into the audio noise reduction module to perform noise reduction processing on the first audio data based on the audio noise reduction parameters to obtain the first noise reduction audio data. At this time, the user's voice in the first audio data will be clearer and more prominent.
[0051] It should be noted that when the first audio data is the first audio data collected when the audio processing device is in an awake state, the audio noise reduction parameters applied are the default noise reduction parameters used when the audio processing device is in a standby state to reduce the noise of the audio data containing the wake-up text, and the audio noise reduction parameters generated based on the voiceprint features corresponding to the noise reduction audio data of the audio data containing the wake-up text.
[0052] If the first audio data is not the first audio data collected when the audio processing device is in an awake state, the audio noise reduction parameter applied is the audio noise reduction parameter updated based on the voiceprint feature of the previous audio data after noise reduction processing.
[0053] S102: Acquire instruction data contained in the first noise reduction audio data, and determine an instruction recognition probability of the instruction data based on a preset instruction set.
[0054] Specifically, a preset text conversion algorithm is called to convert the first noise reduction audio data into text instructions to obtain the instruction data contained in the first noise reduction audio data, and the instruction data is segmented and keyword extracted to separate the core first instruction word. The first instruction word is then similarity calculated with each preset instruction word in the preset instruction set to finally determine the recognition probability of the instruction data, wherein the instruction set is a preset set of operation commands that can be recognized and executed.
[0055] S103: If the command recognition probability is greater than or equal to a preset probability threshold, obtain the voiceprint feature corresponding to the first noise reduction audio data.
[0056] Specifically, if the instruction recognition probability is greater than or equal to a preset probability threshold, the first noise reduction audio data is framed, that is, the first noise reduction audio data is divided into multiple segmented audio data of preset lengths, and a smoothing function is called to smooth each segmented audio data to reduce the spectral distortion of each segmented audio data.
[0057] Furthermore, based on the audio denoising model executed by the audio denoising module, the voiceprint features of the first denoised audio data are determined. During this process, the audio denoising model can quantify the voiceprint features of the first denoised audio data using traditional acoustic feature extraction methods such as Mel-frequency cepstral coefficients or linear prediction cepstral parameters. Alternatively, the voiceprint representation of the first denoised audio data can be determined using the audio denoising model's convolutional layers or attention mechanism. These two methods can be used independently or in combination, and the specific choice is flexibly adjusted based on the actual scenario's requirements for feature differentiation and computational efficiency, and are not specifically limited here.
[0058] S104: Update the audio noise reduction parameters for the user in the audio noise reduction model based on the voiceprint features of the first noise reduction audio data.
[0059] Specifically, a human voice model is established by analyzing the voiceprint features of the first noise reduction audio data, and the frequency band gain and noise suppression threshold are calculated in combination with a preset noise reduction algorithm. The gain parameters and noise suppression threshold parameters of the audio noise reduction parameters are updated based on the obtained frequency band gain and noise suppression threshold.
[0060] Please also refer to Figure 4, Figure 4 This is a scene diagram of an audio data processing method provided by an embodiment of the present application. Figure 4 As shown: the audio processing device obtains audio data when it is in the awake state, and performs noise reduction processing on the obtained audio data based on the audio noise reduction parameters to generate noise-reduced audio data; then it is determined whether the instruction recognition probability of the noise-reduced audio data is greater than the set probability threshold. If the instruction recognition probability is less than the set probability threshold, the audio processing device is controlled to enter the standby state to save energy consumption.
[0061] If the command recognition probability is greater than or equal to the set probability threshold, the voiceprint features of the noise reduction audio data are extracted, and the audio noise reduction parameters are updated based on the voiceprint features.
[0062] As can be seen above, denoising audio data using audio noise reduction parameters yields high-fidelity noise-reduced audio data, improving the recognition rate of user audio data in high-noise environments. By updating the audio noise reduction parameters in real time based on the voiceprint characteristics of the noise-reduced audio data when the command recognition probability of the noise-reduced audio data meets preset conditions, the adaptability of the audio processing device in dynamic scenarios is enhanced, achieving closed-loop optimization of voiceprint characteristics and audio noise reduction parameters, thereby significantly improving the noise reduction performance and speech recognition accuracy of the audio processing device.
[0063] To ensure that the audio processing device can provide basic noise reduction capabilities for audio data in standby mode, see Figure 5 , Figure 5 This is a flow chart of an audio data processing method provided by an embodiment of the present application. Figure 5 As shown, the method of the embodiment of the present application may include the following steps S201-S204.
[0064] S201: If the audio processing device is in a standby state, a default noise reduction parameter is called to perform noise reduction on second audio data input by a user to obtain second noise-reduced audio data.
[0065] In the embodiment of the present application, the default noise reduction parameters refer to the fixed audio noise reduction parameters executed by the audio noise reduction module according to the general scene adaptation rules, including a fixed noise suppression threshold and a fixed filter baseline configuration. The purpose of the default noise reduction parameters is to quickly and universally eliminate common environmental noise when the audio processing device is in standby mode, ensure the initial audio quality, and provide a clear signal for subsequent wake-up detection. The default noise reduction parameters mainly include a fixed noise suppression threshold and a fixed filter baseline configuration, wherein the fixed noise suppression threshold is used to determine which audio is "noise". All audio below the noise suppression threshold will be retained, and audio above the noise suppression threshold will be suppressed or filtered. The fixed filter baseline configuration pre-configures the frequency range and strength of the filter to eliminate noise audio in a specific frequency band.
[0066] Specifically, when the audio processing device is in the standby state, the second audio data input by the user is processed using the preset default noise reduction parameters to quickly filter the basic environmental noise of the second audio data and generate the second noise-reduced audio data after noise reduction.
[0067] S202: Convert the second noise-reduced audio data into noise-reduced audio text.
[0068] Specifically, the conversion algorithm is called to convert the second noise-reduced audio data into a noise-reduced audio text in text form, that is, the noise-reduced audio text is obtained by analyzing the speech signal features in the second noise-reduced audio data and matching the corresponding text content.
[0069] S203: If there is a preset wake-up text in the noise reduction audio text, control the audio processing device to enter a wake-up state.
[0070] Specifically, the noise reduction audio text is matched with a preset wake-up text. If the noise reduction audio text contains any preset wake-up text, the audio device is switched from the standby state to the wake-up state.
[0071] Exemplarily, the preset wake-up text includes text A and text B. The first noise reduction audio text is "What to eat for lunch today"; the second noise reduction audio text is "Text B, what's the weather like outside". Since the second noise reduction audio text contains the preset wake-up text B, the audio processing device is controlled to enter the wake-up state.
[0072] S204: Determine audio noise reduction parameters based on the voiceprint features of the second noise reduction audio data.
[0073] Specifically, the second noise-reduced audio data is framed, that is, divided into multiple segments of preset length. A smoothing function is then applied to each segment to reduce spectral distortion. Furthermore, a short-time Fourier transform is applied to calculate the spectrum of each segment, i.e., the energy distribution of each segment at different frequencies. After the aforementioned processing, the spectra of each segment are merged to obtain the spectral features corresponding to the second noise-reduced audio data.
[0074] After obtaining the spectral features, the spectral features need to be fused with the fundamental frequency and resonance peaks of the second noise reduction audio data. The fundamental frequency, resonance peak, and spectral features are normalized to eliminate the dimensional differences between the different features, so that the Hertz unit of the fundamental frequency, the frequency value of the resonance peak, and the energy value of the spectrum are unified to the same order of magnitude. After normalization, the three types of features are fused into a high-dimensional vector through weighted splicing or a neural network model, and then redundant information is removed with the help of principal component analysis, ultimately obtaining the voiceprint features of the second noise reduction audio data. Furthermore, by analyzing the voiceprint features, a human voice model is established, and the frequency band gain and noise suppression threshold are dynamically calculated by combining spectral subtraction, Wiener filtering, or deep learning algorithms to generate audio noise reduction parameters.
[0075] Please also refer to Figure 6 , Figure 6 This is a scene diagram of an audio data processing method provided by an embodiment of the present application. Figure 6 As shown: After first obtaining the audio data, if it is detected that the audio processing device is in standby mode, the default noise reduction parameters are selected to perform noise reduction processing on the audio data to generate noise reduction audio, and the noise reduction audio is further converted into noise reduction audio text, and it is detected whether there is a preset wake-up text in the noise reduction audio text. If there is a preset wake-up text in the noise reduction audio text, the audio noise reduction parameters of the audio processing device are determined based on the voiceprint features of the noise reduction audio, and finally a closed-loop processing link is formed from basic noise reduction to precise adaptation.
[0076] From the above, we can see that by using low-complexity default noise reduction parameters for preliminary noise reduction when the audio processing device is in standby mode, the computing load is reduced while maintaining basic voice intelligibility. Subsequently, a dual verification mechanism of text conversion and wake-up text matching is used to ensure the accuracy of the wake-up operation and avoid invalid power consumption caused by false triggering. After successful wake-up, the noise reduction parameters are dynamically adjusted based on the real-time extracted voiceprint features. By optimizing the frequency domain suppression rules of the target voice in a targeted manner, the adaptive suppression capability of environmental noise is enhanced while retaining key voice information, thereby significantly improving the robustness of voice recognition and noise reduction efficiency in complex scenarios.
[0077] In actual scenarios, the voice wake-up method may be affected by the environment and may fail to wake up the audio processing device. Therefore, other methods are needed to enhance human-computer interaction. Figure 7 , Figure 7 This is a flow chart of an audio data processing method provided by an embodiment of the present application. Figure 7 As shown, the method of the embodiment of the present application may include the following steps S301-S302.
[0078] In an embodiment of the present application, a user can achieve linkage with an audio processing device through an image acquisition device set up in a real-world scene. The image acquisition device is used to capture the user's gestures or facial features in the real-world scene in real time, and use edge computing or cloud-based algorithms to identify whether the user has triggered a preset wake-up gesture or whether the user's facial features meet the preset facial features.
[0079] It should be noted that preset facial features refer to authorized user facial information pre-registered through biometric technology, such as facial key points, 3D contours, or iris features. Preset wake-up gestures are specific gesture commands defined by the user or the system, such as drawing a circle in the air or spreading five fingers. Both preset facial features and preset wake-up gestures are algorithmically modeled and stored on the audio processing device or cloud service.
[0080] S301: If the user triggers a preset wake-up gesture, the audio processing device is controlled to enter a wake-up state.
[0081] Specifically, an image acquisition device captures a real-time video stream containing the user. A gesture recognition algorithm is then run through edge computing or cloud services to extract the spatiotemporal characteristics of gestures from the video stream, such as hand motion trajectory and shape changes. If the recognition result matches a preset wake-up gesture, a wake-up command is sent to the audio processing device via the internet or local area network, causing it to enter a wake-up state.
[0082] For example, the user makes a "palm facing forward" gesture to the image acquisition device. After recognition, the image acquisition device determines that the "palm facing forward" gesture is a preset wake-up gesture. At this time, a response signal is sent to the state controller module of the audio processing device to put the audio processing device into a wake-up state.
[0083] S302: If the facial features of the user meet the preset facial features, the audio processing device is controlled to enter the awake state.
[0084] Specifically, when the user enters the field of view of the image acquisition device, the image acquisition device locks the user through face detection, extracts the user's facial features, and compares the user's facial features with the stored preset facial feature database for similarity. If the comparison is successful, a response signal is sent to the audio processing device to put the audio processing device into a wake-up state.
[0085] As can be seen above, dual biometric cross-verification enhances identity security, while gesture recognition compensates for the limitations of facial recognition in low light or under obstruction. This dual-modal dynamic complementary mechanism effectively addresses the issues of false triggering and rejection of single recognition in complex environments. It balances low latency with high robustness, enabling contactless and natural interaction while reducing power consumption through algorithm optimization. Ultimately, this creates a secure, reliable, and adaptable intelligent wake-up system for multiple scenarios, significantly improving device interaction efficiency and user experience.
[0086] In a feasible implementation, an audio data processing method provided in an embodiment of the present application includes calling a preset voiceprint extraction model to determine a voiceprint feature corresponding to the first noise reduction audio data.
[0087] Specifically, the voiceprint encoder of the audio denoising model can quantify the voiceprint features of the first denoised audio data by extracting traditional acoustic feature extraction methods such as Mel-frequency cepstral coefficients or linear prediction cepstral parameters from the first denoised audio data. Alternatively, the voiceprint features of the first denoised audio data can be determined using the audio denoising model's convolutional layer or attention mechanism. These two methods can be used independently or in combination, with the specific choice being flexibly adjusted based on the actual scenario's requirements for feature differentiation, computational efficiency, and privacy protection, and are not specifically limited here.
[0088] As can be seen from the above, traditional voiceprint feature extraction methods are suitable for scenarios with high real-time requirements; voiceprint feature extraction based on convolutional layers or attention mechanisms can enhance resistance to noise interference and cross-channel variations. By dynamically setting the voiceprint feature extraction method based on the actual application scenario, we can not only reduce computational overhead through lightweight features, but also improve voiceprint feature recognition accuracy in complex environments by leveraging deep representation, achieving a dynamic balance between extraction efficiency and effect.
[0089] In a feasible implementation, the method of the embodiment of the present application may include determining instruction words contained in instruction data, matching each instruction word with each preset instruction word in the instruction set, and obtaining an instruction recognition probability of the instruction data.
[0090] Specifically, the input instruction data is segmented and keyword extracted to separate the core instruction words, and then the similarity is calculated with each preset instruction word in the preset instruction set to finally determine the instruction matching probability of the instruction data. The instruction set is a preset set of operation commands that can be recognized and executed, and is not specifically limited here.
[0091] For example, taking air conditioning control as an example, the user instruction data is "lower two degrees". The instruction data is segmented through natural language processing to obtain "lower" and "two degrees". Non-key modifiers such as "two degrees" are filtered out and retained as numerical parameters, and the core instruction word "lower" is extracted; then semantic matching is performed with the preset instruction set, such as "increase temperature", "lower temperature", and "turn on and off": the word vector model is used to map "lower" to a high-dimensional vector, and its cosine similarity with "lower temperature" and "increase temperature" is calculated respectively. Among them, the cosine similarity value range is -1 to 1, and the larger the cosine similarity value, the more similar it is. Assuming that the cosine similarity is 0.85 and 0.1 respectively, the recognition probability of "lower temperature" is 85%; combined with the parameter "two degrees", the final operation instruction is generated, triggering the air conditioning temperature control module to lower the set temperature by 2 degrees.
[0092] From the above, we can see that using natural language processing technology to implement a closed-loop process of command segmentation, semantic vector matching, and parameter parsing can effectively improve the fault tolerance of command recognition in complex scenarios.
[0093] To ensure the execution efficiency and save resource overhead. Figure 8 , Figure 8 This is a flow chart of an audio data processing method provided by an embodiment of the present application. Figure 8 As shown, the method of the embodiment of the present application may include the following steps S401-S404.
[0094] S401: Generate a device control instruction based on a first instruction word of instruction data, and determine an instruction execution device for executing the device control instruction.
[0095] Specifically, the core operation intention corresponding to the first instruction word of the instruction data is extracted through semantic analysis, and the device name in the first instruction word is determined to determine the instruction execution device.
[0096] For example, when the first instruction word is "dim the bedroom lights", semantic analysis is performed to determine that the device control instruction is "dim the lights", and based on "bedroom lights", the instruction execution device is determined to be the smart lamp in the bedroom.
[0097] S402 , obtaining a device preference parameter of the user for the instruction execution device, and writing the device preference parameter into the device control instruction.
[0098] Specifically, after determining the device to execute the command, the user's pre-set device preference parameters will be called up, including the user's historical adjustments to the operating parameters of the command execution device, such as light color temperature, brightness level, and lighting mode, and dynamically embedded into the basic control command. For example, if the user's historical settings prefer "warm light mode", the color temperature and brightness parameters will be added to the "turn on the lights" command to form a complete command, such as "turn on the bedroom lights to warm light, brightness 50%." By converting general commands into personalized commands, the burden of repeated settings on users is reduced.
[0099] S403: Send the device control instruction to the instruction execution device.
[0100] Specifically, the device control instruction is transmitted to the instruction execution device through a communication protocol, such as Wi-Fi or Bluetooth communication protocol, so that the instruction execution device executes the corresponding control operation of the device control instruction.
[0101] For example, when the device control instruction is "dim the bedroom lights" and the instruction execution device is the smart lamp in the bedroom, "dim the bedroom lights" will be sent to the smart lamp in the bedroom. After receiving the device control instruction, the smart lamp will reduce the brightness of the light by one level.
[0102] S404: If no audio data is obtained within a preset time period, the audio processing device is controlled to enter a standby state.
[0103] Specifically, the sound collection module continuously monitors whether the user continues to generate audio data and starts a countdown mechanism. If the user is not detected generating audio data within the preset time, it is determined that the user has no intention of interaction. A feedback signal is then generated and sent to the state control module to put the audio processing device into standby mode, thereby reducing the power consumption of the audio processing device while retaining basic wake-up capabilities.
[0104] As can be seen above, by integrating user preference parameters into control commands and issuing them for execution, device operation is closely aligned with user habits, reducing the need for manual intervention. Through the timeout automatic standby mechanism, the device proactively reduces power consumption when there is no effective interaction, reducing energy consumption in non-essential modules and extending battery life. At the same time, it retains basic wake-up functionality, avoiding the tedious task of repeated wake-up operations while ensuring the device is always responsive, balancing energy conservation needs with the immediacy of interaction.
[0105] In a feasible implementation, the method of the embodiment of the present application may include obtaining a first instruction set corresponding to the voiceprint feature, determining a second instruction word from the first instruction set based on the instruction data; generating a second device control instruction based on the second instruction word, and sending the second device control instruction to the instruction execution device.
[0106] In an embodiment of the present application, a corresponding first instruction set is established based on the user's voiceprint features, wherein the first instruction set records a set of device control instructions generated by the user corresponding to the voiceprint features during a session with the audio processing device.
[0107] Specifically, if the audio processing device is in the awake state and fails to receive any device control commands from the user within a preset time period, the device matches the first instruction set based on the voiceprint characteristics of the wake-up statement. A second instruction word is obtained from the first instruction set, and its match with the instruction in the instruction data exceeds a preset match threshold. A device control instruction is generated based on the second instruction word, and the device control instruction is sent to the instruction execution device. The determination of the instruction device is described in step S401 above and will not be repeated here.
[0108] In a feasible implementation, the method of an embodiment of the present application may include obtaining a second instruction set corresponding to a preset wake-up text; determining a third instruction word from the second instruction set based on instruction data, and generating a third device control instruction based on the third instruction word; and sending the third device control instruction to an instruction execution device.
[0109] In an embodiment of the present application, the user can actively set a personalized preset wake-up text. When the preset wake-up text is triggered, the device control instructions generated by the user when in a conversation with the audio processing device will be summarized and recorded to generate a second instruction set, and a mapping relationship will be established between the preset wake-up text and the second instruction set at the same time.
[0110] Specifically, if it is detected that the user fails to issue a specific device control instruction within a preset time period after triggering the preset wake-up text, a second instruction set corresponding to the preset wake-up text is obtained. A third instruction word is obtained from the second instruction set, and its instruction match with the instruction word in the instruction data is greater than a preset match threshold. A device control instruction is generated based on the third instruction word, and the device control instruction is sent to the instruction execution device. The determination of the instruction device is described in step S401 above and will not be repeated here.
[0111] For example, the instruction word in the instruction data is "turn on the air conditioner" and the third instruction word that matches it is "turn on the cooling mode of the air conditioner and set the temperature to 24 degrees."
[0112] In a feasible implementation, the method of the embodiment of the present application may include controlling the audio processing device to enter a standby state if the instruction recognition probability is less than a preset probability threshold.
[0113] Specifically, when the probability of command recognition is less than the preset probability threshold, it is considered that the user's command is issued incorrectly or the user is in an incorrect recognition scenario. At this time, a feedback signal is generated and sent to the state control module to make the audio processing device enter the standby state, so as to reduce the power consumption of the audio processing device while retaining the basic wake-up capability.
[0114] For example, there is no instruction execution device for "air purifier" in the real scenario. When the execution data is "turn on the air purifier", since there is no instruction for the air purifier in the preset instruction set, it is considered that the instruction recognition probability is less than the probability threshold. At this time, the audio processing device is controlled to enter the standby state.
[0115] From the above, we can see that by actively intercepting low-credibility instructions, device misoperation due to misidentification can be avoided, reducing safety risks and user troubles; at the same time, combined with the design of quickly switching to standby mode, it further reduces the waste of resources caused by invalid instruction processing, ensuring that the device only invests full computing power when the instruction is clear, thereby improving the overall operating efficiency and reliability of the system.
[0116] In a feasible implementation, the method of the embodiment of the present application may include collecting ambient environmental noise and adjusting audio noise reduction parameters based on a noise frequency value of the ambient environmental noise.
[0117] Specifically, the audio acquisition module collects noise signals from the surrounding environment in real time, and uses spectrum analysis technology to decompose the noise into sound wave components of different frequencies to identify the main interference frequencies, such as continuous low-frequency humming or high-frequency sharp noise. Subsequently, the parameters of the filter in the audio noise reduction parameters are dynamically adjusted according to the noise frequency distribution: if low-frequency noise is detected to be dominant, the signal suppression strength of the low-frequency band is enhanced; if high-frequency noise is significant, the attenuation amplitude of the high-frequency band is increased. This process
[0118] For example, if it is detected that the ambient noise is mainly composed of mid- and low-frequency human voices and background music, the audio noise reduction parameters in the 200Hz-1000Hz frequency band will be automatically enhanced to effectively filter out the background music, while retaining the high-frequency part where the human voice is located, so that the human voice is more prominent and clear.
[0119] From the above, we can see that the adaptive algorithm ensures that the audio noise reduction parameters always match the current environmental noise characteristics, thereby improving the collection clarity of the user's audio data.
[0120] In a feasible implementation, the method of the embodiment of the present application may include: an audio noise reduction model includes an audio encoder, a feature extractor, and a feature decoder; calling the audio encoder to encode the first audio data to obtain initial audio features corresponding to the first audio data; calling the feature extractor and extracting the user's audio features from the initial audio features based on audio noise reduction parameters; calling the feature decoder to decode the user audio features to obtain first noise reduction audio data.
[0121] Specifically, the input first audio data is compressed and encoded using an audio encoder to generate initial audio features corresponding to the first audio data. A feature extractor invokes audio noise reduction parameters used for the user to perform deep feature analysis on the first audio features, thereby extracting user audio features representing the user audio from the first audio features. A feature decoder is invoked to decode the user audio features to obtain first noise-reduced audio data.
[0122] See Figure 9 , Figure 9 This is a scene diagram of an audio data processing method provided by an embodiment of the present application. Figure 9As shown, the audio encoder is invoked to encode the first audio data to obtain the first audio feature. Meanwhile, the voiceprint extraction model determines the voiceprint features of the second noise-reduced audio data and sends them to the audio noise reduction model. The audio noise reduction model generates user-specific audio noise reduction parameters based on the voiceprint features. The feature extractor uses the audio noise reduction parameters to extract feature data from the first audio feature to obtain first feature audio data. The feature encoder decodes the first feature audio data to obtain first noise-reduced audio data.
[0123] In a feasible implementation, the method of the embodiment of the present application may include sending the first noise reduction audio data to a voiceprint extraction device, the voiceprint extraction device is used to extract voiceprint features from the first noise reduction audio data, and return the voiceprint features to the audio processing device.
[0124] In this embodiment of the present application, the voiceprint extraction device can be another audio processing device running a voiceprint extraction model or a cloud service device running a voiceprint extraction model. The voiceprint extraction device extracts voiceprint features from the first noise reduction audio data and returns the extracted voiceprint features to the audio processing device. The audio processing device updates the audio noise reduction parameters for the user in the audio noise reduction model based on the received voiceprint features.
[0125] In the embodiment of the present application, the task of voiceprint feature extraction is distributed to the voiceprint extraction device for execution, thereby reducing the task load of the audio processing device and effectively improving the execution noise reduction efficiency of the audio processing device.
[0126] based on Figure 1 The following is a schematic diagram of the scene. Figure 10 , the audio data processing device provided in the embodiment of the present application is introduced in detail. It should be noted that, Figure 10 The audio data processing device in the present application is used to execute Figure 2-Figure 9 For the convenience of explanation, only the part related to the embodiment of the present application is shown. For the specific technical details not disclosed, please refer to the present application. Figure 2-Figure 9 In the embodiment shown, the audio data processing device 500 may include an audio noise reduction unit 501, an instruction parsing unit 502, a feature extraction unit 503, and a parameter updating unit 504, specifically as follows:
[0127] An audio noise reduction unit 501 is configured to obtain first audio data input by a user and to perform noise reduction on the first audio data using a preset audio noise reduction model to obtain first noise-reduced audio data, wherein the audio noise reduction model includes audio noise reduction parameters used for the user when the audio processing device is in an awake state;
[0128] An instruction parsing unit 502 is configured to obtain instruction data contained in the first noise reduction audio data and determine an instruction recognition probability of the instruction data based on a preset instruction set;
[0129] The feature extraction unit 503 is configured to obtain a voiceprint feature corresponding to the first noise reduction audio data if the command recognition probability is greater than or equal to a preset probability threshold;
[0130] The parameter updating unit 504 is configured to update the audio noise reduction parameters for the user in the audio noise reduction model based on the voiceprint features of the first noise reduction audio data.
[0131] In some embodiments, the audio noise reduction unit 501 further includes a state acquisition unit, a text acquisition unit, a state control subunit, and a parameter generation unit.
[0132] a state acquisition unit, configured to, if the audio processing device is in a standby state, call default noise reduction parameters to perform noise reduction on the second audio data input by the user to obtain second noise-reduced audio data;
[0133] a text acquisition unit, configured to convert the second noise reduction audio data into noise reduction audio text;
[0134] A state control subunit, configured to control the audio processing device to enter a wake-up state if a preset wake-up text exists in the noise reduction audio text;
[0135] The parameter generating unit is configured to determine the audio noise reduction parameters based on the voiceprint features of the second noise reduction audio data.
[0136] In some embodiments, the audio noise reduction unit 501 further includes a first detection unit and a second detection unit.
[0137] A first detection unit is configured to control the audio processing device to enter a wake-up state if the user triggers a preset wake-up gesture;
[0138] The second detection unit is configured to control the audio processing device to enter a wake-up state if the facial features of the user meet the preset facial features.
[0139] In some embodiments, the feature extraction unit 503 further includes an audio feature extraction unit and a feature fusion unit.
[0140] The feature extraction subunit is used to call a preset voiceprint extraction model to determine the voiceprint features corresponding to the first noise reduction audio data.
[0141] In some embodiments, the instruction parsing unit 502 further includes a recognition probability acquisition unit.
[0142] The recognition probability acquisition unit is used to determine the instruction words contained in the instruction data, match each instruction word with each preset instruction word in the instruction set, and obtain the instruction recognition probability of the instruction data.
[0143] In some embodiments, the feature extraction unit 503 further includes an instruction word parsing unit and an instruction issuing unit.
[0144] an instruction word parsing unit, configured to generate a device control instruction based on a first instruction word of the instruction data, and determine an instruction execution device for executing the device control instruction;
[0145] The instruction issuing unit is used to issue the device control instruction to the instruction execution device.
[0146] In some embodiments, the feature extraction unit 503 further includes a first instruction word generating unit and a first instruction word issuing unit.
[0147] a first instruction word generating unit, configured to obtain a first instruction set corresponding to the voiceprint feature, and determine a second instruction word from the first instruction set based on the instruction data;
[0148] The first instruction word issuing unit is used to generate a second device control instruction based on the second instruction word, and issue the second device control instruction to the instruction execution device.
[0149] In some embodiments, the feature extraction unit 503 further includes a second instruction word set acquisition unit, a second instruction word generation unit, and a second instruction word issuing unit.
[0150] A second instruction word set acquisition unit, configured to acquire a second instruction set corresponding to a preset wake-up text;
[0151] a second instruction word generating unit, configured to determine a third instruction word from the second instruction set based on the instruction data, and generate a third device control instruction based on the third instruction word;
[0152] The second instruction word issuing unit is used to issue the third device control instruction to the instruction execution device.
[0153] In some embodiments, the audio noise reduction unit 501 further includes a preference parameter acquisition unit.
[0154] The preference parameter acquisition unit is used to acquire the device preference parameters of the user for the instruction execution device and write the device preference parameters into the device control instruction.
[0155] In some embodiments, the feature extraction unit 503 further includes an audio duration monitoring unit.
[0156] The audio duration monitoring unit is used to control the audio processing device to enter the standby state if no audio data is obtained within the preset duration.
[0157] In some embodiments, the feature extraction unit 503 further includes a feature extraction subunit.
[0158] The feature extraction unit subunit is used to control the audio processing device to enter a standby state if the instruction recognition probability is less than a preset probability threshold.
[0159] In some embodiments, a noise adjustment unit is further included.
[0160] The noise adjustment unit is used to collect ambient noise and adjust the audio noise reduction parameters based on the noise frequency value of the ambient noise.
[0161] In some embodiments, the audio data processing apparatus 500 further includes an encoding unit, a feature data extraction unit, and a decoding unit.
[0162] an encoding unit, configured to call an audio encoder to encode the first audio data to obtain initial audio features corresponding to the first audio data;
[0163] A feature data extraction unit, configured to call a feature extractor and extract a user audio feature of the user from the initial audio feature based on the audio noise reduction parameter;
[0164] The decoding unit is configured to call a feature decoder to decode the user audio feature to obtain first noise reduction audio data.
[0165] In some embodiments, the feature extraction unit 503 is a first task issuing unit.
[0166] The first task issuing unit is configured to upload the first noise reduction audio data to a cloud service, and the cloud service is configured to extract voiceprint features from the first noise reduction audio data and return the voiceprint features to the audio processing device.
[0167] In some embodiments, the feature extraction unit 503 is a second task issuing unit.
[0168] The second task issuing unit is configured to send the first noise reduction audio data to the voiceprint extraction device, which is configured to extract voiceprint features from the first noise reduction audio data and return the voiceprint features to the audio processing device.
[0169] In the embodiments of the present application, audio data is denoised using audio noise reduction parameters to obtain high-fidelity noise-reduced audio data, thereby improving the recognition rate of user audio data in high-noise environments. By updating the audio noise reduction parameters in real time based on the voiceprint characteristics of the noise-reduced audio data when the command recognition probability of the noise-reduced audio data meets preset conditions, the adaptability of the audio processing device in dynamic scenarios is improved, and closed-loop optimization of voiceprint characteristics and audio noise reduction parameters is achieved, thereby significantly improving the noise reduction performance and speech recognition accuracy of the audio processing device.
[0170] In addition, the audio data processing device provided in the above embodiment and an embodiment of an audio data processing method belong to the same concept, and the implementation process thereof is detailed in the method embodiment and will not be repeated here.
[0171] The serial numbers of the embodiments of the present application are for descriptive purposes only and do not represent the merits of the embodiments. In some cases, the actions or steps described in the claims may be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0172] See Figure 11 , which is a structural diagram of an audio data processing device provided in an embodiment of the present application. Figure 11 As shown, the audio data processing device 600 includes a processor 601 and a memory 602. The processor 601 is electrically connected to the memory 602.
[0173] The processor 601 is the control center of the audio data processing device 600 and may include one or more processing cores. The processor 601 utilizes various interfaces and lines to connect the various parts of the entire audio data processing device, and executes various functions of the audio data processing device and processes data by running or calling computer programs stored in the memory 602, as well as calling data stored in the memory 602, thereby performing overall management and control of the audio data processing device. Optionally, the processor 601 may be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), or programmable logic array (PLA). The processor 601 may integrate one or more combinations of a CPU, a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user pages, and applications; the GPU is responsible for rendering and drawing display content; and the modem is used to handle wireless communications. It is understandable that the above-mentioned modem may also not be integrated into the processor 601 and may be implemented separately through a communication chip.
[0174] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and audio data processing by running the computer programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, computer programs required for at least one function, etc.; the data storage area may store data created based on the use of the audio data processing device.
[0175] In addition, the memory 602 may include a high-speed random access memory and a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.
[0176] In the embodiment of the present application, the processor 601 in the audio data processing device 600 loads instructions corresponding to one or more computer program processes into the memory 602 according to the following steps, and the processor 601 runs the computer program stored in the memory 602 to implement various functions as follows:
[0177] Obtaining first audio data input by a user, and performing noise reduction on the first audio data by calling audio noise reduction parameters to obtain first noise-reduced audio data, where the audio noise reduction parameters are noise reduction parameters used when the audio processing device is in an awake state;
[0178] Obtaining instruction data contained in the first noise reduction audio data, and determining an instruction recognition probability of the instruction data based on a preset instruction set;
[0179] If the command recognition probability is greater than or equal to a preset probability threshold, obtaining a voiceprint feature corresponding to the first noise reduction audio data;
[0180] The audio noise reduction parameters are updated based on the voiceprint features of the first noise reduction audio data.
[0181] Optionally, after obtaining the audio data input by the user, the processor 601 specifically performs the following: if the audio processing device is in standby state, calling the default noise reduction parameters to perform noise reduction on the second audio data input by the user to obtain second noise-reduced audio data; converting the second noise-reduced audio data into noise-reduced audio text; if there is a preset wake-up text in the noise-reduced audio text, controlling the audio processing device to enter the wake-up state; and determining the audio noise reduction parameters based on the voiceprint features of the second noise-reduced audio data.
[0182] Optionally, before executing the call of audio noise reduction parameters to perform noise reduction on the first audio data to obtain the first noise-reduced audio data, the processor 601 specifically performs the following: if the user triggers a preset wake-up gesture, the audio processing device is controlled to enter a wake-up state; if the user's facial features meet the preset facial features, the audio processing device is controlled to enter a wake-up state.
[0183] Optionally, when executing the acquisition of the voiceprint feature corresponding to the first noise reduction audio data, the processor 601 specifically performs: calling a preset voiceprint extraction model to determine the voiceprint feature corresponding to the first noise reduction audio data.
[0184] Optionally, the processor 601 determines the instruction recognition probability of the instruction data based on the preset instruction set, and specifically performs: determining the instruction words contained in the instruction data, matching each instruction word with each preset instruction word in the instruction set, and obtaining the instruction recognition probability of the instruction data.
[0185] Optionally, the processor 601 executes, specifically: generating a device control instruction based on the first instruction word of the instruction data, and determining an instruction execution device for executing the device control instruction; and sending the device control instruction to the instruction execution device.
[0186] Optionally, after executing the acquisition of the voiceprint feature corresponding to the first noise reduction audio data, the processor 601 specifically performs the following: acquiring the first instruction set corresponding to the voiceprint feature, determining the second instruction word from the first instruction set based on the instruction data; generating a second device control instruction based on the second instruction word, and sending the second device control instruction to the instruction execution device.
[0187] Optionally, after executing if there is a preset wake-up text in the noise reduction audio text, the processor 601 specifically executes: obtaining a second instruction set corresponding to the preset wake-up text; after executing if the instruction recognition probability is greater than or equal to a preset probability threshold, the processor 601 specifically executes: determining a third instruction word from the second instruction set based on the instruction data, generating a third device control instruction based on the third instruction word; and sending the third device control instruction to the instruction execution device.
[0188] Optionally, before sending the device control instruction to the instruction execution device, the processor 601 specifically performs the following steps: obtaining a device preference parameter of the user for the instruction execution device, and writing the device preference parameter into the device control instruction.
[0189] Optionally, after executing and sending the device control instruction to the instruction execution device, the processor 601 specifically performs: if no audio data is obtained within a preset time period, controlling the audio processing device to enter a standby state.
[0190] Optionally, after determining the instruction recognition probability of the instruction data based on the preset instruction set, the processor 601 specifically performs: if the instruction recognition probability is less than a preset probability threshold, controlling the audio processing device to enter a standby state.
[0191] Optionally, the processor 601 is further configured to collect ambient environmental noise and adjust audio noise reduction parameters based on a noise frequency value of the ambient environmental noise.
[0192] Optionally, the processor 601 calls a preset audio noise reduction model to perform noise reduction on the first audio data to obtain first noise-reduced audio data, and specifically performs: calling an audio encoder to encode the first audio data to obtain initial audio features corresponding to the first audio data; calling a feature extractor and extracting user audio features of the user from the initial audio features based on audio noise reduction parameters; calling a feature decoder to decode the user audio features to obtain first noise-reduced audio data.
[0193] Optionally, when executing to obtain the voiceprint features corresponding to the first noise reduction audio data, the processor 601 specifically performs: sending the first noise reduction audio data to a voiceprint extraction device, the voiceprint extraction device is used to extract the voiceprint features of the first noise reduction audio data, and returning the voiceprint features to the audio processing device.
[0194] In the embodiments of the present application, audio data is denoised using audio noise reduction parameters to obtain high-fidelity noise-reduced audio data, thereby improving the recognition rate of user audio data in high-noise environments. By updating the audio noise reduction parameters in real time based on the voiceprint characteristics of the noise-reduced audio data when the command recognition probability of the noise-reduced audio data meets preset conditions, the adaptability of the audio processing device in dynamic scenarios is improved, and closed-loop optimization of voiceprint characteristics and audio noise reduction parameters is achieved, thereby significantly improving the noise reduction performance and speech recognition accuracy of the audio processing device.
[0195] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer, the computer executes the above-mentioned related method steps to implement an audio data processing method provided by the above-mentioned embodiment.
[0196] In addition, the device provided in the embodiment of the present application can specifically be a chip, component or module, and the chip may include a connected processor and memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute an audio data processing method provided in the above embodiment.
[0197] An embodiment of the present application also provides a computer-readable storage medium, which stores computer program code. When the computer program code runs on a computer, the computer executes the above-mentioned related method steps to implement an audio data processing method provided by the above-mentioned embodiment.
[0198] An embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement an audio data processing method provided in the above embodiment.
[0199] Among them, the device, computer-readable storage medium, computer program product or chip provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0200] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0201] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the coupling or direct coupling or communication connection between the related ones shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0202] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for processing audio data, characterized in that: Applied to audio processing equipment, including: Obtaining first audio data input by a user, and invoking a preset audio noise reduction model to perform noise reduction on the first audio data to obtain first noise-reduced audio data, wherein the audio noise reduction model includes audio noise reduction parameters used for the user when the audio processing device is in an awake state; Obtaining instruction data contained in the first noise reduction audio data, and determining an instruction recognition probability of the instruction data based on a preset instruction set; If the command recognition probability is greater than or equal to a preset probability threshold, obtaining a voiceprint feature corresponding to the first noise reduction audio data; The audio noise reduction parameters for the user in the audio noise reduction model are updated based on the voiceprint features of the first noise reduction audio data.
2. The method according to claim 1, characterized in that After obtaining the first audio data input by the user, the method further includes: If the audio processing device is in a standby state, calling a default noise reduction parameter to perform noise reduction on the second audio data input by the user to obtain second noise-reduced audio data; converting the second noise-reduced audio data into noise-reduced audio text; If a preset wake-up text exists in the noise reduction audio text, controlling the audio processing device to enter a wake-up state; An audio noise reduction parameter is determined based on the voiceprint feature of the second noise reduction audio data.
3. The method according to claim 1, characterized in that Before calling a preset audio noise reduction model to perform noise reduction on the first audio data to obtain first noise-reduced audio data, the method further includes: If the user triggers a preset wake-up gesture, controlling the audio processing device to enter a wake-up state; or, If the facial features of the user meet the preset facial features, the audio processing device is controlled to enter the awakening state.
4. The method according to claim 1, wherein The obtaining of the voiceprint feature corresponding to the first noise reduction audio data includes: A preset voiceprint extraction model is called to determine the voiceprint features corresponding to the first noise reduction audio data.
5. The method according to claim 1, wherein The determining of the instruction recognition probability of the instruction data based on a preset instruction set includes: Determine the instruction words included in the instruction data, match each of the instruction words with each of the preset instruction words in the instruction set, and obtain the instruction recognition probability of the instruction data.
6. The method according to claim 1 or 5, characterized in that If the instruction recognition probability is greater than or equal to a preset probability threshold, the method further includes: generating a device control instruction based on the first instruction word of the instruction data, and determining an instruction execution device for executing the device control instruction; The device control instruction is sent to the instruction execution device.
7. The method according to claim 1, characterized in that After obtaining the voiceprint feature corresponding to the first noise reduction audio data, the method further includes: Obtaining a first instruction set corresponding to the voiceprint feature, and determining a second instruction word from the first instruction set based on the instruction data; A second device control instruction is generated based on the second instruction word, and the second device control instruction is sent to an instruction execution device.
8. The method according to claim 2, wherein if a preset wake-up text exists in the noise reduction audio text, the method further comprises: Obtaining a second instruction set corresponding to the preset wake-up text; If the instruction recognition probability is greater than or equal to a preset probability threshold, the method further includes: determining a third instruction word from the second instruction set based on the instruction data, and generating a third device control instruction based on the third instruction word; The third device control instruction is sent to the instruction execution device.
9. The method according to claim 6, characterized in that Before sending the device control instruction to the instruction execution device, the method further includes: Obtain device preference parameters of a user for the instruction execution device, and write the device preference parameters into the device control instruction.
10. The method according to claim 6, characterized in that After sending the device control instruction to the instruction execution device, the method further includes: If no audio data is obtained within a preset time period, the audio processing device is controlled to enter a standby state.
11. The method according to claim 1, wherein After determining the instruction recognition probability of the instruction data based on the preset instruction set, the method further includes: If the instruction recognition probability is less than a preset probability threshold, the audio processing device is controlled to enter a standby state.
12. The method according to claim 1 or 2, characterized in that The method further comprises: Ambient noise is collected, and the audio noise reduction parameter is adjusted based on a noise frequency value of the ambient noise.
13. The method according to claim 1, wherein The audio noise reduction model includes an audio encoder, a feature extractor and a feature decoder; The calling a preset audio noise reduction model to perform noise reduction on the first audio data to obtain first noise-reduced audio data includes: Invoking the audio encoder to encode the first audio data to obtain initial audio features corresponding to the first audio data; Calling the feature extractor and extracting the user audio feature of the user from the initial audio feature based on the audio noise reduction parameter; The feature decoder is called to decode the user audio feature to obtain the first noise reduction audio data.
14. The method according to claim 1, wherein The obtaining of the voiceprint feature corresponding to the first noise reduction audio data includes: The first noise reduction audio data is sent to a voiceprint extraction device, which is used to extract voiceprint features from the first noise reduction audio data and return the voiceprint features to the audio processing device.
15. An audio data processing device, characterized in that: include: an audio noise reduction unit, configured to obtain first audio data input by a user, and invoke a preset audio noise reduction model to perform noise reduction on the first audio data to obtain first noise-reduced audio data, wherein the audio noise reduction model includes audio noise reduction parameters used when the audio processing device is in an awake state; an instruction parsing unit, configured to obtain instruction data contained in the first noise reduction audio data, and determine an instruction recognition probability of the instruction data based on a preset instruction set; a feature extraction unit, configured to obtain a voiceprint feature corresponding to the first noise reduction audio data if the command recognition probability is greater than or equal to a preset probability threshold; A parameter updating unit is configured to update the audio noise reduction parameters for the user in the audio noise reduction model based on the voiceprint features of the first noise reduction audio data.
16. An audio data processing device, characterized in that: The audio data processing device comprises: a memory for storing executable program code; A processor is configured to call and run the executable program code from the memory, so that the audio data processing device performs the audio data processing method according to any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed, the audio data processing method according to any one of claims 1 to 14 is implemented.
18. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the audio data processing method according to any one of claims 1 to 14 is implemented.
Citation Information
Cited By
Multi-modal office assistant system
CN121053985A