Audio processing method and electronic equipment

By using different configuration parameters to process audio data in the audio processing method, the problem of voice details loss in the prior art is solved, and the accuracy of AI speech recognition is improved while optimizing the human ear hearing experience.

CN120048272APending Publication Date: 2025-05-27LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510246076.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the process of reducing background noise, existing speech signal processing technologies may lead to the loss of some details of the speech signal, especially the weakening of high-frequency components or syllable features, thereby affecting the accuracy of AI speech recognition.

Method used

By using different configuration parameters in the audio processing method to process the audio data, it is ensured that sufficient speech details are retained for AI model recognition without affecting the auditory experience of the human ear. Specifically, the first configuration parameter is used to optimize the human ear hearing experience, and the second configuration parameter is adjusted for the recognition requirements of the AI ​​model.

Benefits of technology

It realizes that while maintaining speech signal clarity and human ear hearing experience, sufficient speech details are retained to improve the accuracy of AI speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048272A_ABST
    Figure CN120048272A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method, which comprises the steps of obtaining audio data based on a first target instruction, the audio data being first sound data acquired by an audio acquisition device in real time, and the audio data being used for processing the audio data based on a first configuration parameter; based on a second target instruction, the audio data is processed by a second configuration parameter and converted into second sound data, and the second target audio data acts on a target model indicated by the second target instruction; wherein the first configuration parameter is different from the second configuration parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio data processing, and more particularly, to an audio processing method and an electronic device. Background Art

[0002] Voice signal processing technology is widely used in scenarios such as telephone calls, conference systems, voice assistants, and intelligent voice recognition. Existing voice processing technologies mainly optimize the user's auditory experience by reducing background noise and enhancing the clarity of human voices. However, excessive noise reduction may cause some details of the voice signal to be lost, especially the high-frequency components or syllable features are weakened, thus affecting the accuracy of AI voice recognition. Summary of the Invention

[0003] In view of this, the present disclosure provides an audio processing method and an electronic device.

[0004] One aspect of the present disclosure provides an audio processing method, including: obtaining audio data based on a first target instruction, where the audio data is first sound data collected in real time by an audio collection device; the audio data is used to process the audio data based on a first configuration parameter; processing the audio data into second sound data based on a second target instruction with a second configuration parameter, and the second target audio data acts on a target model indicated by the second target instruction; wherein, the first configuration parameter is different from the second configuration parameter.

[0005] According to an embodiment of the present disclosure, the audio processing method is applied to a first electronic device, and the first sound data is first sound data generated by an audio collection device wirelessly connected to the first electronic device based on a third configuration parameter for real-time collected ambient audio data, and the third configuration parameter belongs to the audio collection device.

[0006] According to an embodiment of the present disclosure, the second configuration parameter is obtained by the following operations: obtaining the third configuration parameter; obtaining the corresponding second configuration parameter according to the third configuration parameter, the second configuration parameter is different from the third configuration parameter, and the second configuration parameters corresponding to different third configuration parameters are different.

[0007] According to an embodiment of the present disclosure, the second configuration parameter is obtained by the following operations: obtaining a noise threshold parameter of the target model, where the noise threshold parameter represents the lowest noise level that the model can recognize; obtaining the second configuration parameter according to the noise threshold parameter, and the second configuration parameter is different from the third configuration parameter.

[0008] According to an embodiment of the present disclosure, the second configuration parameter is obtained by the following operations: obtaining environmental information; obtaining the second configuration parameter according to the environmental information, and the second configuration parameter is different from the third configuration parameter.

[0009] According to an embodiment of the present disclosure, the first configuration parameter includes a first noise reduction parameter, and the second configuration parameter includes a second noise reduction parameter. Processing the audio data with the second configuration parameter to convert it into second sound data includes: reducing the noise of the audio data based on the second noise reduction parameter, and the degree of noise reduction of reducing the noise of the audio data with the second noise reduction parameter is less than the degree of noise reduction of reducing the noise of the audio data with the first noise reduction parameter.

[0010] According to an embodiment of the present disclosure, the target model is used for voice authentication. The audio data at least includes a first part, and the first part represents the vocal characteristics in the audio data. Processing the audio data with the second configuration parameter to convert it into second sound data includes: reducing the noise of the first part based on the second noise reduction parameter, and the degree of noise reduction of reducing the noise of the first part with the second noise reduction parameter is less than the degree of noise reduction of reducing the noise of the first part with the first noise reduction parameter.

[0011] According to an embodiment of the present disclosure, based on the target model, according to the second sound data, generate the output content corresponding to the first sound data; obtain the confidence of the output content in real time; in response to the confidence being lower than the confidence threshold, adjust the second configuration parameter.

[0012] According to an embodiment of the present disclosure, in response to starting to process the audio data based on the first configuration parameter, save the audio data in real time to obtain third sound data; generate the output content corresponding to the first sound data, including: based on the target model, according to the second sound data and the third sound data, generate the output content corresponding to the first sound data.

[0013] Another aspect of the present disclosure provides an audio processing device, including: a first acquisition module, configured to obtain audio data based on a first target instruction, where the audio data is the first sound data collected in real time by an audio acquisition device, and the audio data is used to process the audio data based on the first configuration parameter; and a first processing module, configured to process the audio data with the second configuration parameter to convert it into second sound data based on a second target instruction, and the second target audio data acts on the target model indicated by the second target instruction; where the first configuration parameter is different from the second configuration parameter.

[0014] Another aspect of the present disclosure provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the audio processing method of any one of the foregoing embodiments.

[0015] Another aspect of the present disclosure provides a computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the audio processing method according to any one of the foregoing embodiments.

[0016] Another aspect of the present disclosure provides a computer program product, including a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the operations of the audio processing method in any of the foregoing embodiments are implemented. Description of the Drawings

[0017] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0018] Figure 1 A flowchart of the audio processing method according to an embodiment of the present disclosure is schematically shown;

[0019] Figure 2 A scenario diagram of the audio processing method according to an embodiment of the present disclosure is schematically shown;

[0020] Figure 3 A flowchart of obtaining a second configuration parameter in the audio processing method according to an embodiment of the present disclosure is schematically shown;

[0021] Figure 4 Another flowchart of obtaining a second configuration parameter in the audio processing method according to an embodiment of the present disclosure is schematically shown;

[0022] Figure 5 Another flowchart of obtaining a second configuration parameter in the audio processing method according to an embodiment of the present disclosure is schematically shown;

[0023] Figure 6 Another flowchart of the audio processing method according to an embodiment of the present disclosure is schematically shown;

[0024] Figure 7 Another flowchart of the audio processing method according to an embodiment of the present disclosure is schematically shown;

[0025] Figure 8 Another flowchart of the audio processing method according to an embodiment of the present disclosure is schematically shown;

[0026] Figure 9 Another flowchart of the audio processing method according to an embodiment of the present disclosure is schematically shown;

[0027] Figure 10 A block diagram of the audio processing device according to an embodiment of the present disclosure is schematically shown and

[0028] Figure 11 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Detailed Embodiments

[0029] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0030] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0031] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0032] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0033] In the embodiments of the present disclosure, in terms of the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the involved data (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and to safeguard the security of user personal information, network security, and national security.

[0034] The embodiments of the present disclosure provide an audio processing method, including: obtaining audio data based on a first target instruction, where the audio data is first sound data collected in real time by an audio acquisition device, and the audio data is used to process the audio data based on a first configuration parameter; processing the audio data into second sound data based on a second configuration parameter according to a second target instruction, and the second target audio data acts on a target model indicated by the second target instruction; where the first configuration parameter is different from the second configuration parameter.

[0035] Figure 1A flowchart of an audio processing method according to an embodiment of the present disclosure is schematically shown.

[0036] As Figure 1 shown, the audio processing method may at least include operations S110 to S120.

[0037] In operation S110, based on a first target instruction, audio data is obtained. The audio data is the first sound data collected in real time by an audio acquisition device. Among them, the audio data is used to process the audio data based on a first configuration parameter.

[0038] The first target instruction indicates control information for the electronic device to perform audio data acquisition and processing operations. This instruction can be generated by the device's control module, software application, operating system, or user input, and triggers the audio acquisition device to start acquiring the first sound data in the environment. The first target instruction can be triggered periodically or triggered based on an external event or when a specific condition is met. The first target instruction characterizes that the audio data is used for a first specific function, such as a call function, a singing function, a live broadcast function, etc.

[0039] For example, the first target instruction is generated or triggered by user input, and the audio acquisition is triggered through physical buttons, touch operations, or voice commands on devices such as smartphones, smart headphones, and voice assistants. For example, the first target instruction is an application program instruction, and the conferencing software automatically sends an instruction at the start of a call, asking the device to start recording or optimizing the audio. For example, the first target instruction is automatically triggered by the system. After the device detects a specific scenario (such as entering a call mode, a voice recognition mode, a noise reduction mode), it automatically sends an instruction to start audio processing. For example, the first compensation instruction is an external signal, and a control signal is received from an external device through Bluetooth, Wi-Fi, or other communication protocols, instructing the device to enter the audio acquisition mode.

[0040] The audio data refers to the first sound data obtained in real time by the audio acquisition device under a predetermined trigger condition. This data is the original or processed audio signal directly collected based on the current environmental sound source. The acquisition process of the audio data does not depend on pre-stored audio files, but dynamically captures the sound information at the current moment and is available for subsequent processing or transmission.

[0041] The audio acquisition device may include, but is not limited to, a microphone array, a headset microphone, an internal or external recording device of a smart device, etc., and can start audio acquisition based on the first target instruction.

[0042] The first target instruction characterizes that the audio data is used for a first specific function after the audio data is processed based on a first configuration parameter, such as a call function, a singing function, a live broadcast function, etc.

[0043] The first specific function refers to an application function in which audio data is processed with the first configuration parameters and then directly listened to by the human ear to meet the human ear's auditory perception requirements. For audio data that meets the human ear's auditory perception requirements, the processing objective is to improve sound quality, enhance speech clarity, reduce environmental noise interference, and optimize auditory comfort, so that the audio has a natural and smooth listening experience in various acoustic environments; the processing of such audio data can include noise reduction, echo suppression, automatic gain control, audio equalization, dynamic range compression, reverberation suppression, etc., to enhance the audio quality and make it more suitable for the perceptual characteristics of the human ear.

[0044] The first configuration parameters refer to a set of parameters used to control the processing method of audio data. This set of parameters is applicable to processing audio data to obtain specific audio effects, for example, a natural listening experience for direct listening by the human ear. The first configuration parameters can include, but are not limited to, noise reduction parameters, echo suppression parameters, automatic gain control parameters, audio equalization parameters, echo cancellation parameters, and so on.

[0045] Among them, the noise reduction parameter is used to determine the degree of suppression of background noise to reduce the interference of environmental noise on the human ear's auditory perception. The echo suppression parameter is used to reduce or eliminate echo interference caused by the acoustic environment and avoid the impact of echoes generated by room reflections or device speaker feedback during calls or audio playback on the listening experience. The automatic gain control parameter is used to dynamically adjust the gain of the audio signal to compensate for volume changes and ensure stable audio output. The audio equalization parameter is used to adjust the gain of different frequency bands to optimize the sound quality of speech.

[0046] Processing audio data based on the first configuration parameters means that after the audio data is real-time collected, before it is transmitted to the target processing device, target device, or target application program, it is processed according to the first configuration parameters to optimize the audio quality and make it suitable for the first specific function of the target processing device, target device, or target application program.

[0047] Specifically, for example, in a call scenario, the call sound and environmental sound real-time collected by the audio acquisition device are obtained, and the audio data is processed based on the first configuration parameters, so that the clarity of the speech signal is enhanced, environmental noise is suppressed, echoes and reverberations are reduced, the volume remains stable and adapts to different context changes, and the naturalness and intelligibility of the speech are optimized. Finally, audio data suitable for the perceptual characteristics of the human ear is obtained and sent to the call target through the call software.

[0048] In operation S120, based on the second target instruction, the audio data is processed with the second configuration parameters to be converted into second sound data, and the second sound data acts on the target model indicated by the second target instruction.

[0049] The second target instruction refers to an instruction activated by a specific trigger condition and different from the first target instruction. The instruction is used to control the processing method of audio data. The instruction can be triggered by user interaction, system events or preset rules.

[0050] The second sound data acts on the target model indicated by the second target instruction, represents the second sound data processed based on the second target instruction and the second configuration parameter, and is input into the target model to perform a second specific function different from the aforementioned first specific function, such as AI recognition, analysis or reasoning tasks, etc. The second sound data still maintains an audio format that can be heard by the human ear, but it undergoes specific signal optimization to enhance the target model's ability to parse audio information.

[0051] The target model can be an audio-to-text model that converts the speech content in the audio signal into corresponding text output.

[0052] The target model can also be an AI model that can directly process audio and directly generate corresponding outputs based on audio data. For example, it can generate environmental information about the sound source based on whether the audio contains gunshots, explosions, drums, etc.; for example, it can convert the input English audio into Chinese audio; for example, it can adjust the sound in the input audio to make the sound in the audio have funny sound effects. For example, it can analyze the emotional tendency of the audio data.

[0053] According to an embodiment of the present disclosure, the first configuration parameter is different from the second configuration parameter. For example, the first configuration parameter and the second configuration parameter may differ in specific parameter settings. Even if they are applicable to the same or similar audio processing technologies, the parameter adjustment methods, applicable scopes, and optimization degrees of the two are different to adapt to their respective audio processing goals. For example, the first configuration parameter is used to optimize the audio data to meet the auditory perception needs of the human ear and make the audio clearer, more natural, and understandable; while the second configuration parameter is used to process the audio data to meet the input requirements of the target model and ensure that the audio data retains as complete information as possible without affecting the AI ​​recognition or analysis capabilities. The two can use the same processing method (such as noise reduction, echo suppression, volume adjustment, etc.), but their parameter settings are different to achieve their respective optimization goals.

[0054] For example, in a call application, the noise reduction level based on the first configuration parameter is set high to remove environmental noise so that the human ear can clearly hear the speaker's voice. In a speech recognition task, the noise reduction level based on the second configuration parameter is set low to avoid excessive noise reduction affecting speech features, thereby affecting the speech recognition accuracy of the AI ​​model.

[0055] Specifically, for example, in a conference call application, the adjustment range of the automatic gain control based on the first configuration parameter is relatively large to ensure that the voice volume of each speaker is uniform and avoid sudden changes. In a speech emotion analysis task, the adjustment range of the automatic gain control based on the second configuration parameter is relatively small to avoid excessive changes to emotional features such as the pitch and speech rate of the speech, thereby maintaining the accurate recognition of emotional changes by the AI.

[0056] The first specific function corresponding to the first target instruction and the second specific function corresponding to the second target instruction can be executed simultaneously. That is, within the same time period, the audio data can be processed through the first configuration parameter to meet the auditory needs of the human ear, and can also be processed through the second configuration parameter to meet the input requirements of the AI model.

[0057] For example, in a conference call application, the device can execute call optimization and speech recognition functions simultaneously. On the one hand, the audio data can be processed through the first target instruction and the first configuration parameter to enhance the voice clarity, noise reduction, and echo suppression during the call, optimize the call quality, and provide it to the human ear. On the other hand, the audio data can also be processed through the second target instruction and the second configuration parameter to adapt the audio data to the speech recognition model and implement the real-time speech-to-text function to record the meeting content.

[0058] Specifically, for example, in a conference call application, after the audio acquisition device real-time acquires the first sound data, two processing channels are started simultaneously. The first processing channel performs voice clarity enhancement, noise reduction, echo suppression, etc. on the sound data based on the first configuration parameter, removes background noise through high-intensity noise reduction technology to ensure clear voice signal transmission, and makes the speech of the speaker more understandable to the human ear through frequency band optimization and dynamic range compression, and transmits the processed audio to the call component of the meeting. The second processing channel processes the same audio data based on the second configuration parameter, retains more speech details through lower-intensity noise reduction processing, and transmits it to the speech recognition model, enabling the speech recognition model to perform tasks such as more accurate speech-to-text processing or emotion analysis.

[0059] Figure 2 A scenario diagram of the audio processing method according to an embodiment of the present disclosure is schematically shown.

[0060] Based on the foregoing embodiments, the audio processing method is applied to a first electronic device. The first sound data is the first sound data generated by an audio acquisition device connected to the first electronic device based on a third configuration parameter for real-time acquired environmental audio data, and the third configuration parameter belongs to the audio acquisition device.

[0061] As Figure 2As shown, the scene includes a first electronic device 210 and an audio acquisition device 220. The audio acquisition device 220 is connected to the first electronic device 210 and is used to collect environmental audio data in real time. The audio acquisition device 220 generates first sound data based on a third configuration parameter. The third configuration parameter refers to a parameter used in the audio acquisition device to control and adjust various acquisition and processing behaviors of the environmental audio data collected in real time. For example, the third configuration parameter may include but is not limited to noise reduction parameters, echo suppression parameters, automatic gain control parameters, audio equalization parameters, echo cancellation parameters, etc. For example, the audio acquisition device may be an external Bluetooth headset, a wired headset, an external microphone, etc.

[0062] The first electronic device 210 further processes the first sound data through the first configuration parameter and / or the second configuration parameter. The processed audio data can be used for voice calls, voice recognition, identity authentication, or other voice applications. Through the processing of the first electronic device 210, the audio data can not only optimize the sound quality, but also provide high-quality input data for subsequent applications.

[0063] The network 230 is used to provide a medium for a communication link between the first electronic device 210 and the audio collection device 220. The network 230 may include various connection types, such as wired and wireless communication.

[0064] The audio acquisition device has processed the ambient audio data through the third configuration parameters to obtain the first sound data. If the first sound data adjusted by the third configuration parameters is directly used for AI model recognition, or for calls, some problems will arise, especially when the two are performed at the same time, that is, when AI recognition is performed while talking. Specifically, the third configuration parameters of the audio acquisition device are usually optimized for the human ear's hearing, such as noise reduction, gain adjustment, etc. These processes may remove or blur some audio features that are critical to AI model recognition (such as high-frequency syllables or speech details). If the optimized audio is directly sent to the AI ​​speech recognition system, it may cause recognition errors or reduced accuracy. The AI ​​system relies on the details and time-frequency features in the audio, and strong noise reduction and other processes may cause the loss of these features, especially in the case of large background noise or complex speech, which will greatly affect the recognition results.

[0065] Therefore, in order to ensure the accuracy and quality of the AI ​​model recognition and call tasks, the audio acquisition device uses the third configuration parameters to process the ambient audio data. After obtaining the first sound data, the electronic device uses the first configuration parameters and / or the second configuration parameters to further process the audio data according to the different requirements of the task, ensuring that AI model recognition and calls can be carried out simultaneously without interfering with each other, and achieving their respective ideal processing effects.

[0066] Figure 3 Schematically shows a flowchart of obtaining a second configuration parameter in an audio processing method according to an embodiment of the present disclosure.

[0067] As Figure 3 shown, on the basis of the foregoing embodiment, obtaining the second configuration parameter may include operations S310 to S320.

[0068] In operation S310, a third configuration parameter is obtained. Audio configuration parameters (third configuration parameters) related to its model and / or working mode are extracted from an audio acquisition device (such as a headset, a microphone, etc.). Different models of audio acquisition devices may have different hardware designs and processing capabilities, and audio acquisition devices in different working modes may also have different audio configuration parameters.

[0069] For example, the audio acquisition device of a certain headset may have multiple noise reduction modes built in, and automatically selects the most suitable mode according to the noise intensity of the surrounding environment. If in a noisy environment, the headset will automatically switch to the strong noise reduction mode to reduce environmental noise and enhance call quality, but this may sacrifice some high-frequency details; while in a quiet environment, the headset may select the light noise reduction mode to retain more human voice characteristics.

[0070] In operation S320, according to the third configuration parameter, a corresponding second configuration parameter is obtained. The second configuration parameter is different from the third configuration parameter, and the second configuration parameters corresponding to different third configuration parameters are different.

[0071] After obtaining the current working configuration of an audio acquisition device (such as a headset), second configuration parameters that match it and are used for subsequent audio processing are further generated according to these configuration parameters. The third configuration parameter is usually determined by the built-in processing algorithm, hardware characteristics or preset parameters of the audio acquisition device, and involves settings in aspects such as noise reduction, gain, and frequency response. The second configuration parameter is a parameter for further optimizing audio data to achieve a second specific function (such as AI recognition, voice verification, etc.) according to these device characteristics. Since different devices (i.e., different third configuration parameters) have different performances in audio processing, each third configuration parameter of a device corresponds to a different second configuration parameter to ensure that the audio data can achieve the best effect during the processing.

[0072] For example, the operating modes of a certain pair of headphones include a "strong noise reduction mode" and a "light noise reduction mode". If the headphones select the strong noise reduction mode (corresponding to a specific third configuration parameter), then the second configuration parameter may reduce the noise reduction effect (such as reducing the noise suppression intensity) according to the requirements of AI speech recognition, so as to retain more voice details for the AI system to process. If the headphones select the light noise reduction mode (another different third configuration parameter), then the second configuration parameter may slightly increase the noise reduction amplitude compared to the former, so as to remove noise as much as possible on the premise of ensuring the voice quality.

[0073] According to an embodiment of the present disclosure, the audio processing method further includes generating and sending a third instruction in response to a second target instruction, where the third instruction is used to adjust the audio collection device to a target operating mode, and the noise reduction degree of the third configuration parameter corresponding to the audio collection device in the target operating mode is less than the noise reduction degree of the third configuration corresponding to the audio collection device in the non-target operating mode. Since the audio collection device may have different audio configuration parameters in different operating modes, and the audio collection device processes the ambient audio data based on the third configuration parameter corresponding to the current mode, the quality of the human voice in the generated first sound data has been affected. To make up for this impact, a third instruction can be generated and sent to the audio collection device to adjust the operating mode of the audio collection device, so that the noise reduction of the audio collection device after adjusting the operating mode is reduced or there is no noise reduction.

[0074] For example, after receiving the second target instruction indicating that the audio data will be used for the AI text recognition model, a corresponding third instruction is generated and sent. After receiving the third instruction, the audio collection device converts the operating mode of the audio collection device to the "non-noise reduction mode", and adjusts the third configuration parameter corresponding to the audio collection device to the parameter for completely collecting ambient sound, so as to obtain the first sound data with undamaged human voice but with a certain amount of ambient noise, and then transmit it to the device. Before inputting it into the target model, the first sound data with undamaged human voice but with a certain amount of ambient noise can be directly processed based on the corresponding second configuration parameter, so as to extract the human voice information to the greatest extent from the first sound data with undamaged human voice but with a certain amount of ambient noise, and avoid irreversible impact on the human voice information caused by the built-in noise reduction of the audio collection device.

[0075] According to an embodiment of the present disclosure, the audio processing method further includes, in response to the target model being in a target state, generating and sending a third instruction, the third instruction being used to adjust the audio acquisition device to a target working mode, the noise reduction degree of the third configuration parameter corresponding to the audio acquisition device in the target working mode being less than the noise reduction degree of the third configuration corresponding to the audio acquisition device in a non-target working mode. In addition to receiving the second target instruction mentioned in the aforementioned embodiment, the timing for generating the third instruction may also be when the target model is in a target working state, such as in a speech recognition state, then directly generating the third instruction to control the audio acquisition device to switch to a corresponding working mode, so as to reduce or cancel the influence of the third configuration parameter of the audio acquisition device on the ambient audio data.

[0076] Figure 4 Another flowchart of obtaining the second configuration parameter in the audio processing method according to an embodiment of the present disclosure is schematically shown.

[0077] like Figure 4 As shown, based on the foregoing embodiment, obtaining the second configuration parameter may include operations S410 to S420.

[0078] In operation S410, the noise threshold parameter of the target model is obtained, and the noise threshold parameter characterizes the minimum noise level that the model can recognize. The noise threshold parameter of the target model can be the minimum background noise level that the target model can tolerate during normal recognition. This parameter defines the minimum noise intensity allowed in the audio data. When the background noise in the audio is lower than the threshold, the recognition accuracy of the model may be affected, resulting in recognition failure or misjudgment. This is because some AI speech recognition models do not perform best under completely noise-free or extremely low noise conditions. On the contrary, they can more effectively recognize speech information when working under certain background noise, because these models usually rely on certain noise characteristics for distinction and optimization during training.

[0079] For example, some AI speech recognition systems, such as deep learning-based models, may perform poorly in environments with little background noise. When the noise is reduced too low, the model may produce recognition errors or "drift" due to the lack of sufficient environmental audio features (such as changes in the spectrum). Therefore, the model's noise threshold parameter indicates this minimum noise value. When the threshold is exceeded, the model can still recognize the speech content normally, but below this threshold, the model may not adapt or lose effective recognition ability.

[0080] For example, if the noise threshold of a certain target model is set to -60 dB, this means that when the background noise is lower than -60 dB, the model may exhibit poor recognition accuracy due to the lack of sufficient environmental features or details of the voice signal, and may even fail to correctly recognize the voice. On the contrary, if the background noise is higher than -60 dB, the model can stably perform voice recognition through these noise features.

[0081] In operation S420, according to the noise threshold parameter, a second configuration parameter is obtained, and the second configuration parameter is different from the third configuration parameter. During audio processing, based on the noise threshold parameter of the target model, the second configuration parameter suitable for the model is dynamically adjusted and generated or selected. The difference between this second configuration parameter and the third configuration parameter indicates that it is optimized for the recognition requirements of the AI model, rather than simply relying on the built-in processing parameters of the audio acquisition device (such as headphones, microphones, etc.). In other words, the noise threshold parameter is used as a reference to guide how to adjust the processing method of audio data, so as to ensure that the audio input can meet the best recognition conditions of the AI model.

[0082] In another embodiment, obtaining the second configuration parameter in the audio processing method may further include: obtaining the corresponding second configuration parameter according to the target model, and the second configuration parameters corresponding to different target models are different.

[0083] In the embodiments of the present disclosure, according to the specific requirements of the device hardware and the target model, the corresponding second configuration parameter has been preset for each target model at the factory. Specifically, when the device leaves the factory, it will determine a set of appropriate second configuration parameters according to the target application scenario (such as voice recognition, voice authentication, etc.) and the requirements of the AI model used, and store these parameters in the device to ensure that when the user uses the device, the device can perform optimized processing according to these preset parameters.

[0084] For example, for a smart device, at the factory, the manufacturer can preset the second configuration parameter according to the noise threshold parameter of the target AI model. Assuming that the noise threshold of this model is -60 dB, the device will be configured with a noise reduction level suitable for this noise threshold, such as set to "medium" at the factory. In this way, even in a relatively noisy environment, the AI model can effectively perform voice recognition through the background noise features in the audio signal. When the user uses the device, the device will automatically apply these preset second configuration parameters, so as to ensure that the audio input can meet the best recognition conditions of the AI model and avoid the decline in recognition accuracy caused by too low or too high noise.

[0085] The second configuration parameter can also be set differently according to the model data of various models collected in advance. For example, the feature and requirement information of different speech recognition AI models can be collected in advance, and different second configuration parameters can be set according to the requirements of different models for recognition accuracy, noise threshold, background noise processing ability, etc. Based on this, the device can automatically adjust the processing method of the audio signal according to the specific AI model installed, ensuring that the audio input can be best adapted to the target AI model, thereby optimizing the accuracy and fluency of speech recognition.

[0086] Specifically, for example, the same intelligent device may support multiple different AI speech recognition models when leaving the factory. For instance, the device can install a voice assistant model and a voice control system model at the same time. Suppose the noise threshold of the voice assistant model is relatively low and is suitable for processing scenarios with relatively low background noise, while the voice control system model needs to maintain a high recognition accuracy in a high-noise environment. To meet the requirements of these different models, the device will preset different second configuration parameters for each model when leaving the factory.

[0087] On the basis of the embodiment shown in Figure 4 the method for obtaining the audio processing method may further include: adding specific noise to the audio data in response to the noise in the audio data being lower than the noise threshold. When the background noise intensity in the audio data collected by the audio acquisition device is lower than the lowest noise level acceptable to the target model (i.e., the noise threshold parameter), specific background noise needs to be added to the audio data to restore a certain background noise without interfering with the voice signal, so as to optimize the speech recognition process of the AI model and avoid the reduction of the model recognition accuracy or recognition failure due to too low background noise.

[0088] Figure 5 Another flowchart of obtaining the second configuration parameter of the audio processing method according to an embodiment of the present disclosure is schematically shown.

[0089] As Figure 5 shown, on the basis of the foregoing embodiment, obtaining the second configuration parameter may include operations S510 to S520.

[0090] In operation S510, environmental information is obtained. The environmental information can be various external data and features used to describe the current audio acquisition environment, and these data and features can reflect the background conditions of audio acquisition. The environmental information includes but is not limited to: environmental noise level, environmental type (such as indoor or outdoor), background noise source, surrounding physical space characteristics (such as crowd density, traffic flow, etc.), and inferences or judgments based on these environmental characteristics (such as the noisy or quiet degree of the current environment).

[0091] The acquisition of environmental information can be achieved through various means. For example, environmental information can be collected based on positioning and / or microphones and / or sensors. Specifically, for example, the un-denoised environmental sounds collected by the microphone, the scene speculation based on positioning and sensors, etc. These information can be used to determine the current environmental type.

[0092] In operation S520, according to the environmental information, a second configuration parameter is obtained, and the second configuration parameter is different from the third configuration parameter. Based on the collected environmental information (such as noise level, environmental type, noise source, etc.), the device dynamically adjusts the configuration parameters of audio processing. Specifically, the second configuration parameter is selected according to the environmental information to adjust factors such as the noise reduction intensity and gain in the audio processing process, ensuring that the audio received by the AI speech recognition model adapts to the current environment and maximizing the speech recognition effect. The second configuration parameter is different from the third configuration parameter to adapt to different processing requirements and objectives, ensuring that under different environmental conditions, the audio data can better meet the requirements of the speech recognition model.

[0093] For example, in the scenario where the mobile phone is connected to headphones, the headphones (audio acquisition device) perform strong noise reduction processing (processed with the third configuration parameter) in a street environment, significantly reducing the background noise in the audio but also taking away some details of the human voice. The mobile phone terminal obtains environmental information based on positioning and / or microphones and / or sensors, and will judge that the current environment is a noisy street according to the environmental information, so a weaker noise reduction parameter (the second configuration parameter) is selected to avoid further loss of the already affected human voice characteristics during subsequent processing, thus ensuring that the speech recognition model can better restore the details of the human voice.

[0094] For example, in a relatively quiet coffee shop environment, the noise reduction amplitude of the headphones is small, so the human voice is better retained. At this time, after the mobile phone terminal recognizes the environmental information, an appropriate noise reduction parameter (the second configuration parameter) is selected. Because the human voice is better retained at this time, a slightly stronger noise reduction level relative to the previous example can be appropriately selected to further remove the noise.

[0095] Figure 6 Another flowchart of the audio processing method according to an embodiment of the present disclosure is schematically shown.

[0096] As Figure 6 shown, on the basis of the foregoing embodiment, the first configuration parameter includes a first noise reduction parameter, the second configuration parameter includes a second noise reduction parameter, processing the audio data with the second configuration parameter is converted into second sound data, and operation S120 may include operation S610.

[0097] In operation S610, the audio data is denoised based on the second noise reduction parameter, and the degree of noise reduction of denoising the audio data with the second noise reduction parameter is less than the degree of noise reduction of denoising the audio data with the first noise reduction parameter.

[0098] The first noise reduction parameter is used for strong noise reduction of audio data, usually for noise reduction during audio output on headphones or other devices. The purpose of this noise reduction process is to optimize the user's auditory experience by maximizing the removal of background noise to enhance the clarity of the human voice. In contrast, the second noise reduction parameter is used for noise reduction when the audio data is transmitted to the speech recognition model, and the intensity of noise reduction is lower than that of the first noise reduction parameter. The lower noise reduction degree of the second noise reduction parameter enables the speech recognition model to better retain the details of the speech signal in the audio and avoid information loss caused by excessive noise reduction.

[0099] According to an embodiment of the present disclosure, the second noise reduction parameter is greater than a preset noise reduction parameter threshold, and the noise reduction degree of noise-reducing the audio data with the second noise reduction parameter is less than the noise reduction degree of noise-reducing the audio data with the first noise reduction parameter. The preset noise reduction parameter threshold is a predetermined standard, usually representing the lowest intensity of noise reduction. The second noise reduction parameter must be greater than this threshold to ensure that the noise reduction effect is sufficiently effective. If the noise reduction intensity is too low, there may be too much noise in the second sound data, resulting in the speech recognition model being unable to accurately recognize the speech signal. Specifically, the intensity setting of the second noise reduction parameter needs to balance two objectives: avoiding excessive noise reduction (i.e., if it is too much lower than the threshold, it will cause too much audio noise and affect the model accuracy), and ensuring that the noise reduction intensity is sufficient to remove the background noise, thereby improving the accuracy of speech recognition.

[0100] Figure 7 Another flowchart of the audio processing method according to an embodiment of the present disclosure is schematically shown.

[0101] As Figure 7 shown, on the basis of the foregoing embodiment, the target model is used for voice authentication, the audio data includes at least a first part, the first part represents the vocal characteristics in the audio data, and operation S120 may include operation S710.

[0102] In operation S710, the first part is noise-reduced based on the second noise reduction parameter, and the noise reduction degree of noise-reducing the first part with the second noise reduction parameter is less than the noise reduction degree of noise-reducing the first part with the first noise reduction parameter.

[0103] The target model can be a model used to authenticate an individual based on the speech characteristics in the audio data. The target model analyzes specific characteristics (such as voiceprint, speech characteristics, formants, etc.) in the audio data to compare with the pre-stored speech template or reference data in the database, thereby confirming the identity corresponding to the human voice in the audio data. For example, when a user issues an instruction to the speech recognition system through a smart audio device, the target model analyzes the user's speech, extracts the speech characteristics (such as formants) therein and matches them with the registered voice template, thereby confirming the user's identity and verifying whether to allow access to the smart home control permission.

[0104] In an audio signal, there are not only speech components, but also other non-speech signals such as background noise and ambient sound. To effectively extract and analyze human voices, the audio signal can be divided into different parts. The "first part" is specifically separated from these signals and represents the part with speech characteristics. It includes the data part in the speech signal that can best reflect the personalized characteristics of the speaker, such as the frequency components in the speech, the distribution of formants, the pronunciation patterns of the speech, etc.

[0105] When processing audio data, for the first part of the audio data (i.e., the part containing human voice characteristics), when using the second noise reduction parameter for noise reduction, the intensity of noise reduction is weak, maintaining more speech details and characteristics. In contrast, in noise reduction for the human ear's listening perception, the first noise reduction parameter has a higher noise reduction intensity for the first part to remove more background noise. The second noise reduction parameter protects more speech characteristics (such as formants, high-frequency components of speech, etc.) and avoids damaging the key information in the speech due to excessive noise reduction. Especially in scenarios such as speech verification and recognition, it helps to improve the accuracy of the target model for speech verification and recognition.

[0106] Figure 8 Another flowchart of the audio processing method according to an embodiment of the present disclosure is schematically shown.

[0107] As Figure 8 shown, on the basis of the foregoing embodiment, the audio processing method may further include operations S810 to S830.

[0108] In operation S810, based on the target model, according to the second sound data, the output content corresponding to the first sound data is generated. The output content may be the result generated by a generative model (such as a speech recognition, speech-to-text, or speech analysis model) through analyzing and processing the first sound data. For example, in a speech recognition system, the first sound data may be the audio data after noise reduction and enhancement, and the generated output content may be the text transcription of the audio content, the parsing result of a voice command, or other voice-related information. The output content may also be the verification result output by voice identity verification.

[0109] For example, the first sound data is the audio from a telephone conference. After processing, the generative speech recognition model converts the audio into text content. For example, a sentence in the audio "Let's start discussing the next topic" is processed by the model, and the output content is the text "Let's start discussing the next topic".

[0110] In operation S820, the confidence level of the output content is obtained in real time. After the target model processes the audio data and generates the output result, the accuracy of the output content is evaluated to obtain a confidence value. The confidence value is usually used to measure the reliability of the generated result and reflects the confidence of the model in the output content.

[0111] In operation S830, in response to the confidence level being lower than the confidence threshold, the second configuration parameter is adjusted. When the confidence level is lower than the preset confidence threshold (usually a numerical limit, such as 0.7), it indicates that there is a relatively large uncertainty in the accuracy of the result, and the second configuration parameter will be automatically adjusted according to this feedback.

[0112] For example, on a noisy street, the noise reduction system of the headphones has processed the surrounding noise, making the human voice in the audio transmitted to the mobile phone side clearer and the background noise less. However, due to the large street noise, the headphones need to adopt a stronger noise reduction strategy to remove the noise, so the quality of the human voice has suffered a relatively strong loss. And if this data has been adjusted by the second configuration parameter before being input into the target model and has undergone a relatively large amount of noise reduction again, then the confidence level of the output content of the target model will be relatively small. When it is detected that the confidence level is small, the second configuration parameter is readjusted. For example, the noise reduction amplitude is reduced to improve the accuracy of model recognition.

[0113] Figure 9 Another flowchart of the audio processing method according to an embodiment of the present disclosure is schematically shown.

[0114] As Figure 9 shown, on the basis of the foregoing embodiment, the audio processing method may include operations S910 to S920.

[0115] In operation S910, in response to starting to process the audio data based on the first configuration parameter, the audio data is saved in real time to obtain the third sound data. For example, when starting to process the audio data based on the first configuration parameter, it means the start of a call, recording, or meeting. Then, from this time point, the audio data is saved in real time as the third sound data. It should be noted that the audio data saved here in real time is the audio data before being processed by the first configuration parameter, rather than the audio data after being processed by the first configuration parameter.

[0116] In operation S920, based on the target model, the output content corresponding to the first sound data is generated according to the second sound data and the third sound data. When the target model starts to execute the corresponding function, it is not necessarily at the same moment when processing the audio data based on the first configuration parameter. It may be at a certain moment after starting to process the audio data based on the first configuration parameter. At this time, the second sound data is incomplete relative to the first sound data. For example, during a call, the user suddenly remembers to turn on the AI recognition function and asks the AI model to summarize the content of this call. Then, when the AI recognition function is turned on, the second sound data obtained in real time is incomplete relative to the entire call.

[0117] By combining the saved audio data (the third sound data) with the audio data processed in real time (the second sound data), it is ensured that even if the first target instruction and the second target instruction are not synchronized, the target model can obtain complete audio data.

[0118] According to the embodiments of the present disclosure, before generating the output content corresponding to the first sound data based on the target model according to the second sound data and the third sound data, it may further include: processing the third sound data according to the second configuration parameter. After processing the third sound data in the same way as the second sound data, the audio input to the model can be made unified, enabling the model to process the second sound data and the third sound data more accurately.

[0119] Figure 10 A block diagram of an audio processing apparatus according to an embodiment of the present disclosure is schematically shown.

[0120] As Figure 10 shown, the audio processing apparatus 1000 may include a first obtaining module 1010 and a first processing module 1020.

[0121] The first obtaining module 1010 is configured to obtain audio data based on a first target instruction. The audio data is the first sound data collected in real time by an audio collection device. Among them, the audio data is used to process the audio data based on a first configuration parameter. In some embodiments, the first obtaining module 1010 may be configured to perform the operation S110 in the above audio processing method, which will not be elaborated here.

[0122] The first processing module 1020 is configured to process the audio data into second sound data based on a second target instruction with a second configuration parameter. The second target audio data acts on the target model indicated by the second target instruction, where the first configuration parameter is different from the second configuration parameter. In some embodiments, the first processing module 1020 may be configured to perform the operation S120 in the above audio processing method, which will not be elaborated here.

[0123] According to an embodiment of the present disclosure, the audio processing device may include an audio parameter acquisition module, and the audio parameter acquisition module may include a second acquisition module and a third acquisition module.

[0124] The second acquisition module is used to acquire a third configuration parameter. In some embodiments, the second acquisition module may be used to perform operation S310 in the above audio processing method, which will not be elaborated here.

[0125] The third acquisition module is used to obtain a corresponding second configuration parameter according to the third configuration parameter. The second configuration parameter is different from the third configuration parameter, and the second configuration parameters corresponding to different third configuration parameters are different. In some embodiments, the third acquisition module may be used to perform operation S320 in the above audio processing method, which will not be elaborated here.

[0126] According to an embodiment of the present disclosure, the audio processing device may include a fourth acquisition module and a fifth acquisition module.

[0127] The fourth acquisition module is used to acquire a noise threshold parameter of the target model. The noise threshold parameter characterizes the lowest noise level that the model can recognize. In some embodiments, the fourth acquisition module may be used to perform operation S410 in the above audio processing method, which will not be elaborated here.

[0128] The fifth acquisition module is used to obtain a second configuration parameter according to the noise threshold parameter. The second configuration parameter is different from the third configuration parameter. In some embodiments, the fifth acquisition module may be used to perform operation S420 in the above audio processing method, which will not be elaborated here.

[0129] According to an embodiment of the present disclosure, the audio processing device may include a sixth acquisition module and a seventh acquisition module.

[0130] The sixth acquisition module is used to acquire environmental information. In some embodiments, the sixth acquisition module may be used to perform operation S510 in the above audio processing method, which will not be elaborated here.

[0131] The seventh acquisition module is used to obtain a second configuration parameter according to the environmental information. The second configuration parameter is different from the third configuration parameter. In some embodiments, the seventh acquisition module may be used to perform operation S520 in the above audio processing method, which will not be elaborated here.

[0132] According to an embodiment of the present disclosure, the first configuration parameter includes a first noise reduction parameter, the second configuration parameter includes a second noise reduction parameter, and the audio data is processed with the second configuration parameter to be converted into second sound data. The first processing module may include a first noise reduction module.

[0133] The first noise reduction module is used to reduce the noise of the audio data based on the second noise reduction parameter. The degree of noise reduction of the audio data with the second noise reduction parameter is less than the degree of noise reduction of the audio data with the first noise reduction parameter. In some embodiments, the first noise reduction module may be used to perform the operation S610 in the above audio processing method, which will not be elaborated here.

[0134] According to an embodiment of the present disclosure, the target model is used for voice authentication. The audio data at least includes a first part, and the first part characterizes the vocal characteristics in the audio data. The first processing module may include a second noise reduction module.

[0135] The second noise reduction module is used to reduce the noise of the first part based on the second noise reduction parameter. The degree of noise reduction of the first part with the second noise reduction parameter is less than the degree of noise reduction of the first part with the first noise reduction parameter. In some embodiments, the second noise reduction module may be used to perform the operation S710 in the above audio processing method, which will not be elaborated here.

[0136] According to an embodiment of the present disclosure, the audio processing device may include a first generation module, a ninth acquisition module, and a first adjustment module.

[0137] The first generation module is used to generate the output content corresponding to the first sound data based on the target model and according to the second sound data. In some embodiments, the first generation module may be used to perform the operation S810 in the above audio processing method, which will not be elaborated here.

[0138] The ninth acquisition module is used to acquire the confidence of the output content in real time. In some embodiments, the ninth acquisition module may be used to perform the operation S820 in the above audio processing method, which will not be elaborated here.

[0139] The first adjustment module is used to adjust the second configuration parameter in response to the confidence being lower than the confidence threshold. In some embodiments, the first adjustment module may be used to perform the operation S830 in the above audio processing method, which will not be elaborated here.

[0140] According to an embodiment of the present disclosure, the audio processing device may include a first saving module and a second generation module.

[0141] The first saving module is used to save the audio data in real time in response to starting to process the audio data based on the first configuration parameter, and obtain the third sound data. In some embodiments, the first saving module may be used to perform the operation S910 in the above audio processing method, which will not be elaborated here.

[0142] The second generation module is used to generate output content corresponding to the first voice data based on the target model according to the second voice data and the third voice data. In some embodiments, the second generation module may be used to perform operation S920 in the above audio processing method, which will not be elaborated here.

[0143] According to embodiments of the present disclosure, any plurality of modules, sub-modules, units, and sub-units, or at least part of the functions of any of them can be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure can be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, such as hardware or firmware, or can be implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.

[0144] For example, any plurality of the first acquisition module 1010 and the first processing module 1020 can be combined and implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units can be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to embodiments of the present disclosure, at least one of the first acquisition module 1010 and the first processing module 1020 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, such as hardware or firmware, or can be implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the first acquisition module 1010 and the first processing module 1020 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.

[0145] It should be noted that the data processing system part in the embodiments of the present disclosure corresponds to the data processing method part in the embodiments of the present disclosure. For the description of the data processing system part, please refer to the data processing method part specifically, and details will not be repeated here.

[0146] Figure 11 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 11 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0147] As Figure 11 shown, the electronic device 1100 according to an embodiment of the present disclosure includes a processor 1101, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1102 or the program loaded from the storage section 1108 into the random access memory (RAM) 1103. The processor 1101 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), and so on. The processor 1101 can also include on-board memory for caching purposes. The processor 1101 can include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiments of the present disclosure.

[0148] In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are stored. The processor 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. The processor 1101 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 1102 and / or the RAM 1103. It should be noted that the program can also be stored in one or more memories other than the ROM 1102 and the RAM 1103. The processor 1101 can also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.

[0149] According to an embodiment of the present disclosure, the electronic device 1100 may further include an input / output (I / O) interface 1105, and the input / output (I / O) interface 1105 is also connected to the bus 1104. The electronic device 1100 may further include one or more of the following components connected to the input / output (I / O) interface 1105: an input portion 1106 including a keyboard, a mouse, etc.; an output portion 1107 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 1108 including a hard disk, etc.; and a communication portion 1109 including a network interface card such as a LAN card, a modem, etc. The communication portion 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the input / output (I / O) interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed so that a computer program read from thereon is installed into the storage portion 1108 as needed.

[0150] According to an embodiment of the present disclosure, the method flow according to the embodiment of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication portion 1109, and / or installed from the removable medium 1111. When the computer program is executed by the processor 1101, the above functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0151] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.

[0152] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0153] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the ROM 1102 and / or RAM 1103 and / or ROM 1102 and RAM 1103 described above.

[0154] An embodiment of the present disclosure also includes a computer program product, which includes a computer program. The computer program contains program code for executing the method provided by the embodiment of the present disclosure. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the control method provided by the embodiment of the present disclosure.

[0155] When the computer program is executed by the processor 1101, the above functions defined in the system / apparatus of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0156] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 1109, and / or installed from the removable medium 1111. The program code included in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above. According to the embodiments of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0158] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. An audio processing method, comprising: Based on the first target instruction, audio data is obtained, where the audio data is first sound data collected in real time by an audio collection device; The audio data is used to process the audio data based on a first configuration parameter; Based on a second target instruction, the audio data is processed with a second configuration parameter to convert into second sound data, the second sound data acting on a target model indicated by the second target instruction; The first configuration parameter is different from the second configuration parameter.

2. The method according to claim 1 is applied to a first electronic device, wherein the first sound data is generated by the audio acquisition device wirelessly connected to the first electronic device based on a third configuration parameter for real-time collected ambient audio data, and the third configuration parameter belongs to the audio acquisition device.

3. According to the method of claim 2, the second configuration parameter is obtained by the following operation: Acquire the third configuration parameter; According to the third configuration parameter, the corresponding second configuration parameter is obtained, the second configuration parameter is different from the third configuration parameter, and different third configuration parameters respectively correspond to different second configuration parameters.

4. According to the method of claim 2, the second configuration parameter is obtained by the following operation: Acquire a noise threshold parameter of the target model, wherein the noise threshold parameter represents the lowest noise level that can be recognized by the model; The second configuration parameter is obtained according to the noise threshold parameter, and the second configuration parameter is different from the third configuration parameter.

5. According to the method of claim 2, the second configuration parameter is obtained by the following operation: Obtain environmental information; The second configuration parameter is obtained according to the environment information, where the second configuration parameter is different from the third configuration parameter.

6. The method according to claim 1, wherein the first configuration parameter comprises a first noise reduction parameter, the second configuration parameter comprises a second noise reduction parameter, and the converting the audio data into the second sound data by processing the audio data with the second configuration parameter comprises: The audio data is denoised based on the second noise reduction parameter, and a noise reduction degree of the audio data denoised by the second noise reduction parameter is less than a noise reduction degree of the audio data denoised by the first noise reduction parameter.

7. The method according to claim 6, wherein the target model is used for voice authentication, and the audio data comprises at least a first part, wherein the first part represents a human voice feature in the audio data; The converting the audio data into second sound data by processing the audio data with the second configuration parameter includes: The first portion is denoised based on the second denoise parameter, and a denoise degree of the first portion denoised by the second denoise parameter is less than a denoise degree of the first portion denoised by the first denoise parameter.

8. The method according to claim 1, further comprising: Based on the target model, generating output content corresponding to the first sound data according to the second sound data; Acquiring the confidence of the output content in real time; In response to the confidence being below a confidence threshold, the second configuration parameter is adjusted.

9. The method according to claim 1, further comprising: In response to starting to process the audio data based on the first configuration parameter, saving the audio data in real time to obtain third sound data; Based on the target model, output content corresponding to the first sound data is generated according to the second sound data and the third sound data.

10. An electronic device comprising An audio acquisition device, used for acquiring audio data; processor, at least one processor; as well as A memory connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the following operations: Based on the first target instruction, audio data is obtained, where the audio data is first sound data collected in real time by an audio collection device; the audio data is used to process the audio data based on the first configuration parameter; Based on a second target instruction, the audio data is processed with a second configuration parameter to convert into second sound data, the second target audio data acting on a target model indicated by the second target instruction; The first configuration parameter is different from the second configuration parameter.