An audio processing method and related equipment

By converting audio signals into frequency domain signals and identifying and attenuating interference audio domain parameters, the problem of slow response speed of devices in noisy environments is solved, enabling devices to respond quickly to voice commands.

CN122090859APending Publication Date: 2026-05-26ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG GEELY HLDG GRP CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, devices respond slowly when receiving voice commands in noisy environments, resulting in untimely device wake-up or control.

Method used

The audio signal to be processed is converted into a frequency domain signal. By determining the spectral characteristic value and sparse characteristic value, the interference audio domain parameters are identified and attenuated to obtain the audio signal used to wake up or control the device function.

Benefits of technology

It achieves real-time suppression of interference sounds and improves the device's response rate to voice commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090859A_ABST
    Figure CN122090859A_ABST
Patent Text Reader

Abstract

The application provides an audio processing method and related equipment, the audio processing method comprises the following steps: converting a first audio signal to be processed into a first frequency domain signal containing a plurality of frequency domain parameters; determining a spectral feature value and a sparse feature value of each of the frequency domain parameters in the first frequency domain signal; determining a target frequency domain parameter from the plurality of frequency domain parameters according to the spectral feature value and the sparse feature value, and determining an attenuation parameter of the target frequency domain parameter; performing gain reduction of the attenuation parameter on the target frequency domain parameter in the first frequency domain signal to obtain a second frequency domain signal, and performing time domain conversion on the second frequency domain signal to obtain a second audio signal. In the application, the response speed of the equipment to the voice instruction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method and related equipment. Background Technology

[0002] With the development of intelligent technology in devices, users can wake up or control the functions of the devices using voice. However, when a user speaks, the presence of noise in the environment or the device's own output prompts can cause interference in the user's voice, affecting the waking up or control of the device's functions.

[0003] In exemplary technologies, audio editing tools or plugins are configured in the device to suppress interference sounds in the speech. However, these audio editing tools or plugins suppress interference sounds offline, which means that the user's voice cannot wake up or control the functions in the device in a timely manner, resulting in a slow response rate of the device to voice commands. Summary of the Invention

[0004] Based on the aforementioned technological status, this application provides an audio processing method and related equipment to solve the problem of slow response rate of devices to voice commands.

[0005] To achieve the above-mentioned technical objectives, this application proposes the following technical solution:

[0006] In a first aspect, this application provides an audio processing method, including: The first audio signal to be processed is converted into a first frequency domain signal containing multiple frequency domain parameters; Determine the spectral feature value and sparse feature value of each frequency domain parameter in the first frequency domain signal. The spectral feature value is used to indicate the energy flatness of the frequency band in which the frequency domain parameter is located, and the sparse feature value is used to indicate the sparsity of the frequency domain parameter. Based on the spectral feature value and the sparse feature value, a target frequency domain parameter is determined among multiple frequency domain parameters, and an attenuation parameter of the target frequency domain parameter is determined, wherein the spectral feature value of the target frequency domain parameter is greater than a spectral threshold, and the sparse feature value of the target frequency domain parameter is greater than a sparse threshold. The target frequency domain parameter in the first frequency domain signal is reduced by the attenuation parameter to obtain a second frequency domain signal, and the second frequency domain signal is converted into a time domain to obtain a second audio signal. The second audio signal is used to wake up or control the function of the device.

[0007] In some implementations, determining the attenuation parameter of the target frequency domain parameter includes: Determine a first difference between the spectral feature value and the spectral threshold, and determine a second difference between the sparse feature value and the sparse threshold; A target probability value is determined based on the first difference, the second difference, the weight corresponding to the spectral feature value, and the weight corresponding to the sparse feature value. The target probability value is used to indicate the probability that the target frequency domain parameter is a frequency domain feature of the interfering audio. The attenuation parameter is determined based on the target probability value, and the target probability value is positively correlated with the attenuation parameter.

[0008] In some implementations, determining the attenuation parameter of the target frequency domain parameter includes: The preset value is set as the attenuation parameter of the target frequency domain parameter.

[0009] In some implementations, the step of reducing the attenuation parameter of the target frequency domain parameter in the first frequency domain signal to obtain the second frequency domain signal includes: For the target frequency domain parameter in the first frequency domain signal, the attenuation parameter is reduced by the gain, and the frequency domain parameter to be processed adjacent to the target frequency domain parameter is determined; The target parameter corresponding to the frequency domain parameter to be processed is processed so that the difference between the first target parameter and the second target parameter is less than a preset difference, so as to obtain the second frequency domain signal. The first target parameter is the frequency or decibel corresponding to the frequency domain parameter to be processed after processing, and the second target parameter is the frequency or decibel corresponding to the target frequency domain parameter after loss.

[0010] In some implementations, determining the sparse feature value of each frequency domain parameter in the first frequency domain signal includes: Determine the sparse vector corresponding to the frequency domain parameters; Based on the number of frames corresponding to the frame-segmentation of the first audio signal, the sparse vector is normalized to obtain the sparse feature values ​​of the frequency domain parameters.

[0011] In some implementations, determining the spectral characteristic value of each frequency domain parameter in the first frequency domain signal includes: Obtain the lower limit frequency value, upper limit frequency value, and frequency width of the frequency band interval where the frequency domain parameter is located; The spectral characteristic value is determined based on the lower limit frequency value, the upper limit frequency value, the frequency width, and the amplitude of the frequency domain parameter.

[0012] In some implementations, determining the target frequency domain parameter from among a plurality of frequency domain parameters based on the spectral feature value and the sparse feature value includes: Obtain the rate of change of the square of the amplitude of the frequency domain parameter; Based on the rate of change, the spectral feature value, and the sparse feature value, a target frequency domain parameter is determined from among a plurality of frequency domain parameters, wherein the rate of change corresponding to the target frequency domain parameter is greater than a change threshold.

[0013] Secondly, this application provides an audio processing apparatus, comprising: A conversion module is used to convert the first audio signal to be processed into a first frequency domain signal containing multiple frequency domain parameters; The first determining module is used to determine the spectral feature value and the sparse feature value of each frequency domain parameter in the first frequency domain signal. The spectral feature value is used to indicate the energy flatness of the frequency band in which the frequency domain parameter is located, and the sparse feature value is used to indicate the sparsity of the frequency domain parameter. The second determining module is used to determine a target frequency domain parameter from among a plurality of frequency domain parameters based on the spectral feature value and the sparse feature value, and to determine the attenuation parameter of the target frequency domain parameter, wherein the spectral feature value of the target frequency domain parameter is greater than a spectral threshold, and the sparse feature value of the target frequency domain parameter is greater than a sparse threshold. The attenuation module is used to reduce the target frequency domain parameter in the first frequency domain signal by the attenuation parameter to obtain a second frequency domain signal, and to perform time-domain conversion on the second frequency domain signal to obtain a second audio signal. The second audio signal is used to wake up or control the function of the device.

[0014] Thirdly, this application provides an audio processing apparatus, including a memory and a processor, wherein, The memory is connected to the processor and is used to store programs; The processor is configured to implement the audio processing method as described in the first aspect or any implementation thereof by running a program in the memory.

[0015] Fourthly, this application provides a vehicle, characterized in that the vehicle includes an audio processing device that implements the audio processing method as described in the first aspect or any implementation thereof.

[0016] Fifthly, this application provides a computer program product, which, when executed by a processor, implements the audio processing method as described in the first aspect or any implementation thereof.

[0017] In a sixth aspect, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the audio processing method as described in the first aspect or any implementation thereof.

[0018] This application provides an audio processing method and related device, which converts a first audio signal to be processed into a first frequency domain signal containing multiple frequency domain parameters, determines the spectral feature value and sparse feature value of each frequency domain parameter in the first frequency domain signal, determines a target frequency domain parameter among the multiple frequency domain parameters based on the spectral feature value and sparse feature value, determines the attenuation parameter of the target frequency domain parameter, and performs a gain reduction on the attenuation parameter of the target frequency domain parameter in the first frequency domain to obtain a second frequency domain signal, and performs time domain conversion on the second frequency domain signal to obtain a second audio signal used to wake up or control the function of the device. In this application, the spectral characteristic value of the target frequency domain parameter is greater than the spectral threshold, and the sparsity characteristic value of the target frequency domain parameter is greater than the sparsity threshold. This indicates that the energy flatness of the frequency band where the target frequency domain parameter is located is relatively large and the sparsity of the target frequency domain parameter is relatively small. Therefore, the target frequency domain parameter is the frequency domain characteristic of the interference sound. Thus, the target frequency domain parameter is attenuated to suppress the interference sound in the audio signal, thereby achieving real-time suppression of interference sound in the audio signal. This enables the device to quickly wake up or control the function based on the real-time suppressed audio signal, improving the device's response rate to voice commands. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0020] Figure 1 A flowchart of an audio processing method provided in this application embodiment Figure 1 .

[0021] Figure 2 A flowchart of an audio processing method provided in this application embodiment Figure 2 .

[0022] Figure 3 A flowchart of an audio processing method provided in this application embodiment Figure 3 .

[0023] Figure 4 A flowchart of an audio processing method provided in this application embodiment Figure 4 .

[0024] Figure 5 A flowchart of an audio processing method provided in this application embodiment Figure 5 .

[0025] Figure 6 This is a schematic diagram of the functional modules of an audio processing device provided in an embodiment of this application.

[0026] Figure 7 This is a structural schematic diagram of a vehicle provided in an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] It should be noted that the user information (including but not limited to electrical equipment information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0029] With the development of intelligent technology in devices, users can wake up or control functions via voice. However, when a user speaks, ambient noise or the device's own output tones can introduce interference, affecting the ability to wake up or control the device's functions. The following example, using a vehicle as an example, illustrates the impact of interference on a vehicle.

[0030] The microphones in the vehicle's cabin collect the sound inside the cabin, process it through algorithms, and then use it for human-to-human interaction (such as making calls), human-computer interaction (such as voice assistants), and local amplification (such as karaoke without a microphone).

[0031] Of the sounds collected by the vehicle, the driver's and passengers' voices are the target signals, while other sound signals can be considered "interference" signals. For example, when the vehicle turns, a turn signal tone is played in the cabin, similar to a "click" sound; a similar tone is also played when operating the cabin's control panel. These interference signals are called "turn signal tone interference."

[0032] The characteristics of turn signal tones are that their playback and cessation times are unpredictable and directly related to the vehicle's driving status, making them a typical form of non-steady-state noise. This renders noise reduction schemes such as spectral subtraction and Wiener filtering, which are designed for steady-state noise, unsuitable. The extremely short duration of turn signal tones, typically between a few milliseconds and tens of milliseconds, significantly reduces the performance of adaptive tracking algorithms. The lack of a defined pitch, making it impossible to sing the turn signal tones using "Do, Re, Mi," renders pitch-tracking algorithms ineffective. Furthermore, the envelope of the turn signal tones only has an "attack" and a very rapid "decay," lacking significant "sustain" and "release," which places high demands on the real-time processing capabilities of the algorithms.

[0033] Because of the aforementioned characteristics of turn signal audible tones, although their overall duration is limited, improper processing can impact certain applications of the in-cabin microphones. For example, when using a voice assistant, if the turn signal audible tone coincides with the voice command and the tone is not properly processed, its significant short-term energy can mask the voice command. Similarly, during microphone-less karaoke, if the turn signal audible tone picked up by the microphone is not properly processed and is played back by the cabin speakers, the residual frequency of the tone in this positive feedback loop can cause persistent ringing if it happens to be in the loop phase. If the amplitude is too large, it can even cause feedback.

[0034] Currently, depending on the perspective from which the problem is viewed, different approaches are being taken to deal with "interference similar to turn signal warning sounds".

[0035] From a music processing perspective: Turn prompts are click sounds, and different de-click techniques exist to eliminate them. Some convenient audio editing software includes automatic click removal effects that can remove click sounds from selected audio files. Professional audio restoration and enhancement software boasts powerful de-click (click removal) modules that allow precise control over the sensitivity and processing range of the restoration, even distinguishing between large, medium, and small click sounds; however, these require the purchase of a separate de-click plugin. Other software allows manual selection of click sounds for removal; or other commercial audio editing software or plugins support de-click functionality. The de-click functions in these commercial software programs are mostly for audio file editing and are typically offline operations.

[0036] From the perspective of speech enhancement: Considering non-steady-state noise, neural network-based denoising models can be used to suppress it. This requires incorporating noise data similar to turn signal audible sounds during the training process. However, no dedicated denoising model has been found specifically for this type of noise. Given that the envelope of a turn signal audible sound only has an "onset" and an extremely rapid "decay," classic denoising schemes model it as impulse noise. However, this is inaccurate. The mathematical expression for impulse noise is a delta function, with an infinitely large amplitude, infinitely small width, and an area of ​​1. The Fourier transform of the delta function is a constant, meaning it has the same energy across all frequencies—that is, its spectrum is flat and spans the entire frequency range. However, the spectrum of a turn signal audible sound is not full-range and lasts for tens of milliseconds. Therefore, denoising models are not very effective at suppressing short-duration interference. Furthermore, for general neural network denoising models, the larger the model size and the more complex the network structure, the stronger the denoising capability and the better the suppression effect on "interference similar to turn signal audible sounds." However, the cockpit chip has limited computing and storage resources for voice processing, which cannot meet the requirements of noise reduction models with large size and complex structure.

[0037] In view of this, the embodiments of this application aim to provide an audio processing method and related device, which converts a first audio signal to be processed into a first frequency domain signal containing multiple frequency domain parameters, determines the spectral feature value and sparse feature value of each frequency domain parameter in the first frequency domain signal, determines a target frequency domain parameter among the multiple frequency domain parameters based on the spectral feature value and sparse feature value, determines the attenuation parameter of the target frequency domain parameter, and performs a gain reduction on the attenuation parameter of the target frequency domain parameter in the first frequency domain to obtain a second frequency domain signal, and performs time domain transformation on the second frequency domain signal to obtain a second audio signal used to wake up or control the function of the device. In this application, the spectral characteristic value of the target frequency domain parameter is greater than the spectral threshold, and the sparsity characteristic value of the target frequency domain parameter is greater than the sparsity threshold. This indicates that the energy flatness of the frequency band where the target frequency domain parameter is located is relatively large and the sparsity of the target frequency domain parameter is relatively small. Therefore, the target frequency domain parameter is the frequency domain characteristic of the interference sound. Thus, the target frequency domain parameter is attenuated to suppress the interference sound in the audio signal, thereby achieving real-time suppression of interference sound in the audio signal. This enables the device to quickly wake up or control the function based on the real-time suppressed audio signal, improving the device's response rate to voice commands.

[0038] The audio processing method proposed in this application will be described in detail below with reference to various embodiments.

[0039] Reference Figure 1 , Figure 1 A flowchart of an audio processing method provided in this application embodiment Figure 1 .like Figure 1 As shown, the audio processing method provided in this embodiment includes: Step S101: Convert the first audio signal to be processed into a first frequency domain signal containing multiple frequency domain parameters.

[0040] In this embodiment, the executing entity is an audio processing device. The audio processing device can be a component in a vehicle used for processing audio, the vehicle itself, or any terminal device capable of suppressing interference in the audio. For ease of description, the term "device" will be used to refer to the audio processing device below.

[0041] The device collects the user's voice, which is defined as the first audio signal to be processed. This first audio signal carries noise; therefore, it is a noisy signal in the time domain, which can be expressed as: y(t) = x(t) + n(t), where x is the voice signal without interference, n is the noise signal, y is the first audio signal, and t is the discrete-time indicator. The noise signal n can represent various additive noises. When the first audio signal is used to wake up or control functions in the vehicle, additive noise refers to "interference sounds similar to turn signal alerts." To avoid confusion, "click" is used instead of n. Therefore, the noisy signal in the time domain is transformed into: y(t) = x(t) + click(t).

[0042] After obtaining the first audio signal, a frequency domain transformation is performed on it to obtain a frequency domain signal containing multiple frequency domain parameters. This frequency domain signal is defined as the first frequency domain signal. The first frequency domain signal is represented as: Y(f) = X(f) + Click(f), where Y, X, and Click represent the Fourier transforms of the first audio signal, the speech signal, and click, respectively, and f is the frequency domain identifier. That is, Y(f) = X(f) + Click(f) is obtained by using the short-time Fourier transform (STFT) of y(t) = x(t) + click(t). It should be noted that the STFT requires windowing, which is not reflected in Y(f) = X(f) + Click(f).

[0043] Step S102: Determine the spectral characteristic value and sparse characteristic value of each frequency domain parameter in the first frequency domain signal. The spectral characteristic value is used to indicate the energy flatness of the frequency band in which the frequency domain parameter is located, and the sparse characteristic value is used to indicate the sparsity of the frequency domain parameter.

[0044] After obtaining the first frequency domain signal, analysis is performed based on the frequency domain characteristics of Click.

[0045] Specifically, a characteristic of howling is that the initial stage of howling is characterized by a surge of signal with a wide spectral distribution, resembling a vertical line in the frequency distribution on the spectrogram. Click signals exhibit similar characteristics. This key frequency domain property is spectral sparsity: within a single time frame, the vector containing the STFT amplitude coefficients will have very few large elements in Y, while during a click, it will have elements with equal or approximately equal amplitudes in certain frequency bands f. This means that this vector is less sparsy in the time frame where a click occurs than in the time frame without a click. Therefore, the frequency domain characteristics of the interference tone in Y are less sparsy.

[0046] Another frequency domain characteristic of click signals can be measured using spectral flatness. Spectral flatness is calculated as the ratio of the geometric mean to the arithmetic mean of Y. As mentioned before, the spectrum of a click signal is not flat across the entire frequency range. However, the short-duration, high-energy nature of the click signal still results in approximate flatness in certain frequency intervals. In other words, if the energy flatness is relatively high within the frequency interval containing the frequency parameter, then this frequency parameter can be confirmed as a frequency domain characteristic of interference.

[0047] As can be seen from the above, the frequency domain characteristics of Click are: low sparsity and high energy flatness in the frequency band.

[0048] To address this, the device determines the spectral eigenvalues ​​and sparsity eigenvalues ​​of each frequency domain parameter f in the first frequency domain signal. The spectral eigenvalues ​​refer to the energy flatness of the frequency band in which the frequency domain parameter f resides, while the sparsity eigenvalues ​​indicate the sparsity of the frequency domain parameter f. The sparsity eigenvalues ​​can be determined from the sparse vector of the frequency domain parameter, and the spectral eigenvalues ​​can be calculated using the same method as for spectral flatness.

[0049] Step S103: Based on the spectral feature value and the sparse feature value, determine the target frequency domain parameter among multiple frequency domain parameters, and determine the attenuation parameter of the target frequency domain parameter, wherein the spectral feature value of the target frequency domain parameter is greater than the spectral threshold, and the sparse feature value of the target frequency domain parameter is greater than the sparse threshold.

[0050] After obtaining the spectral eigenvalues ​​and sparse eigenvalues ​​corresponding to each frequency domain parameter f, a target frequency domain parameter is determined from among the various spectral parameters based on these eigenvalues ​​and sparse eigenvalues. The spectral eigenvalues ​​of the target frequency domain parameters are greater than a spectral threshold, and the sparse eigenvalues ​​are less than a sparse threshold. The spectral threshold and sparse threshold can be any suitable values; for example, the spectral threshold could be 0.9, and the sparse threshold 0.1. It can be understood that the device uses frequency domain parameters with spectral eigenvalues ​​greater than the spectral threshold and sparse eigenvalues ​​less than the sparse threshold as target frequency domain parameters.

[0051] Since the spectral eigenvalue of the target frequency domain parameter is greater than the spectral threshold, and the sparse eigenvalue of the target frequency domain parameter is less than the sparse threshold, the target frequency domain parameter satisfies the frequency domain characteristics of Click. Therefore, if the speech segment corresponding to the target frequency domain parameter is found to contain interference, it is necessary to suppress the frequency corresponding to the target frequency domain parameter. To this end, the device determines the attenuation parameter corresponding to the target frequency domain parameter. The attenuation parameter can be a complex frequency domain mask, i.e., Mask(f). Mask(f) can be a preset value, such as 0.1, which is used as the attenuation parameter of the target frequency domain parameter.

[0052] Step S104: Reduce the attenuation parameter of the target frequency domain parameter in the first frequency domain signal to obtain the second frequency domain signal, and perform time domain conversion on the second frequency domain signal to obtain the second audio signal. The second audio signal is used to wake up or control the function of the device.

[0053] After determining the attenuation parameter, the target frequency domain parameter in the first frequency domain signal is subtracted by the attenuation parameter to obtain the second frequency domain signal. In one example, the attenuation parameter can be used to subtract the decibel level of the target frequency domain parameter. For example, if Mask(f) is 0.1, corresponding to -20dB, then the decibel level of the target frequency domain parameter is subtracted by 20bB. In another example, the attenuation parameter can be used to subtract the frequency of the target frequency domain parameter. For example, if Mask(f) is 0.1, then the frequency of the target frequency domain parameter is modified to 0.1 times the frequency.

[0054] After subtracting the gain from each target frequency domain parameter in the first frequency domain signal, the second frequency domain signal is obtained. The device then performs a time-domain transformation on the second frequency domain signal, that is, performs an inverse short-time Fourier transform on the second frequency domain signal, to obtain the second audio signal. The second audio signal can be used to wake up or control the device's functions.

[0055] In this embodiment, the interference is suppressed by modeling the spectral characteristics of "interference similar to turn signal warning sounds" and then performing complex frequency domain masking. This is a real-time processing scheme that improves the device's response rate to voice commands. In addition, modeling based on the spectral characteristics of "interference similar to turn signal warning sounds" rather than simply modeling it as impulse noise improves the accuracy of mathematical modeling. Furthermore, this embodiment is not based on data-trained neural network noise reduction, which avoids the data requirements during model training and reduces the computational and storage requirements when deploying chips in the vehicle's cockpit.

[0056] In this embodiment, the first audio signal to be processed is converted into a first frequency domain signal containing multiple frequency domain parameters. The spectral characteristic value and sparse characteristic value of each frequency domain parameter in the first frequency domain signal are determined. Based on the spectral characteristic value and sparse characteristic value, a target frequency domain parameter is determined among the multiple frequency domain parameters, and an attenuation parameter for the target frequency domain parameter is determined. The target frequency domain parameter in the first frequency domain is then reduced by the attenuation parameter to obtain a second frequency domain signal. The second frequency domain signal is then time-domain transformed to obtain a second audio signal used for waking up or controlling the device. In this embodiment, the spectral characteristic value of the target frequency domain parameter is greater than a spectral threshold, and the sparse characteristic value of the target frequency domain parameter is greater than a sparse threshold. This indicates that the energy flatness of the frequency band containing the target frequency domain parameter is relatively high, and the sparsity of the target frequency domain parameter is relatively low. Therefore, the target frequency domain parameter represents the frequency domain characteristic of interference sounds. Attenuating the target frequency domain parameter suppresses interference sounds in the audio signal, achieving real-time suppression of interference sounds in the audio signal. This allows the device to quickly wake up or control functions based on the real-time suppressed audio signal, improving the device's response rate to voice commands.

[0057] Reference Figure 2 , Figure 2 A flowchart of an audio processing method provided in this application embodiment Figure 2 ,based on Figure 1 In the embodiment shown, step S103 includes: Step S201: Determine the first difference between the spectral feature value and the spectral threshold, and determine the second difference between the sparse feature value and the sparse threshold.

[0058] In this embodiment, the device determines a first difference between the spectral feature value and the spectral threshold, and determines a second difference between the sparse feature value and the sparse threshold.

[0059] Step S202: Determine the target probability value based on the first difference, the second difference, the weights corresponding to the spectral feature values ​​and the weights corresponding to the sparse feature values. The target probability value is used to indicate the probability that the target frequency domain parameter is a frequency domain feature of the interfering audio.

[0060] After determining the first difference and the second difference, the target probability value is determined based on the weights corresponding to the first difference, the second difference, the spectral feature values, and the sparse feature values. The target probability value refers to the probability that the target frequency domain parameter is a frequency domain feature of the interfering audio.

[0061] For example, the target probability value can be represented by the following formula:

[0062] in, This refers to the click-or-not probability, which is the target probability value; H indicates the click-or-not (whether there is a notification sound) state. The weight is represented by w; the feature width of the mapping function is represented by M; and the mapping function is represented by tanh or sigmoid. These are sparse eigenvalues; These are frequency domain eigenvalues; The sparse threshold; This is the spectral threshold.

[0063] The above formula can be obtained through neural network training, that is... w, M, and T (frequency domain threshold and sparse threshold) can all be obtained through neural network training.

[0064] Step S203: Determine the attenuation parameter based on the target probability value. The target probability value and the attenuation parameter are positively correlated.

[0065] Once the target probability value is obtained, the attenuation parameter can be determined based on the target probability value. The larger the target probability value of the target frequency domain parameter, the larger the attenuation parameter will be. In other words, the target probability value and the attenuation parameter are positively correlated.

[0066] In this embodiment, a target probability value corresponding to the target frequency domain parameter is determined, and an attenuation parameter is determined based on the target probability value, so that the device adaptively attenuates the target frequency domain parameter based on the probability that the audio corresponding to the target frequency domain parameter is interference.

[0067] Figure 3 A flowchart of an audio processing method provided in this application embodiment Figure 3 ,based on Figure 1 or Figure 2 In the embodiment shown, step S104 includes: Step S301: Reduce the attenuation parameter of the target frequency domain parameter in the first frequency domain signal, and determine the frequency domain parameter to be processed that is adjacent to the target frequency domain parameter.

[0068] In this embodiment, applying a gain to the target frequency domain parameter using an attenuation parameter can easily lead to spectral distortion. For example, if the frequency domain parameters f that meet the conditions appear at intervals in the frequency domain segment, a sawtooth-shaped gain will occur if the gain is applied to these interval frequency domain parameters f.

[0069] To address this, we can take the target frequency domain parameters as the center and apply corresponding gain reduction in the adjacent frequency domains on the left and right. This gain reduction can be achieved through notch filtering.

[0070] Specifically, the device first performs attenuation loss reduction on the target frequency domain parameter in the first frequency domain signal. Then, the device determines the frequency domain parameters adjacent to the target frequency domain parameter as the frequency domain parameters to be processed; that is, the frequency domain parameters adjacent to the target frequency domain parameter on the left and right are used as the frequency domain parameters to be processed. For example, the frequency domain parameters to be processed and the target frequency domain parameter on two adjacent frequency bands are represented as follows: [f F1_min-2 f F1_min-1 f F1_min f F1_min+1 f F1_min+2 ]; where f F1_min For target frequency domain parameters.

[0071] Step S302: Process the target parameter corresponding to the frequency domain parameter to be processed so that the difference between the first target parameter and the second target parameter is less than a preset difference, so as to obtain the second frequency domain signal. Here, the first target parameter is the frequency or decibel corresponding to the processed frequency domain parameter to be processed, and the second target parameter is the frequency or decibel corresponding to the target frequency domain parameter after loss.

[0072] After determining the frequency domain parameters to be determined, the target parameters of the frequency domain parameters to be determined are processed, that is, the frequency or decibel of the frequency domain parameters to be determined is processed, so that the difference between the first target parameter and the second target parameter is less than a preset difference, and the second frequency domain signal is obtained. Here, the first target parameter is the frequency or decibel corresponding to the processed frequency domain parameters, and the second target parameter is the frequency or decibel corresponding to the target frequency domain parameters after loss.

[0073] In this embodiment, after the target frequency domain parameter is reduced, the frequency or decibel of the frequency domain parameter to be processed adjacent to the target frequency domain parameter is also reduced to avoid distortion in the frequency domain.

[0074] Figure 4 A flowchart of an audio processing method provided in this application embodiment Figure 4 .based on Figures 1 to 3 In any of the embodiments shown, step S101 includes: Step S401: Determine the sparse vector corresponding to the frequency domain parameters.

[0075] Step S402: Based on the number of frames corresponding to the frame processing of the first audio signal, normalize the sparse vector to obtain the sparse feature values ​​of the frequency domain parameters.

[0076] In this embodiment, the device determines the sparse vector corresponding to the frequency domain parameters, and normalizes the sparse vector based on the number of frames processed by the first audio signal to obtain the sparse feature values ​​of the frequency domain parameters.

[0077] For example, the output of the STFT, as a vector of time (represented by time t) over a certain frequency domain parameter f, is expressed by the following formula:

[0078] Where M represents the number of frames, the above formula represents the vector composed of the STFT output results of the past M frames of data in the frequency domain parameter f.

[0079] Therefore, the normalized eigenvalues ​​measuring sparsity over the frequency domain parameter f are expressed by the formula:

[0080] For each frequency domain parameter f, the sparse eigenvalue F1 can be calculated using the normalized formula described above. Considering the temporal characteristics of the interference signal "click," the frame number M should not be too large. If the frame length is 8ms, M should not exceed 10; if the frame length is 16ms, M should not exceed 5, in order to achieve fast tracking of the click signal. The sparse eigenvalue F1 ranges from 0 to 1. The closer it is to 0, the sparser the frequency domain parameter f, and the more it matches the frequency domain characteristics of the click signal.

[0081] In this embodiment, the sparse feature values ​​corresponding to the frequency domain parameters can be accurately obtained by using the sparse vector of the frequency domain parameters and the number of frames corresponding to the frame-segmentation of the first audio signal.

[0082] Figure 5 A flowchart of an audio processing method provided in this application embodiment Figure 5 .based on Figures 1 to 4 In any of the embodiments shown, step S101 includes: Step S501: Obtain the lower limit frequency value, upper limit frequency value, and frequency width of the frequency band interval where the frequency domain parameters are located.

[0083] Step S502: Determine the spectral characteristic value based on the lower limit frequency value, the upper limit frequency value, the frequency width, and the amplitude of the frequency domain parameters.

[0084] In this embodiment, the device obtains the lower limit frequency value, upper limit frequency value, and frequency width of the frequency band interval where the frequency domain parameter is located, and then determines the spectral characteristic value based on the lower limit frequency value, upper limit frequency value, frequency width, and amplitude of the frequency domain parameter.

[0085] For example, the click signal, despite its short duration and high energy, still results in an approximately flat frequency range. Therefore, the spectral flatness can be measured using the spectral characteristics of piecewise flatness. The spectral flatness (spectral characteristic value) of the frequency domain parameters for the current frame is expressed as:

[0086] in, b is the amplitude of the frequency parameter. l and b u is the frequency index of the upper and lower boundaries of frequency band interval b, and w is the width of frequency band interval b. The spectral characteristic value F2 ranges from 0 to 1. The closer it is to 1, the flatter the frequency band interval is, and the more it conforms to the frequency domain characteristics of the click signal. The frequency band interval b can be determined based on the number of points in the Fourier transform in the STFT. For example, for a signal with a sampling rate of 16 kHz, when the number of points in the Fourier transform is 256, the width of frequency band interval b is 800 Hz; when the number of points in the Fourier transform is 1024, the width of frequency band interval b is 400 Hz.

[0087] In this embodiment, the spectral characteristic value is accurately determined by the lower limit frequency value, upper limit frequency value, and frequency width of the frequency band interval where the frequency domain parameter is located.

[0088] In one embodiment, the frequency domain characteristics of "interference similar to turn signal beep sounds" are not limited to the two mentioned above. For example, based on the characteristic that the click signal only has an "onset" and extremely rapid "fading," in the time domain, for each frequency domain f... Tracking can be performed, for example, if the rate of change (which can be measured by derivative or difference) of the signal is too large for 2 or 3 consecutive frames and exceeds a certain threshold, it can also be used as a characteristic of the frequency domain f of the click signal.

[0089] To this end, the device obtains the rate of change of the square of the amplitude of the frequency domain parameter, i.e. The rate of change. For example, the rate of change can be calculated by squared amplitudes of frequency domain parameters over two or three consecutive frames.

[0090] Based on the rate of change, spectral eigenvalues, and sparsity eigenvalues, the target frequency domain parameter is determined from multiple frequency domain parameters. Specifically, the frequency domain parameter whose rate of change is greater than a change threshold, whose spectral eigenvalue is greater than a spectral threshold, and whose sparsity eigenvalue is less than a sparsity threshold is identified as the target frequency domain parameter.

[0091] In this embodiment, the frequency domain parameters are accurately determined to be the target frequency domain parameters that conform to the click signal by using the amplitude squared change rate, spectral characteristic value, and sparse characteristic value of the frequency domain parameters.

[0092] Corresponding to the above-described audio processing method, this application also provides an audio processing apparatus. Figure 6 This is a schematic diagram of a module of an audio processing device provided in an embodiment of this application. The audio processing device 600 provided in this embodiment includes: The conversion module 610 is used to convert the first audio signal to be processed into a first frequency domain signal containing multiple frequency domain parameters. The first determining module 620 is used to determine the spectral characteristic value and the sparse characteristic value of each frequency domain parameter in the first frequency domain signal. The spectral characteristic value is used to indicate the energy flatness of the frequency band in which the frequency domain parameter is located, and the sparse characteristic value is used to indicate the sparsity of the frequency domain parameter. The second determining module 630 is used to determine a target frequency domain parameter from multiple frequency domain parameters based on spectral feature values ​​and sparse feature values, and to determine the attenuation parameter of the target frequency domain parameter, wherein the spectral feature value of the target frequency domain parameter is greater than a spectral threshold, and the sparse feature value of the target frequency domain parameter is greater than a sparse threshold. The attenuation module 640 is used to attenuate the target frequency domain parameter in the first frequency domain signal to obtain the second frequency domain signal, and then perform time domain conversion on the second frequency domain signal to obtain the second audio signal. The second audio signal is used to wake up or control the function of the device.

[0093] In some implementations, the audio processing device 600 is also used for: Determine the first difference between the spectral eigenvalues ​​and the spectral threshold, and determine the second difference between the sparse eigenvalues ​​and the sparse threshold; The target probability value is determined based on the weights corresponding to the first difference, the second difference, the spectral feature value, and the sparse feature value. The target probability value is used to indicate the probability that the target frequency domain parameter is a frequency domain feature of the interfering audio. The attenuation parameter is determined based on the target probability value, and the target probability value and the attenuation parameter are positively correlated.

[0094] In some implementations, the audio processing device 600 is also used for: The preset value is set as the attenuation parameter of the target frequency domain parameter.

[0095] In some implementations, the audio processing device 600 is also used for: For the target frequency domain parameter in the first frequency domain signal, the attenuation parameter is reduced, and the frequency domain parameter to be processed adjacent to the target frequency domain parameter is determined. The target parameter corresponding to the frequency domain parameter to be processed is processed so that the difference between the first target parameter and the second target parameter is less than a preset difference, so as to obtain the second frequency domain signal. Here, the first target parameter is the frequency or decibel corresponding to the frequency domain parameter to be processed after processing, and the second target parameter is the frequency or decibel corresponding to the target frequency domain parameter after loss.

[0096] In some implementations, the audio processing device 600 is also used for: Determine the sparse vector corresponding to the frequency domain parameters; Based on the number of frames corresponding to the frame-segmentation of the first audio signal, the sparse vector is normalized to obtain the sparse feature values ​​of the frequency domain parameters.

[0097] In some implementations, the audio processing device 600 is also used for: Obtain the lower limit frequency value, upper limit frequency value, and frequency width of the frequency domain parameter within the frequency band interval; The spectral characteristic values ​​are determined based on the lower limit frequency value, the upper limit frequency value, the frequency bandwidth, and the amplitude of the frequency domain parameters.

[0098] In some implementations, the audio processing device 600 is also used for: Obtain the rate of change of the square of the amplitude of the frequency domain parameter; Based on the rate of change, spectral characteristic value, and sparse characteristic value, the target frequency domain parameter is determined from multiple frequency domain parameters, and the rate of change corresponding to the target frequency domain parameter is greater than the change threshold.

[0099] The audio processing apparatus and audio processing method provided in the above embodiments of this application belong to the same concept and can execute the audio processing method provided in any of the above embodiments of this application. They have the corresponding functional modules and beneficial effects for executing the audio processing method. Technical details not described in detail in this embodiment can be found in the specific processing content of the audio processing method provided in the above embodiments of this application, and will not be repeated here.

[0100] The functions implemented by each module in the audio processing device can be implemented by the same or different processors, and this application embodiment does not limit this.

[0101] It should be understood that the modules in the above audio processing device can be implemented by a processor calling firmware. For example, the system includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of each module of the device. The processor can be a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal to the device or external to the system. Alternatively, the modules in the system can be implemented as hardware circuits. By designing the hardware circuits, some or all of the module functions can be implemented. The hardware circuits can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above modules are implemented by designing the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through a configuration file, thereby implementing the functions of some or all of the above modules. All modules of the above audio processing device can be implemented entirely by a processor calling firmware, entirely by hardware circuits, or partially by a processor calling firmware with the remaining parts implemented by hardware circuits.

[0102] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.

[0103] As can be seen, each module in the above audio processing device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor types.

[0104] Furthermore, the modules in the above audio processing device can be integrated in whole or in part, or they can be implemented independently. In one implementation, these modules are integrated together and implemented in the form of a System-on-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the modules of the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.

[0105] This application provides a structural schematic diagram of a vehicle, see [link / reference] Figure 7 As shown, the vehicle includes a memory 700 and a processor 710; wherein the memory 700 is connected to the processor 710 and is used to store programs; the processor 710 is used to implement the audio processing method disclosed in any of the above embodiments by running the programs stored in the memory 700.

[0106] Specifically, the vehicle may also include: a bus, a communication interface 720, an input device 730, an output device 740, and an audio processing device 750. The vehicle may also include a data transceiver module, an image monitoring module, and a signal monitoring module.

[0107] The processor 710, memory 700, communication interface 720, input device 730, output device 740, and audio processing device 750 are interconnected via a bus. Among them: A bus can include a pathway for transmitting information between various components in a vehicle.

[0108] The processor 710 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0109] The processor 710 may include a main processor, as well as a baseband chip, modem, etc.

[0110] The memory 700 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 700 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0111] Input device 730 may include a device for receiving data and information input by a user, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0112] Output device 740 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0113] The communication interface 720 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0114] The processor 710 executes the program stored in the memory 700 and calls other devices, which can be used to implement the various steps of any of the audio processing methods provided in the above embodiments of this application.

[0115] It should be noted that the vehicle can be an in-vehicle terminal, mobile phone, wearable device or server, etc.; or it can be a vehicle that includes an in-vehicle terminal, etc.

[0116] This application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in the memory through the data interface to execute the audio processing method described in any of the above embodiments. For details of the processing and its beneficial effects, please refer to the above-described embodiments of the audio processing method.

[0117] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the audio processing methods according to various embodiments of this application as described in any of the above embodiments of this specification.

[0118] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the power device, as a standalone firmware package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0119] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor to perform the steps of the audio processing method according to various embodiments of this application described in any of the above embodiments of this specification, specifically implementing the steps of the above audio processing method.

[0120] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0121] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0122] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0123] The units of the apparatus in the various embodiments of this application can be merged, divided, and deleted according to actual needs.

[0124] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0125] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0126] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or as firmware functional modules or sub-modules.

[0127] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer firmware, or a combination of both. To clearly illustrate the interchangeability of hardware and firmware, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or firmware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0128] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly using hardware, firmware units executed by a processor, or a combination of both. The firmware unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0129] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0130] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An audio processing method, characterized in that, include: The first audio signal to be processed is converted into a first frequency domain signal containing multiple frequency domain parameters; Determine the spectral feature value and sparse feature value of each frequency domain parameter in the first frequency domain signal. The spectral feature value is used to indicate the energy flatness of the frequency band in which the frequency domain parameter is located, and the sparse feature value is used to indicate the sparsity of the frequency domain parameter. Based on the spectral feature value and the sparse feature value, a target frequency domain parameter is determined among multiple frequency domain parameters, and an attenuation parameter of the target frequency domain parameter is determined, wherein the spectral feature value of the target frequency domain parameter is greater than a spectral threshold, and the sparse feature value of the target frequency domain parameter is greater than a sparse threshold. The target frequency domain parameter in the first frequency domain signal is reduced by the attenuation parameter to obtain a second frequency domain signal, and the second frequency domain signal is converted into a time domain to obtain a second audio signal. The second audio signal is used to wake up or control the function of the device.

2. The audio processing method according to claim 1, characterized in that, The determination of the attenuation parameter of the target frequency domain parameter includes: Determine a first difference between the spectral feature value and the spectral threshold, and determine a second difference between the sparse feature value and the sparse threshold; A target probability value is determined based on the first difference, the second difference, the weight corresponding to the spectral feature value, and the weight corresponding to the sparse feature value. The target probability value is used to indicate the probability that the target frequency domain parameter is a frequency domain feature of the interfering audio. The attenuation parameter is determined based on the target probability value, and the target probability value is positively correlated with the attenuation parameter.

3. The audio processing method according to claim 1, characterized in that, The determination of the attenuation parameter of the target frequency domain parameter includes: The preset value is set as the attenuation parameter of the target frequency domain parameter.

4. The audio processing method according to claim 1, characterized in that, The step of reducing the attenuation parameter of the target frequency domain parameter in the first frequency domain signal to obtain the second frequency domain signal includes: For the target frequency domain parameter in the first frequency domain signal, the attenuation parameter is reduced by the gain, and the frequency domain parameter to be processed adjacent to the target frequency domain parameter is determined; The target parameter corresponding to the frequency domain parameter to be processed is processed so that the difference between the first target parameter and the second target parameter is less than a preset difference, so as to obtain the second frequency domain signal. The first target parameter is the frequency or decibel corresponding to the frequency domain parameter to be processed after processing, and the second target parameter is the frequency or decibel corresponding to the target frequency domain parameter after loss.

5. The audio processing method according to claim 1, characterized in that, Determining the sparse feature value of each frequency domain parameter in the first frequency domain signal includes: Determine the sparse vector corresponding to the frequency domain parameters; Based on the number of frames corresponding to the frame-segmentation of the first audio signal, the sparse vector is normalized to obtain the sparse feature values ​​of the frequency domain parameters.

6. The audio processing method according to claim 1, characterized in that, Determining the spectral feature value of each frequency domain parameter in the first frequency domain signal includes: Obtain the lower limit frequency value, upper limit frequency value, and frequency width of the frequency band interval where the frequency domain parameter is located; The spectral characteristic value is determined based on the lower limit frequency value, the upper limit frequency value, the frequency width, and the amplitude of the frequency domain parameter.

7. The audio processing method according to any one of claims 1-6, characterized in that, The step of determining the target frequency domain parameter from among multiple frequency domain parameters based on the spectral feature value and the sparse feature value includes: Obtain the rate of change of the square of the amplitude of the frequency domain parameter; Based on the rate of change, the spectral feature value, and the sparse feature value, a target frequency domain parameter is determined from among a plurality of frequency domain parameters, wherein the rate of change corresponding to the target frequency domain parameter is greater than a change threshold.

8. An audio processing apparatus, characterized in that, Including memory and processor, among which, The memory is connected to the processor and is used to store programs; The processor is used to implement the audio processing method as described in any one of claims 1-7 by running the program in the memory.

9. A vehicle, characterized in that, The vehicle includes an audio processing device that implements the audio processing method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the audio processing method as described in any one of claims 1-7.

11. A computer program product, characterized in that, When the computer program is executed by the processor, it implements the audio processing method as described in any one of claims 1-7.