Data processing method and system and computer equipment
By converting the audio signal into the frequency domain and enhancing the process in the frequency domain, the target gain is determined, and the delay problem caused by volume differences in multi-person meetings is solved, achieving a smoother audio output.
Patent Information
- Application Number
- CN202510561141.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-08
AI Technical Summary
In a multi-person conference scenario, due to the large difference in sound volume due to the different distances of participants from the microphone, the existing AI's AGC algorithm causes output delay through overlapping FFT processing, affecting the information transmission efficiency.
The target audio signal is converted into a first audio signal in a frequency domain dimension, and the number of points converted process is consistent with the number of sample points of the target audio signal, avoiding overlapping processing, determining the target gain based on the energy representation of the second audio signal and the energy representation of the target audio signal, and processing the target audio signal based on the target gain.
Reduces processing delay, improves real-time, and the output audio signal is more stable, avoiding the impact of spectrum leakage.
Smart Images

Figure CN120455901A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of signal processing technology, and in particular to a data processing method, system and computer equipment. Background Art
[0002] In multi-person conferences, participants' distances from microphones vary, leading to variations in the volume of sound picked up by the microphones. One solution is to employ an automatic gain control (AGC) method using artificial intelligence (AI). This method controls the volume of input audio signals of varying intensities, keeping the output audio signal within a constant range.
[0003] However, for AGC algorithms involving AI, in order to prevent spectrum leakage, it is necessary to use overlapping methods to implement Fast Fourier Transform (FFT) processing, but this will cause a large delay in the output results, and the played voice will be out of sync with the lip movements of the speakers at the meeting, which in turn affects the efficiency of information transmission. Summary of the Invention
[0004] This application mainly provides a data processing method, system and computer device. The technical solution of this application is implemented as follows:
[0005] In a first aspect, an embodiment of the present application provides a data processing method, comprising:
[0006] obtaining a target audio signal;
[0007] Convert the target audio signal into a first audio signal in the frequency domain; the number of points processed by the conversion is consistent with the number of sampling points of the target audio signal;
[0008] performing enhancement processing on the first audio signal to obtain a second audio signal;
[0009] determining a target gain based on an energy representation of the second audio signal and an energy representation of the target audio signal;
[0010] The target audio signal is processed based on the target gain to obtain a third audio signal.
[0011] In a second aspect, an embodiment of the present application provides a data processing system, including:
[0012] The first module is used to obtain a target audio signal;
[0013] The second module is configured to convert the target audio signal into a first audio signal in the frequency domain; the number of points processed by the conversion is consistent with the number of sampling points of the target audio signal;
[0014] A third module is configured to perform enhancement processing on the first audio signal to obtain a second audio signal;
[0015] The fourth module is configured to determine a target gain based on the energy representation of the second audio signal and the energy representation of the target audio signal, and process the target audio signal based on the target gain to obtain a third audio signal.
[0016] In a third aspect, an embodiment of the present application provides a computer device, comprising at least one audio acquisition component, at least one audio output component, at least one memory, and at least one processor, wherein:
[0017] At least one audio acquisition component, configured to acquire an audio signal;
[0018] a memory for storing a plurality of computer instructions;
[0019] The processor is configured to load and execute computer instructions to implement the following methods:
[0020] Obtaining a target audio signal; converting the target audio signal into a first audio signal in a frequency domain dimension; the number of points processed by the conversion is consistent with the number of sampling points of the target audio signal; performing enhancement processing on the first audio signal to obtain a second audio signal; determining a target gain based on an energy representation of the second audio signal and an energy representation of the target audio signal; processing the target audio signal based on the target gain to obtain a third audio signal;
[0021] The audio output component is configured to output based on the third audio signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flow chart of a speech processing method provided in this application embodiment Figure 1 ;
[0023] Figure 2 A flow chart of a speech processing method provided in this application embodiment Figure 2 ;
[0024] Figure 3 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 1 ;
[0025] Figure 4 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 2 ;
[0026] Figure 5 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 3 ;
[0027] Figure 6A schematic diagram of a data processing method provided in an embodiment of the present application Figure 4 ;
[0028] Figure 7 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 5 ;
[0029] Figure 8 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 6 ;
[0030] Figure 9 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 7 ;
[0031] Figure 10 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 8 ;
[0032] Figure 11 A schematic diagram of the structure of a data processing system provided in an embodiment of the present application;
[0033] Figure 12 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present application;
[0034] Figure 13 A schematic diagram of the composition structure of a second model provided in an embodiment of the present application;
[0035] Figure 14 A schematic diagram of the composition structure of a first model provided in an embodiment of the present application;
[0036] Figure 15 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 9 . DETAILED DESCRIPTION
[0037] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.
[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0039] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0040] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0041] In addition, references to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of such phrases in various places in the specification does not necessarily refer to the same embodiment, nor does it necessarily refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0042] The following is an introduction to the relevant technologies of this application.
[0043] In multi-person conferences or when switching between different scenarios, audio capture can vary due to factors like equipment, environment, and the speakers themselves, leading to inconsistent volume. Specifically, different speakers have varying volume levels, resulting in persistently low or suddenly high volumes. The distance between the speaker and the microphone can also cause the sound to fluctuate. This not only affects the listener's listening experience but can also force them to frequently adjust the device volume, impacting the user's listening experience.
[0044] One solution to this scenario is to use an AGC algorithm to automatically adjust the input audio signal, maintaining the output volume within a constant range regardless of changes in the input audio signal's strength. The AGC algorithm monitors the input audio signal's amplitude in real time and dynamically adjusts the gain factor based on a preset target level, preventing the output audio from being too low or too high. AGC is widely used in areas such as teleconferencing, voice recognition, and music playback, significantly improving audio signal quality and acceptability.
[0045] The following example illustrates an AGC usage scenario. For example, two participants at Location A are in the same conference room, holding a remote video conference with colleagues at Location B. The conference room at Location A contains only one microphone, with one participant sitting one meter from the microphone and the other five meters away. As the two participants at Location A discuss their issue, the voices heard by the colleague at Location B sound extremely loud, while the other sounds extremely quiet. If the AGC algorithm intervenes, it can attenuate the voice signal of the participant at one meter away and amplify the voice signal of the participant at five meters away, ensuring a relatively stable volume for the colleague at Location B.
[0046] At present, one of the algorithms actually used in the above scenario is the AGC algorithm involving AI (also known as the AI voice algorithm). This algorithm extracts features from the input time-domain voice signal after FFT, trains it, and then inputs it into the neural network model to output the audio signal after gain. In order to ensure the smoothness of the processing features of the previous and next frames and prevent spectrum leakage, an overlap must be set between frames when performing FFT. For example, the sampling rate of the input pulse code modulation (PCM) audio signal, also known as the original audio signal, is 16000Hz. When performing audio processing, the FFT length is generally set to twice the frame shift. It can be understood that if the FFT is set to 512 points, the frame shift is 256 points, which overlaps by 50%.
[0047] like Figure 1 As shown, for a 16000Hz voice signal, when the length of a frame of voice signal is 256 points, the length of FFT is set to 512 points, and the overlap length (hoplength) is set to 256 points, after the voice signal 1 is input into the FFT processing module 101, it is necessary to wait for the next frame of audio signal, that is, voice signal 2, and merge these two frames of voice signals to accumulate 512 points. The signal is input into the FFT processing module 101 for FFT processing, and further input into the automatic gain control module 102. The gain is determined based on the AGC algorithm, and the gain acts on the first frame of audio signal that enters the FFT, that is, voice signal 1. In this way, since it is necessary to wait for voice signal 2 before FFT processing can be performed, the AI voice algorithm generates a delay of 16ms.
[0048] like Figure 2As shown, speech processing is performed frame by frame, one frame at a time. One frame of data consists of 256 sampling points, and the time for 256 sampling points is 16ms. 512 sampling points are collected to perform an FFT, of which there is a 50% overlap. It can be understood as follows: the first frame input 1031 is 256 point data, which is less than 512. At this time, 256 zeros will be added to form 512 data, and the first transformation (FFT processing) 1031 will be performed. After the FFT processing is completed, the algorithm will output 256 zeros 1051 or data close to 0; during the second transformation 1042, the 256 point data of the second frame input 1032 and the 256 point data of the first frame input 1031 are combined into 512 point data. After the FFT processing is completed, the data result of 256 points of the first frame output 1052 will be obtained, which corresponds to the processing result of the first frame input 1031. In this way, a delay of 256 samples (16ms) will be generated. Correspondingly, during the third transformation 1043, FFT processing is performed on the second frame input 1032 and the third frame input 1033 to obtain the second frame output 1053; during the fourth transformation 1044, FFT processing is performed on the third frame input 1033 and the fourth frame input 1043 to obtain the third frame output 1054. Subsequent processing can refer to the above process and be carried out in sequence.
[0049] In summary, in the above process, the input voice data is processed in an overlapping manner, which weakens the mutual influence between the various spectra in the signal spectrum and causes the problem of spectrum leakage. However, this FFT processing method increases the delay of the processing result output, affecting the efficiency of information transmission.
[0050] Based on this, an embodiment of the present application provides a data processing method, system and computer device. First, the target audio signal is converted into a first audio signal in the frequency domain dimension, wherein the number of points of the conversion processing is consistent with the number of sampling points of the target audio signal. In this way, the delay caused by overlap is avoided, and the output target gain can be directly applied to the input target audio signal, thereby improving real-time performance; secondly, the first audio signal is enhanced to obtain a second audio signal, and the target gain is further determined based on the energy representation of the second audio signal and the energy representation of the target audio signal, and the target audio signal is processed based on the target gain to obtain the target gain. In this way, not only the output target gain is made more accurate, but also the influence of spectrum leakage on the output third audio signal is avoided, and the output third audio signal is made smoother.
[0051] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0052] In one embodiment of the present application, Figure 3 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 1 .like Figure 3As shown, the method may include:
[0053] S201: Obtain a target audio signal.
[0054] In the embodiment of the present application, the target audio signal may refer to an audio signal to be subjected to processing such as gain processing based on the following steps.
[0055] Among them, the source of the target audio signal can be a voice signal captured in real time by a microphone, or a voice signal recorded and played in advance, or a voice signal obtained by pre-processing the voice signal, or a voice signal obtained by PCM processing the input voice signal, which are not listed here one by one.
[0056] S202: Convert the target audio signal into a first audio signal in the frequency domain.
[0057] The number of points of the conversion processing is consistent with the number of sampling points of the target audio signal.
[0058] In an embodiment of the present application, after obtaining the target audio signal, the time information of the target audio signal can be removed and converted to the frequency domain to obtain the first audio signal. For example, the target audio signal can be converted to the first audio signal in the frequency domain by using methods such as FFT, discrete cosine transform, and scale-invariant feature transform (SIFT).
[0059] It should be noted that, taking the conversion of the target audio signal into the first audio signal in the frequency domain as an example, during the conversion process, the number of points for each conversion process can be set to be consistent with the number of sampling points of the target audio signal. Figure 4 As shown, when the number of sampling points of the target audio signal is 256, the number of points for the conversion processing is also set to 256. Thus, for target audio signal 1 3011, when the first conversion processing 3021 is performed, a 256-point conversion processing is performed on target audio signal 1 3011, and the number of sampling points of the output first audio signal 1 3031 is also 256. For subsequently input target audio signals 2, 3, 4, etc., each target audio signal is independently converted to obtain the corresponding first audio signal 2, first audio signal 3, first audio signal 4, etc. In this way, the number of points for the conversion processing is equal to the number of sampling points of the input target audio signal. There is no audio signal caching and no overlapping processing is required, which avoids waiting for the audio signal of the next frame, thereby reducing the delay of the conversion processing.
[0060] S203: Perform enhancement processing on the first audio signal to obtain a second audio signal.
[0061] It should be noted that the enhancement processing may include processing to improve the clarity of the speech signal in the first audio signal, processing to improve the quality of the first audio signal, etc. The model or algorithm for implementing the enhancement processing is not specifically limited herein. Exemplarily, the enhancement processing may include noise reduction, speech signal enhancement, spectrum adjustment and control, etc.
[0062] like Figure 4 As shown, after the first audio signal 3031 is enhanced, a second audio signal 3041 is obtained. The number of sampling points of the second audio signal is the same as the number of sampling points of the first audio signal. For the first audio signal 2, the first audio signal 3, the first audio signal 4, and subsequent first audio signals, each first audio signal can be enhanced to obtain the corresponding second audio signal 2, the second audio signal 3, the second audio signal 4, and subsequent second audio signals.
[0063] In the embodiment of the present application, the first audio signal is a frequency domain signal that has been converted and processed, and its information is more concentrated in the frequency domain dimension. Therefore, in the process of enhancing the first audio signal, it is easier to discover and enhance the features in the signal, while reducing information loss and improving the quality of the audio signal.
[0064] S204: Determine a target gain based on the energy representation of the second audio signal and the energy representation of the target audio signal.
[0065] It should be noted that during the conversion process of the target audio signal described above, for each target audio signal, the number of sampling points of the first audio signal output by the conversion process is the same as the number of points of the target audio signal, and the conversion process does not need to be overlapped. Although this conversion process is latency-free, if the gain is determined based on the first audio signal through the AGC algorithm, if the conversion process is not overlapped, it may cause speech smoothing and spectral leakage.
[0066] Based on this, in embodiments of the present application, the difference between the energy representation of the output second audio signal and the energy representation of the target audio signal can be used as the gain change of the current target audio signal, also referred to as the target gain. The target gain determined based on this step for the audio signal obtained during the aforementioned conversion process with overlap processing is very close to the target gain determined based on this step for the second audio signal obtained without overlap processing. In other words, when the target gain is determined using the energy representation of the audio signal, the presence or absence of spectral leakage has little impact on the target gain.
[0067] The energy representation of the second audio signal may refer to an actual physical quantity carried in the second audio signal, such as the power or intensity of the sound, and may be associated with the amplitude of the sound. For example, the energy representation of the second audio signal may be a root mean square (RMS) value of the second audio signal, which may be obtained by sequentially performing square, sum, average, and square root operations on each sampling point in the second audio signal.
[0068] In the embodiment of the present application, the meaning and determination method of the energy representation of the target audio signal may refer to the energy representation of the second audio signal.
[0069] For example, when the energy is expressed as an RMS value, the result of subtracting the RMS value of the target audio signal from the RMS value of the second audio signal can be determined as the target gain corresponding to the target audio signal. Figure 4 As shown, a corresponding first target gain can be determined based on the energy representation of the output second audio signal 1 3041 and the energy representation of the corresponding target audio signal 1 3011. Accordingly, a second target gain is determined based on the second audio signal 2 and the target audio signal 2, a third target gain is determined based on the second audio signal 3 and the target audio signal 3, and a fourth target gain is determined based on the second audio signal 4 and the target audio signal 4. Subsequent determination of target gains refers to the aforementioned process.
[0070] S205: Process the target audio signal based on the target gain to obtain a third audio signal.
[0071] In the embodiment of the present application, the target gain can be understood as the amplitude of the change of the target audio signal, including the amplitude of the increased sound, the amplitude of the reduced sound, etc.
[0072] In the embodiment of the present application, the third audio signal may be a processed audio signal output after processing the target audio signal based on the target gain, wherein the processing may be superimposing the target gain on the target audio signal.
[0073] It should be noted that the output target gain has a corresponding relationship with the input target audio signal, that is, the output target gain acts on the input target audio signal. Figure 4As shown, a conversion process is performed based on the target audio signal 13011 to obtain a first audio signal 13031. The first audio signal 13031 is then enhanced to obtain a second audio signal 13041. A first target gain 3051 is determined based on the energy representation of the second audio signal 13041. Finally, the output first target gain 3051 is applied to the input target audio signal 3011 to obtain an output third audio signal 13061. During this process, the number of sampling points of the target audio signal 13011, the first audio signal 13031, the second audio signal 13041, and the third audio signal 13061 are consistent. In subsequent processing, each target audio signal corresponds to the output target gain. Thus, after the above steps are completed, the third audio signal can be obtained and played without waiting for the next audio signal frame, thereby reducing latency.
[0074] Based on the above steps S201 to S205, the following description is given: Figure 5 As shown, the target audio signal is converted into a first audio signal in the frequency domain dimension through frequency domain conversion 307, the first audio signal is enhanced 308 to obtain a second audio signal, a target gain is determined based on the second audio signal and the target audio signal, and the target audio signal is processed with the determined target gain to obtain a third audio signal.
[0075] An embodiment of the present application provides a data processing method. First, a target audio signal is converted into a first audio signal in the frequency domain dimension, wherein the number of points of the conversion processing is consistent with the number of sampling points of the target audio signal. In this way, the delay caused by overlap is avoided, and the output target gain can be directly applied to the input target audio signal, thereby improving real-time performance. Secondly, the first audio signal is enhanced to obtain a second audio signal, and the target gain is further determined based on the energy representation of the second audio signal and the energy representation of the target audio signal, and the target audio signal is processed based on the target gain to obtain the target gain. In this way, not only the output target gain is made more accurate, but also the influence of spectrum leakage on the output third audio signal is avoided, and the output third audio signal is made smoother.
[0076] In another embodiment of the present application, Figure 6 As shown, for the aforementioned step S202, converting the target audio signal into a first audio signal in the frequency domain dimension may include:
[0077] S401: Perform Fourier transform on the target audio signal to obtain corresponding frequency domain features.
[0078] In an embodiment of the present application, the target audio signal can be processed by Fourier transform to decompose the target audio signal into a superposition of sine waves and cosine waves of different frequencies, thereby converting the target audio signal into corresponding frequency domain features.
[0079] Among them, the frequency domain features can more clearly display the frequency components of the target audio signal and the energy distribution at different frequencies, which facilitates the identification of noise, interference components, etc. in the audio signal in the subsequent enhancement processing steps.
[0080] S402: Divide the frequency domain features into multiple sub-bands to obtain a first audio signal containing the multiple sub-bands.
[0081] Among them, the spectrum features are divided into multiple sub-bands. For example, the Equivalent Rectangular Bandwidth (ERB) method can be used to simulate the human perception of signal frequency characteristics and divide the frequency domain features into multiple sub-bands with lower dimensions. This can reduce the number of features and facilitate subsequent enhancement processing.
[0082] In an embodiment of the present application, the target audio signal can be converted into frequency domain features and the frequency domain features can be used as the first audio signal; alternatively, the frequency domain features can be further divided into multiple sub-bands based on the conversion of the target audio signal into the frequency domain to obtain a first audio signal containing multiple sub-bands.
[0083] An embodiment of the present application provides a data processing method that converts a target audio signal into frequency domain features based on Fourier transform, obtains the energy distribution of the target audio signal at different frequencies, and facilitates the rapid identification of noise, interference and other signals; further divides the frequency domain features into sub-bands to obtain a first audio signal containing multiple sub-bands, thereby reducing the number of features and improving processing efficiency.
[0084] In another embodiment of the present application, Figure 7 As shown, for the aforementioned step S203, performing enhancement processing on the first audio signal to obtain the second audio signal may include:
[0085] S501: Perform noise reduction processing on a first audio signal to obtain a first intermediate audio signal.
[0086] In an embodiment of the present application, the first audio signal may include a speech audio signal of a noisy audio signal, and performing noise reduction processing on the first audio signal may include identifying the noise audio signal in the first audio signal, suppressing the noise audio signal in the first audio signal, and enhancing the speech audio signal in the first audio signal to obtain a first intermediate audio signal.
[0087] The noise reduction process may be implemented by an automatic noise cancellation algorithm (ANC), a time domain or frequency filtering algorithm, a pre-trained neural network model, etc., which is not specifically limited here.
[0088] It should be noted that after the target audio signal is converted into a first audio signal in the frequency domain dimension in the aforementioned step, the first audio signal is subjected to noise reduction processing in the frequency domain. This processing method can better utilize and capture frequency information, making specific frequency components, such as noise audio signals, easier to capture and suppress, thereby enhancing the effect of noise reduction processing.
[0089] It should also be noted that, in the embodiment of the present application, when the input first audio signal is a first audio signal containing multiple sub-bands, the output audio signal containing multiple sub-bands can be converted into corresponding frequency domain features to obtain a first intermediate audio signal.
[0090] S502: Convert the first intermediate audio signal into a second intermediate audio signal in the time domain.
[0091] In an embodiment of the present application, the first audio signal can be converted into a second intermediate audio signal in the time domain dimension to solve problems that are difficult to solve from a frequency domain perspective, such as the smoothness, trend, periodicity and other characteristics of the audio signal.
[0092] Among them, the conversion method can be the inverse process of converting the target audio signal into the first audio signal of the frequency domain dimension, for example, it can be Inverse Fast Fourier Transform (IFFT), Inverse Discrete Fourier Transform (IDFT), Laplace transform, etc., and the audio signal is restored from the frequency domain to the time domain by inverse transforming the signal to obtain a second intermediate audio signal.
[0093] S503: Perform gain fitting processing on the second intermediate audio signal to obtain a second audio signal.
[0094] In an embodiment of the present application, a second intermediate audio signal in the time domain is subjected to gain fitting processing in the time domain to obtain a second audio signal. Gain fitting processing may refer to strengthening and fitting the gain of the intermediate process so that the target gain meets the requirements. In other words, the third audio signal obtained by processing the target audio signal with the target gain is closer to the desired effect.
[0095] The gain processing may be implemented based on a pre-trained model or algorithm, which is not specifically limited here.
[0096] An embodiment of the present application provides a data processing method that first performs noise reduction processing on a first audio signal in the frequency domain to obtain a first intermediate audio signal. The first intermediate audio signal is then converted into a second intermediate audio signal in the time domain, and gain fitting is performed on the intermediate audio signal to obtain a second audio signal. In this manner, performing noise reduction processing in the frequency domain makes it easier to capture and suppress interfering signals, while performing gain fitting in the time domain makes it easier to fit the gain to a desired value, thereby improving both processing efficiency and effectiveness.
[0097] In some embodiments, as Figure 8 As shown, for the aforementioned step S503, performing gain fitting processing on the second intermediate audio signal to obtain the second audio signal may include:
[0098] S601: Extract features of the second intermediate audio signal and perform normalization processing to obtain first feature information.
[0099] In the embodiment of the present application, feature extraction is first performed on the second audio signal to obtain features of the second audio signal. Further, normalization is performed on the features of the second audio signal to obtain first feature information.
[0100] Among them, the normalization process is used to accelerate the training process, make the distribution of the data of the first feature information more stable, and reduce the problem of internal covariate shift.
[0101] S602: Perform sparse activation and time series analysis on the first feature information to obtain a first mask sequence.
[0102] The length of the first mask sequence is consistent with the length of the target audio signal.
[0103] In an embodiment of the present application, the first feature information obtained by normalization is further sparsely activated, for example, by means of a rectified linear unit (Parametric Rectified Linear Unit, PReLU). Through this unilateral suppression method, the features in the first feature information can be better mined.
[0104] Furthermore, the first feature information after sparse activation is subjected to time series analysis processing, including processing the time series data in the first feature information, capturing the dependency between the time series data in the first feature information after sparse activation, capturing the time features in the audio signal, and outputting the first mask sequence based on the noise reduction processing of the aforementioned first model.
[0105] The first mask sequence may be a value sequence of the same length as the input target audio signal, each value of the first mask sequence corresponds to a sampling point of the target audio signal, and each value is between 0 and 1, indicating noise reduction information and gain information for the corresponding sampling point of the target audio signal.
[0106] S603: Process the target audio signal based on the first mask sequence, suppress the noise signal in the target audio signal, and gain the speech signal in the target audio signal to obtain a second audio signal.
[0107] In an embodiment of the present application, the value of the first mask sequence is multiplied by the sampling point at the corresponding position in the target audio signal. Based on the different values of the first mask sequence, the noise reduction result of the first model is applied to the target audio signal to suppress the noise signal at the corresponding sampling point. In addition, based on the value of the first mask sequence, the gain fitting result of the second model is applied to the target audio signal to obtain a second audio signal.
[0108] This embodiment of the present application provides a data processing method that sequentially performs feature extraction, normalization, sparse activation, and time series analysis on a second intermediate audio signal to obtain a first mask sequence. The target audio signal is then processed based on the first mask sequence to obtain a second audio signal. In this manner, the first mask sequence retains the aforementioned noise reduction processing results while also taking into account gain information, resulting in an enhanced and smoother speech signal in the third audio signal.
[0109] In some embodiments, as Figure 9 As shown, before the aforementioned step S603, in which the target audio signal is processed based on the first mask sequence, the following steps may also be included:
[0110] S701: Acquire a target level value, where the target level value is used to represent an expected level value.
[0111] In the embodiment of the present application, the target level value represents the user's expected level value for the third audio signal, and is associated with the output volume of the third audio signal.
[0112] The target level value may be input by a user, or generated by a computer device, or obtained by a user based on a recommendation of a computer device.
[0113] S702: Update the first mask sequence based on the target level value to obtain an updated first mask sequence.
[0114] The updated first mask sequence enables the level value of the second audio signal to match the target level value.
[0115] In an embodiment of the present application, after obtaining the first mask sequence, the target level value (also referred to as TargetLevel) can be fused with the first mask sequence. The fusion method can be to perform matrix multiplication (or dot multiplication) on the target level value and the first mask sequence, introduce the influence of the target level value parameter on the third audio signal, and obtain an updated first mask sequence.
[0116] The updated first mask sequence takes into account factors such as noise reduction, gain fitting, and a desired target level value. The updated first mask sequence is fused with the target audio signal, such as by performing a dot product, to obtain a second audio signal.
[0117] This embodiment of the present application provides a data processing method that updates a first mask sequence based on an acquired target level value to obtain an updated first mask sequence. In this way, the target level value is also considered in the process of determining the target gain, so that the third audio signal approaches the target level value, thereby providing a more flexible and effective desired level range for the user.
[0118] In another embodiment of the present application, Figure 10 As shown, for the aforementioned step S203, performing enhancement processing on the first audio signal to obtain the second audio signal may include:
[0119] S801: Downsample and transpose the first audio signal to obtain second feature information.
[0120] In the embodiment of the present application, encoding the first audio signal includes mapping the first audio signal to a high-dimensional embedding and performing downsampling processing to obtain a downsampled first audio signal.
[0121] Furthermore, the down-sampled first audio signal is classified and judged to distinguish noise and speech signals, and is subjected to transposition recovery processing to obtain second feature information.
[0122] It should be noted that, when the first audio signal is a first audio signal including multiple sub-bands, the first audio signal has fewer features, thereby improving processing efficiency.
[0123] S802: Perform feature extraction based on the second feature information to determine a second mask sequence.
[0124] In the embodiment of the present application, feature extraction processing is performed on the second feature information. Feature extraction may be the inverse process of the aforementioned downsampling processing. The second mask sequence corresponding to the second feature information is obtained through feature extraction.
[0125] It should be noted that the second mask sequence can be a value sequence having the same length as the first audio signal and the target audio signal, wherein the value of each bit of the second mask sequence corresponds to the sampling point at the corresponding position of the first audio signal. Exemplarily, the value of each bit can be 0 or 1, where 0 indicates that the corresponding sampling point is a noise signal, and 1 indicates that the corresponding sampling point is a speech signal.
[0126] S803: Process the first audio signal based on the second mask sequence, suppress the noise signal in the first audio signal, and obtain a second audio signal.
[0127] In an embodiment of the present application, the second mask sequence processes the first audio signal, which can be understood as multiplying the second mask sequence by the sampling points at corresponding positions in the first audio signal. Each value in the second mask sequence represents relevant information for noise reduction processing. Therefore, by processing the first audio signal with the second mask sequence, the noise signal in the first audio signal can be suppressed, and the speech signal in the first audio signal can be enhanced to obtain a second audio signal.
[0128] It should be noted that after the second mask sequence is used to process the first audio signal, the first audio signal with sub-bands may be converted into a spectrum based on the reverse process of dividing the frequency domain features into sub-bands to obtain the second audio signal.
[0129] It should also be noted that the process of performing noise reduction processing on the first audio signal to obtain the first intermediate audio signal in step S501 can be handled in accordance with this embodiment. This process may include: downsampling and transposing the first audio signal to obtain second feature information; performing feature extraction based on the second feature information to determine a second mask sequence; and processing the first audio signal based on the second mask sequence to suppress noise signals in the first audio signal to obtain the first intermediate audio signal.
[0130] The present embodiment provides a data processing method that sequentially downsamples, transposes, and extracts features from a first audio signal to obtain a second mask sequence. The first audio signal is then processed based on the second mask sequence to obtain a second audio signal. This method enhances the speech signal in the first audio signal and suppresses the noise signal in the first audio signal, preventing the target gain from being applied to the noise signal, thereby improving the quality of the third audio signal.
[0131] In another embodiment of the present application, in step S201, obtaining the target audio signal includes any of the following:
[0132] Obtaining an input audio signal, and in response to the input audio signal including multiple sound source audios, separating the input audio signal to obtain input sub-audio signals corresponding to the respective sound source audios, and determining the input sub-audio signals as target audio signals;
[0133] An input audio signal is obtained, and in response to the input audio signal including multiple sound source audios, the input audio signal is determined as a target audio signal.
[0134] In an embodiment of the present application, the input audio signal may include multiple sound source audios. For example, in a conference scenario, the same microphone collects multiple sound source audios at the same time. These sound source audios correspond to participants at different distances from the microphone, and the volumes of the multiple sound source audios may be different.
[0135] In this case, in some embodiments, voiceprint recognition can be performed on the input audio signal to separate the input sub-audio signals corresponding to different users (sound sources), and each input sub-audio signal is used as the target audio signal and processed based on the steps in the aforementioned embodiments to obtain a third audio signal corresponding to each input sub-audio signal.
[0136] Alternatively, in this case, in some embodiments, voiceprint recognition can be performed on the input audio signal, the sound source audio with the highest volume in the input audio signal is processed as the target audio signal, and the corresponding third audio signal is output, and other sound source audios are eliminated as background sounds.
[0137] Alternatively, in this case, in some embodiments, the higher volume of one of the sound source audios and the larger difference between the volume of the other sound source audios and the volume of the sound source audio are used as optional judgment conditions, and the input audio signal can be used as the target audio signal.
[0138] Alternatively, the input audio signal may include a source audio, in which case the input audio signal may be used as the target audio signal.
[0139] In the embodiment of the present application, when the input audio signal is separated into multiple target audio signals, after obtaining the third audio signal corresponding to each target audio signal, the third audio signals can be merged to obtain a fourth audio signal. It should be noted that because the embodiment of the present application adopts a streaming processing method, the third audio signals corresponding to the target audio signal of the same frame can be merged to obtain a fourth audio signal.
[0140] The present application provides a data processing method in which the target audio signal can be an input audio signal, or can be an input sub-audio signal corresponding to one of the source audio signals within the input audio signal. This allows users to adaptively select a processing method for the input audio signal based on their actual needs, thereby improving the user experience and the quality of the third audio signal.
[0141] In another embodiment of the present application, Figure 11 This is a structural diagram of a data processing system provided in an embodiment of the present application. Figure 11 As shown, the data processing system includes:
[0142] The first module 1101 is configured to obtain a target audio signal.
[0143] The second module 1102 is configured to convert the target audio signal into a first audio signal in the frequency domain; the number of points processed by the conversion is consistent with the number of sampling points of the target audio signal.
[0144] The third module 1103 is configured to perform enhancement processing on the first audio signal to obtain a second audio signal.
[0145] The fourth module 1104 is configured to determine a target gain based on the energy representation of the second audio signal and the energy representation of the target audio signal, and process the target audio signal based on the target gain to obtain a third audio signal.
[0146] In some embodiments, the second module 1102 is further configured to perform Fourier transform on the target audio signal to obtain corresponding frequency domain features; and divide the frequency domain features into multiple sub-bands to obtain a first audio signal containing multiple sub-bands.
[0147] In an embodiment of the present application, the target module 1102 may include a Fourier transform unit a1 and an equivalent rectangular bandwidth unit a2. The target audio signal is input into the Fourier transform unit a1, and the target audio signal is Fourier transformed to obtain frequency domain features corresponding to the target audio signal; further, the frequency domain features of the target audio signal are input into the equivalent rectangular bandwidth unit a2, and the frequency domain features are divided into multiple sub-bands to obtain a first audio signal containing multiple sub-bands.
[0148] In some embodiments, as Figure 11 As shown, the third module 1103 includes: a first model b1 and a second model b2.
[0149] The first model b1 is configured to perform noise reduction processing on the first audio signal to obtain a first intermediate audio signal, and convert the first intermediate audio signal into a second intermediate audio signal in the time domain;
[0150] Among them, the first model 503 can be a model that can realize noise reduction processing, has the ability to capture noise and extract speech signals, such as a U-NET model, a convolutional neural network (CNN) model, etc.
[0151] In this embodiment of the present application, a first audio signal is input into the first model b1, which performs noise reduction processing to obtain a first intermediate audio signal. Furthermore, the first model b1 may also include a unit for converting the audio signal in the frequency domain into an audio signal in the time domain, such as an IFFT unit, which converts the first intermediate audio signal into a second intermediate audio signal in the time domain.
[0152] The second model b2 is used to perform gain fitting processing on the second intermediate audio signal to obtain a second audio signal.
[0153] The second model 504 may be a model of a neural network structure including convolution and recurrence, such as a convolutional recurrent neural network (CRN).
[0154] In the embodiment of the present application, the second intermediate audio signal is input into the second model b2 for gain fitting processing to obtain the second audio signal. The implementation process of the gain fitting processing can refer to the corresponding method embodiment mentioned above.
[0155] In some embodiments, the second model b2 is also used to extract features of the second intermediate audio signal and perform normalization processing to obtain first feature information; perform sparse activation and timing analysis processing on the first feature information to obtain a first mask sequence; the length of the first mask sequence is consistent with the length of the target audio signal; the target audio signal is processed based on the first mask sequence, the noise signal in the target audio signal is suppressed, and the speech signal in the target audio signal is amplified to obtain a second audio signal.
[0156] In some embodiments, the second model b2 is further used to obtain a target level value, which is used to represent an expected level value; the first mask sequence is updated based on the target level value to obtain an updated first mask sequence; the updated first mask sequence makes the level value of the second audio signal match the target level value.
[0157] In some embodiments, the third module 1103 includes a first model b1; the first model b1 is used to perform noise reduction processing on the first audio signal to obtain a first intermediate audio signal, and convert the first intermediate audio signal into a second intermediate audio signal in the time domain dimension.
[0158] In some embodiments, the first model b1 is further used to downsample and transpose the first audio signal to obtain second feature information; perform feature extraction based on the second feature information to determine a second mask sequence; process the first audio signal based on the second mask sequence to suppress the noise signal in the first audio signal to obtain a second audio signal.
[0159] In some embodiments, the first module 1101 is further used to obtain an input audio signal, in response to the input audio signal including multiple sound source audios, separate the input audio signal to obtain input sub-audio signals corresponding to each sound source audio, and determine the input sub-audio signal as a target audio signal; obtain an input audio signal, and in response to the input audio signal including multiple sound source audios, determine the input audio signal as a target audio signal.
[0160] In some embodiments, Figure 12 This is a hardware structure diagram of a computer device provided in an embodiment of the present application. Figure 12 As shown, the computer device 120 may include: at least one audio acquisition component 1201, at least one audio output component 1202, at least one memory 1203 and at least one processor 1204.
[0161] At least one audio acquisition component 1201, configured to acquire audio signals;
[0162] a memory for storing a plurality of computer instructions;
[0163] The processor is configured to load and execute computer instructions to implement the following methods:
[0164] Obtaining a target audio signal; converting the target audio signal into a first audio signal in a frequency domain dimension; the number of points processed by the conversion is consistent with the number of sampling points of the target audio signal; performing enhancement processing on the first audio signal to obtain a second audio signal; determining a target gain based on an energy representation of the second audio signal and an energy representation of the target audio signal; processing the target audio signal based on the target gain to obtain a third audio signal;
[0165] The audio output component is configured to output based on the third audio signal.
[0166] In some embodiments, the processor may further be configured to: perform Fourier transform on the target audio signal to obtain corresponding frequency domain features; and divide the frequency domain features into multiple sub-bands to obtain a first audio signal containing the multiple sub-bands.
[0167] In some embodiments, the processor can also be used to execute: performing noise reduction processing on the first audio signal to obtain a first intermediate audio signal; converting the first intermediate audio signal into a second intermediate audio signal in the time domain dimension; and performing gain fitting processing on the second intermediate audio signal to obtain a second audio signal.
[0168] In some embodiments, the processor can also be used to execute: extracting features of the second intermediate audio signal and performing normalization processing to obtain first feature information; performing sparse activation and timing analysis processing on the first feature information to obtain a first mask sequence; the length of the first mask sequence is consistent with the length of the target audio signal; processing the target audio signal based on the first mask sequence, suppressing the noise signal in the target audio signal, and gaining the speech signal in the target audio signal to obtain a second audio signal.
[0169] Figure 13 A schematic diagram of the composition structure of a second model provided in an embodiment of the present application. In an embodiment of the present application, the first model and the second model can be deployed on the processor. Figure 13 As shown, the second model is used to input the second intermediate audio signal into a one-dimensional convolutional layer 1 (Conv1d) 1301, which is used to extract features of the second audio signal. Convolutional layer 1 1301 inputs the extracted features of the second audio signal into a normalization module (BatchNormal) 1302 to obtain first feature information.
[0170] Furthermore, the first feature information output by the normalization module 1302 is input into the activation function 1 1303, which enables the neuron to have sparse activation and is used to perform sparse activation on the first feature information, wherein the activation function 1 can be, for example, a rectified linear unit (PReLU).
[0171] Furthermore, the first feature information after sparse activation of the activation function 1 1303 is input into the recurrent neural network 1 (Gated Recurrent Unit, GRU) 1304. The structure of the recurrent neural network 1 1304 may include two GRUs, which can be used to process time series data. By capturing the dependency between the time series data in the first feature information after sparse activation, the time features in the audio signal are captured, and based on the noise reduction processing of the aforementioned first model, the first mask sequence is output.
[0172] In the embodiment of the present application, the second audio signal may be obtained by performing a point product between the obtained first mask sequence and the target audio signal.
[0173] In some embodiments, the processor may further be configured to execute: obtaining a target level value, where the target level value is used to represent an expected level value; updating the first mask sequence based on the target level value to obtain an updated first mask sequence; and enabling the updated first mask sequence to cause the level value of the second audio signal to match the target level value.
[0174] like Figure 13 As shown, after obtaining the first mask sequence, the target level value (also referred to as TargetLevel) can be fused with the first mask sequence. The fusion method can be to perform matrix multiplication (or dot product) on the target level value and the first mask sequence, introduce the influence of the target level value parameter on the third audio signal, and obtain an updated first mask sequence.
[0175] Furthermore, the updated first mask sequence can be sequentially passed through the one-dimensional convolution layer 2 (Conv1d) 1305 for feature extraction, and the updated first mask sequence can be input into the activation function 2 1306 to output the target mask sequence, which is then fused with the target audio signal, such as by dot product, to obtain the second audio signal.
[0176] Among them, activation function 2 1306 can be a hyperbolic tangent activation (tanh) function.
[0177] In some embodiments, the processor can also be used to perform: downsampling and transposition recovery processing of the first audio signal to obtain second feature information; performing feature extraction based on the second feature information to determine a second mask sequence; processing the first audio signal based on the second mask sequence to suppress the noise signal in the first audio signal to obtain a second audio signal.
[0178] It should be noted that the first model can be deployed on the processor, or the first model and the second model can be deployed on the processor as in the aforementioned embodiment. The first model can be used to perform noise reduction processing on the first audio signal, and when the first model can be deployed on the processor, a second audio signal can be obtained, or when the first model and the second model are deployed, a second intermediate audio signal can be obtained and input into the second model for gain fitting processing.
[0179] Figure 14 This is a schematic diagram of the composition structure of a first model provided in an embodiment of the present application. Figure 14 As shown, the encoding module 1401 may include convolution layer (Conv) 1, convolution layer 2 and convolution layer 3.
[0180] In an embodiment of the present application, the downsampled first audio signal is input into the recurrent neural network 2 1402. The structure of the recurrent neural network 2 1402 may, for example, include two GRUs for classification and judgment, distinguishing noise and speech signals, and performing transposition recovery processing to obtain second feature information.
[0181] In this embodiment of the present application, the recurrent neural network 2 inputs the second feature information into the decoding module 1403 to obtain a second mask sequence. The decoding module 1403 may be a mirrored version of the encoding module 1401, restoring the second feature information to its original size. The decoding module 1403 may include a deconvolution layer (DeConv) 1, a deconvolution layer 2, and a deconvolution layer 3.
[0182] Furthermore, a dot product process is performed on the second mask sequence and the first audio signal to suppress noise in the first audio signal and obtain a second audio signal.
[0183] In some embodiments, at least one audio acquisition component, configured to acquire an audio signal, includes any of the following:
[0184] Obtaining an input audio signal, and in response to the input audio signal including multiple sound source audios, separating the input audio signal to obtain input sub-audio signals corresponding to the respective sound source audios, and determining the input sub-audio signals as target audio signals;
[0185] An input audio signal is obtained, and in response to the input audio signal including multiple sound source audios, the input audio signal is determined as a target audio signal.
[0186] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DRRAM). Memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0187] The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0188] Optionally, as another embodiment, the processor is further configured to execute the steps of the method described in any one of the aforementioned embodiments when running the computer program.
[0189] The data processing method and system provided by the embodiment of the present application are described in detail below in conjunction with specific application scenarios. Figure 15 As shown, the data processing method may include the following steps:
[0190] S1501, edit the voice file, segment the sentences, and obtain a training data set.
[0191] In an embodiment of the present application, the training speech signal is trimmed and spliced according to sentences. It should be understood that when speaking a sentence, there is a pause between words, and the gain cannot be changed frequently at this time. In other words, when a person speaks a sentence, it is expected that the gain of the sentence is stable. Therefore, it is very necessary to perform sentence trimming training on the training material, and the effect brought about by the comparison of experimental results is very significant. The training file is trimmed to adjust the length according to the sentence, and the same random gain is applied to the same sentence. The trimming during training can make the memory layer of the model have a memory function for the volume characteristics of the entire sentence, so that the volume between words in the sentence is stable in actual use, and the gain stability of the output third audio signal is improved.
[0192] S1502, build the target model framework and design the loss function.
[0193] In an embodiment of the present application, the target model may include the aforementioned first model, or may include a first model and a second model (in this case, it may be called a two-stage model or an AIAGC model), or may include other models that can enhance the first audio signal, which is not specifically limited here.
[0194] Among them, the first model can be a U-Net network in the frequency domain, which is trained through a loss function and mainly optimizes noise recognition related parameters; the second model can be a CRN model in the time domain, which performs enhanced training on the gain.
[0195] In the embodiment of the present application, the loss function (see formula (1)) can be applied simultaneously to the second model in the time domain and the first model in the frequency domain. Through comparative experiments, it was found that the loss function in the frequency domain is conducive to distinguishing noise and speech signals, and the loss function in the time domain is conducive to fitting the gain to the target level value (also called TargetLevel). Therefore, by combining the advantages of the two loss functions, the optimal weight coefficient is assigned and reconstructed into a comprehensive loss function.
[0196] S1503: Train the target model and tune the target model to obtain the optimal weight.
[0197] The target model is trained based on the above loss function to identify whether the current signal is noise. If it is noise, the signal will be suppressed. If it is speech, the speech gain will be controlled as expected.
[0198] like Figure 11 As shown, the target model can be composed of two parts, where the structure of the first model is as follows Figure 14 As shown in , a simplified U-NET structure can be used to distinguish noise and speech signals; the structure of the second model is as follows Figure 13 As shown, a CRN structure can be used to fit the gain.
[0199] In this embodiment, the TargetLevel parameter is introduced during training. The first mask sequence contains noise reduction and gain information. This first mask sequence is combined with the target level value to introduce the influence of the TargetLevel parameter on the model. The final third audio signal output is made to approach the TargetLevel value. The entire target model training is based on the level approximation of TargetLevel. This parameter provides users with a flexible and effective desired level range.
[0200] The first mask sequence is a value sequence with the same length as the input target audio signal, each value indicates whether the signal component at the corresponding sampling point is noise or target signal, and may also include gain information of the corresponding sampling point.
[0201] S1504: Determine a target gain and apply it to the target audio signal.
[0202] In the embodiment of the present application, the target gain may also be referred to as a gain factor, and the two may be equivalent to or replace each other.
[0203] In an embodiment of the present application, by weakening the output of the target model, only the output of the target model is used as part of the calculation of the gain factor, and then the gain factor is applied to the original input target audio signal. Through the effect comparison test, when the target audio signal contains 256 sampling points, compared with the gain factor determined by performing 512-point FFT and IFFT processing, the gain factor determined by performing 256-point FFT and IFFT processing is almost equal. Therefore, in an embodiment of the present application, the effect of spectrum leakage on the derivation of the gain factor is negligible. In an embodiment of the present application, the FFT is directly set to 256 points. When the first frame of audio data has 256 points, a 256-point FFT is directly performed. The result of the processing output is the result of the 256-point data of the first frame, without delay. But obviously, if the data of these 256 points is directly used as the output, the speech will be uneven and the spectrum will leak because there is no overlap. To avoid this result, only the RMS energy value (level value) of the 256 points of the output second audio signal is calculated. This output RMS value minus the RMS energy value of the input signal (target audio signal) is the gain change of this frame signal (i.e., the target gain). For example: RMS output signal - RMS input signal = 5dB, then this frame signal needs to be increased by 5dB, and this 5dB will be directly applied to the input signal. In other words, when the entire model processes a frame of input data, its only purpose is to obtain a gain change of 5dB, and then apply this 5dB to the input signal. In this way, the third audio signal obtained has zero delay.
[0204] S1505 : Deploy the target model and the algorithm for determining the target gain to the real-time voice stream on the end side.
[0205] The target model and the algorithm for determining the target gain based on the second audio signal output by the target model are deployed on a computer device on the end side that needs to perform processing.
[0206] It should also be noted that the loss function based on the following formula (1) can be applied to the first model for training, or applied to the first model and the second model for joint training:
[0207]
[0208] in, is the enhanced waveform output by the model, y is the clean waveform (which can be understood as the target waveform for training), is the enhanced spectrum, and Y is the clean spectrum, corresponding to the enhanced waveform and the clean waveform respectively.
[0209] Where α and β are weight coefficients. α corresponds to the weight coefficient of the part of the model processed in the frequency domain, and β corresponds to the weight coefficient of the part of the model processed in the time domain. Based on experiments, a good combination of weight coefficients may be α = 0.08, β = 0.4.
[0210] Through experimental comparison, it is found that the loss function in the frequency domain is conducive to distinguishing noise and speech signals, and the loss function in the time domain is conducive to fitting the gain to the target level value. Therefore, the loss function of the above formula (1) combines the advantages of the time domain and frequency domain loss functions.
[0211] Each of the above formulas is calculated as shown in the following formulas (2) to (5):
[0212]
[0213] It should be noted that during the loss function-based training process, the speech signals used for training can be trimmed and spliced into sentences to serve as training audio data. When speaking, there are pauses between words, and the gain cannot be changed frequently at this time. In other words, when a person speaks a sentence, the gain of the sentence is expected to be stable. Therefore, when obtaining training audio data, it is very necessary to trim the speech material. In addition, experimental comparisons have shown that trimming the sentences during the experiment can enable the model's memory layer to remember the volume of the entire sentence. In this way, when the model is used, even if the speech signal input according to the audio frame stream is not trimmed, since the training file is adjusted according to the sentence length, the same target gain will be applied to the same sentence output during use to ensure a stable volume between words in the sentence.
[0214] It is understood that in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system. Furthermore, the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The aforementioned integrated units can be implemented in the form of hardware or software functional modules.
[0215] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0216] Therefore, this embodiment provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by at least one processor, the steps of the method in any one of the above embodiments are implemented.
[0217] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof.
[0218] For software implementation, the techniques described herein can be implemented by modules (e.g., procedures, functions, etc.) that perform the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0219] The above is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
[0220] An embodiment of the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.
[0221] It should be understood that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0222] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, they will not be repeated here.
[0223] It should also be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0224] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0225] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0226] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0227] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0228] The above is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A data processing method, comprising: obtaining a target audio signal; Converting the target audio signal into a first audio signal in the frequency domain; The number of points of the conversion processing is consistent with the number of sampling points of the target audio signal; performing enhancement processing on the first audio signal to obtain a second audio signal; determining a target gain based on an energy representation of the second audio signal and an energy representation of the target audio signal; The target audio signal is processed based on the target gain to obtain a third audio signal.
2. The method according to claim 1, wherein converting the target audio signal into a first audio signal in a frequency domain dimension comprises: Performing Fourier transform on the target audio signal to obtain corresponding frequency domain features; The frequency domain feature is divided into a plurality of sub-bands to obtain a first audio signal containing the plurality of sub-bands.
3. The method according to claim 1, wherein the enhancing process is performed on the first audio signal to obtain the second audio signal, comprising: performing noise reduction processing on the first audio signal to obtain a first intermediate audio signal; Converting the first intermediate audio signal into a second intermediate audio signal in a time domain dimension; Perform gain fitting processing on the second intermediate audio signal to obtain the second audio signal.
4. The method according to claim 3, wherein performing gain fitting processing on the second intermediate audio signal to obtain the second audio signal comprises: extracting features of the second intermediate audio signal and performing normalization processing on the second intermediate audio signal to obtain first feature information; Performing sparse activation and time series analysis on the first feature information to obtain a first mask sequence; the length of the first mask sequence is consistent with the length of the target audio signal; The target audio signal is processed based on the first mask sequence, a noise signal in the target audio signal is suppressed, and a speech signal in the target audio signal is amplified to obtain the second audio signal.
5. The method according to claim 4, before processing the target audio signal based on the first mask sequence, the method further comprises: Obtaining a target level value, where the target level value is used to represent an expected level value; updating the first mask sequence based on the target level value to obtain an updated first mask sequence; The updated first mask sequence enables the level value of the second audio signal to match the target level value.
6. The method according to claim 1, wherein the enhancing process is performed on the first audio signal to obtain the second audio signal, comprising: downsampling and transposing the first audio signal to obtain second feature information; performing feature extraction based on the second feature information to determine a second mask sequence; The first audio signal is processed based on the second mask sequence to suppress a noise signal in the first audio signal to obtain a second audio signal.
7. The method according to any one of claims 1 to 6, wherein obtaining the target audio signal comprises any one of the following: obtaining an input audio signal, and in response to the input audio signal including multiple sound source audios, separating the input audio signal to obtain input sub-audio signals corresponding to the respective sound source audios, and determining the input sub-audio signals as the target audio signals; An input audio signal is obtained, and in response to the input audio signal including multiple sound source audios, the input audio signal is determined as the target audio signal.
8. A data processing system comprising: The first module is used to obtain a target audio signal; A second module is configured to convert the target audio signal into a first audio signal in a frequency domain dimension; The number of points of the conversion processing is consistent with the number of sampling points of the target audio signal; A third module is configured to perform enhancement processing on the first audio signal to obtain a second audio signal; The fourth module is configured to determine a target gain based on the energy representation of the second audio signal and the energy representation of the target audio signal, and process the target audio signal based on the target gain to obtain a third audio signal.
9. The system according to claim 8, wherein the third module comprises: a first model, configured to perform noise reduction processing on the first audio signal to obtain a first intermediate audio signal, and convert the first intermediate audio signal into a second intermediate audio signal in a time domain; The second model is used to perform gain fitting processing on the second intermediate audio signal to obtain the second audio signal.
10. A computer device comprising at least one audio acquisition component, at least one audio output component, at least one memory, and at least one processor, wherein: The at least one audio acquisition component is used to acquire audio signals; The memory is used to store a plurality of computer instructions; The processor is configured to load and execute the computer instructions to implement the following method: Obtaining a target audio signal; converting the target audio signal into a first audio signal in a frequency domain dimension; wherein the number of points processed by the conversion is consistent with the number of sampling points of the target audio signal; performing enhancement processing on the first audio signal to obtain a second audio signal; determining a target gain based on an energy representation of the second audio signal and an energy representation of the target audio signal; processing the target audio signal based on the target gain to obtain a third audio signal; The audio output component is configured to perform output based on a third audio signal.
Citation Information
Patent Citations
Speech signal processing method and device
CN110875049A
Audio signal processing method and device, equipment and medium
CN116524950A
Speech enhancement
GB202318554D0
Low power voice detection
US20140236582A1
Method for training speech enhancement network, method for enhancing speech, and electronic device
US20250391419A1