Speech processing method and apparatus, storage medium, and computer device
By identifying the frequency points of the noise intensification phase in the power spectrum and performing probability correction, the problems of noise residue and distortion in speech frames are solved, and the clarity of speech frames is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU SHIYUAN ELECTRONICS CO LTD
- Filing Date
- 2022-02-22
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies suffer from residual noise and speech distortion when processing voice frames, which affects the voice function of electronic devices.
By identifying target frequency points in the power spectrum that are in the noise gradually increasing stage, the probability of local speech presence is corrected using the global speech presence probability, a gain factor is obtained, and gain processing is performed to generate the target speech frame.
It reduces residual noise and speech distortion in the target speech frame, thus improving the clarity of the speech frame.
Smart Images

Figure CN116682447B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech technology, and more specifically, to a speech processing method, apparatus, storage medium, and computer device. Background Technology
[0002] With the rapid development of voice technology, more and more electronic devices are equipped with voice recognition, voice calls, and voice control functions based on voice technology. Since daily life involves a wide variety of voices, the voice frames captured by electronic devices inevitably contain a certain amount of noise, which can affect the functionality of these devices to some extent. Therefore, in existing technologies, electronic devices perform noise reduction processing on the voice frames after they are captured. Summary of the Invention
[0003] This application provides a speech processing method, apparatus, storage medium, and computer device that can solve the technical problem of how to improve the clarity of a target speech frame.
[0004] In a first aspect, embodiments of this application provide a voice processing method, the method comprising:
[0005] An initial speech frame is acquired from the speech signal, and the initial spectrum and global speech presence probability corresponding to the initial speech frame are obtained. An initial power spectrum is obtained based on the initial spectrum, and the initial power spectrum includes multiple frequency points and the power value of each frequency point among the multiple frequency points.
[0006] Based on the initial power spectrum, each target frequency point that satisfies the noise gradual increase stage is determined among the multiple frequency points, and the local speech existence probability corresponding to each frequency point is obtained.
[0007] Based on the global speech existence probability, the local speech existence probability corresponding to each target frequency point is corrected to obtain the target speech existence probability corresponding to each target frequency point.
[0008] Based on the probability of the presence of target speech corresponding to each target frequency point and the probability of the presence of local speech corresponding to other frequency points, the gain factor corresponding to each frequency point is obtained, wherein the other frequency points are the frequency points among the multiple frequency points that are not in the noise gradually increasing stage.
[0009] Based on the gain factor corresponding to each frequency point, the initial spectrum is processed to obtain the target spectrum, and the target speech frame corresponding to the initial speech frame is generated based on the target spectrum.
[0010] Secondly, embodiments of this application provide a voice processing apparatus, including:
[0011] The power spectrum acquisition module is used to acquire an initial speech frame in a speech signal, acquire the initial spectrum corresponding to the initial speech frame and the global speech presence probability, and acquire an initial power spectrum based on the initial spectrum. The initial power spectrum includes multiple frequency points and the power value of each frequency point among the multiple frequency points.
[0012] The frequency point determination module is used to determine each target frequency point that satisfies the noise gradual increase stage among the plurality of frequency points based on the initial power spectrum;
[0013] The probability acquisition module is used to acquire the probability of local speech presence corresponding to each frequency point;
[0014] The probability correction module is used to perform probability correction on the local speech existence probability corresponding to each target frequency point based on the global speech existence probability, so as to obtain the target speech existence probability corresponding to each target frequency point.
[0015] The factor acquisition module is used to acquire the gain factor corresponding to each frequency point based on the probability of the existence of target speech corresponding to each target frequency point and the probability of the existence of local speech corresponding to other frequency points, wherein the other frequency points are the frequency points among the multiple frequency points that are not in the noise gradually increasing stage.
[0016] The speech frame generation module is used to perform gain processing on the initial spectrum based on the gain factor corresponding to each frequency point to obtain the target spectrum, and generate the target speech frame corresponding to the initial speech frame based on the target spectrum.
[0017] Thirdly, embodiments of this application provide a storage medium storing a computer program adapted to be loaded by a processor and to execute the steps of the above-described method.
[0018] Fourthly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described above.
[0019] In this embodiment, by first identifying the target frequency point in the noise gradually increasing stage of the power spectrum, and then, based on the global speech presence probability of the initial speech frame, the probability of local speech presence corresponding to the target frequency point is specifically corrected, thereby improving the accuracy of the local speech presence probability corresponding to each frequency point. This improves the accuracy of the gain factor calculated based on the local speech presence probability, thereby reducing the residual noise signal and speech distortion in the target speech frame obtained based on the initial speech frame, and improving the clarity of the target speech frame. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A schematic flowchart of a speech processing method provided in an embodiment of this application;
[0022] Figure 2 A schematic flowchart of a speech processing method provided in an embodiment of this application;
[0023] Figure 3 An example diagram illustrating a current speech cycle provided in an embodiment of this application;
[0024] Figure 4 A schematic diagram illustrating an example of a first frequency point sequence provided in an embodiment of this application;
[0025] Figure 5 A schematic diagram illustrating an example of a second frequency point sequence provided in an embodiment of this application;
[0026] Figure 6 A schematic flowchart of a speech processing method provided in an embodiment of this application;
[0027] Figure 7 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application;
[0028] Figure 8 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application;
[0029] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0030] To make the features and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] Existing speech denoising methods involve obtaining the frequency points corresponding to the current speech frame, determining the minimum power value corresponding to each frequency point, calculating the posterior signal-to-noise ratio, prior signal-to-noise ratio, and local speech presence probability based on the minimum power value, obtaining the noise power spectrum corresponding to the current speech frame, obtaining the gain factor corresponding to each frequency point based on the aforementioned parameters, and finally performing gain processing on the current speech frame based on the gain factor corresponding to each frequency point to obtain the target speech frame.
[0032] During voice interaction, various sudden external noises, such as motor vibrations from washing machines and air conditioners, can cause a large amount of noise signals to remain in the target voice frame obtained based on the aforementioned noise reduction process, accompanied by voice distortion and other problems.
[0033] The following will combine Figures 1-6 This paper provides a detailed description of the speech processing method provided in the embodiments of this application.
[0034] Please see Figure 1 This is a flowchart illustrating a speech processing method provided in an embodiment of this application. Figure 1 As shown, the method may include the following steps S101-S105.
[0035] S101, an initial speech frame is acquired from the speech signal, the initial spectrum corresponding to the initial speech frame and the global speech presence probability are obtained, and an initial power spectrum is obtained based on the initial spectrum. The initial power spectrum includes multiple frequency points and the power value of each frequency point among the multiple frequency points.
[0036] In one implementation, the speech signal comprises multiple consecutive speech frames. When processing the speech signal, the speech processing device acquires the speech frames one by one in sequence and processes the currently acquired frame. Therefore, the initial speech frame is the speech frame acquired at the current moment during the continuous acquisition of speech frames. The global speech existence probability refers to the probability of speech presence in a complete initial speech frame. The initial spectrum is the frequency distribution curve, i.e., the frequency spectral density. The initial spectrum refers to a noisy initial spectrum obtained by performing signal processing such as short-time Fourier transform (STFT) on the initial speech frame. This initial spectrum includes multiple frequency points and the amplitude values of each frequency point. An initial power spectrum is obtained based on the initial spectrum, which also includes multiple frequency points and the power values of each frequency point. For example, the amplitude value of frequency a in the initial spectrum is |Y|. a |Then the power value at frequency a in the initial power spectrum is |Y a | 2 .
[0037] The speech processing device includes a speech acquisition device for acquiring speech signals from the environment in which the speech processing device operates. The speech processing device acquires the speech signals from the speech acquisition device and extracts an initial speech frame from the speech signals. Then, speech preprocessing is performed based on the initial speech frame, followed by speech feature extraction based on the preprocessed speech signal to obtain speech features such as pitch, Bark-frequency cepstral coefficients (BFCC), and Mel-scale frequency cepstral coefficients (MFCC). Finally, based on the speech features in the initial speech frame, the global speech presence probability corresponding to the initial speech frame is obtained.
[0038] The speech processing device can also obtain the initial power spectrum corresponding to the initial speech frame based on the speech signal after speech preprocessing.
[0039] S102, based on the initial power spectrum, determine each target frequency point in the noise gradually increasing stage among the multiple frequency points, and obtain the local speech existence probability corresponding to each frequency point.
[0040] In one implementation, the noise gradation phase refers to the phase where the noise energy gradually increases. Optionally, this noise gradation phase can also be a non-stationary noise phase where the noise energy gradually increases, but the magnitude of the increase is not stable. The probability of local speech presence refers to the probability of speech presence in the speech segment corresponding to the frequency point in the initial power spectrum.
[0041] The speech processing device sequentially acquires the power values of each frequency point in the initial power spectrum. Based on the power value of the currently acquired frequency point, it determines whether the currently acquired frequency point is in a noise gradually increasing phase. If the currently acquired frequency point is in a noise gradually increasing phase, it designates the currently acquired frequency point as the target frequency point. Simultaneously, based on the power value of the currently acquired frequency point, it calculates the local speech presence probability of the currently acquired frequency point. This allows it to identify the target frequency point in the initial power spectrum and acquire the local speech presence probabilities corresponding to multiple frequency points in the initial power spectrum.
[0042] S103, based on the global speech existence probability, perform probability correction on the local speech existence probability corresponding to each target frequency point to obtain the target speech existence probability corresponding to each target frequency point.
[0043] In one embodiment, the speech processing device acquires a correction factor, which is a parameter corresponding to the speech processing device and is used to control the degree of probability correction of the probability of local speech presence. Optionally, the correction factor can be a preset parameter or an empirical value of the speech processing device, that is, an empirical value for probability correction of the probability of local speech presence. Further, the value range of the correction factor is (0, 1).
[0044] For example, suppose the probability of local speech presence at the k-th frequency point in the initial power spectrum is p(k, l), and the probability of global speech presence at the initial speech frame is p. global The correction factor is α, and the probability of local speech at the k-th frequency point in the initial power spectrum after correction, i.e., the probability of the target speech, is: It should be noted that l represents the number of frames in the speech signal for the initial speech frame. The correction process for the probability of local speech presence can then be represented by the following formula:
[0045]
[0046] S104, based on the probability of the presence of target speech corresponding to each target frequency point and the probability of the presence of local speech corresponding to other frequency points, obtain the gain factor corresponding to each frequency point, wherein the other frequency points are the frequency points among the plurality of frequency points that are not in the noise gradually increasing stage.
[0047] In one implementation, the gain factor is a parameter used to amplify the power value of a frequency point. The speech processing device divides multiple frequency points in the initial power spectrum into a target frequency point and other frequency points, where the target frequency point is the frequency point in the noise gradation phase, and the other frequency points are the frequency points not in the noise gradation phase.
[0048] The speech processing device first obtains the gain factor corresponding to each frequency point based on the probability of the existence of target speech corresponding to each target frequency point and the probability of the existence of local speech corresponding to other frequency points.
[0049] S105, based on the gain factor corresponding to each frequency point, perform gain processing on the initial spectrum to obtain the target spectrum, and generate the target speech frame corresponding to the initial speech frame based on the target spectrum.
[0050] In this embodiment, by first identifying the target frequency point in the noise gradually increasing stage of the power spectrum, and then, based on the global speech presence probability of the initial speech frame, the probability of local speech presence corresponding to the target frequency point is specifically corrected, thereby improving the accuracy of the local speech presence probability corresponding to each frequency point. This improves the accuracy of the gain factor calculated based on the local speech presence probability, thereby reducing the residual noise signal and speech distortion in the target speech frame obtained based on the initial speech frame, and improving the clarity of the target speech frame.
[0051] Understandably, since the initial power spectrum obtained from the initial speech frame includes multiple frequency points, the correlation processing of the initial power spectrum can be achieved by processing each frequency point. This will be combined with... Figures 2-4 The document provides a detailed introduction to the speech processing methods corresponding to each frequency point.
[0052] Please see Figure 2 This is a flowchart illustrating a speech processing method provided in an embodiment of this application. Figure 2 As shown, the method may include the following steps S201-S211.
[0053] S201, Acquire the initial speech frame from the speech signal.
[0054] In one embodiment, the speech signal includes multiple consecutive speech frames. When the speech processing device performs speech processing on the speech signal, it acquires the speech frames one by one in the order of the speech frames and performs speech processing on the currently acquired speech frame. It can be understood that the initial speech frame is a speech frame acquired at the current moment during the process of the speech processing device continuously acquiring speech frames.
[0055] The voice processing device includes a voice acquisition device for acquiring voice signals from the environment in which the voice processing device is located. The voice processing device acquires the voice signals from the voice acquisition device and extracts initial voice frames from the voice signals.
[0056] S202, using a probability estimation model, obtain the global speech presence probability of the initial speech frame.
[0057] In one implementation, the global speech presence probability refers to the probability of speech presence in a complete initial speech frame. The probability estimation model is a neural network model used to obtain the global speech presence probability in the initial speech frame. This model can be a Gated Recurrent Unit (GRU) model, a Long Short-Term Memory (LSTM) model, or a Time Delay Neural Network (TDNN) model, etc., and is not limited thereto. Furthermore, the probability estimation model is trained based on multiple training speech frames and their corresponding global speech presence probabilities.
[0058] After acquiring the initial speech frame, the speech processing device inputs the initial speech frame into the probability estimation model to obtain the global speech presence probability of the initial speech frame through the probability estimation model.
[0059] Optionally, the probability estimation model will perform relevant processing on the initial speech frame to obtain the speech features corresponding to the initial speech frame, such as pitch, Bark-frequency cepstral coefficients (BFCC), and Mel-scale frequency cepstral coefficients (MFCC), so as to calculate the probability of speech presence in the initial speech frame based on the extracted speech features.
[0060] S203, Perform a Fourier transform on the initial speech frame to obtain the initial spectrum corresponding to the initial speech frame. The initial spectrum includes multiple frequency points and the amplitude values corresponding to the multiple frequency points.
[0061] In one implementation, the Fourier transform can be a short-time Fourier transform (STFT), and the initial spectrum includes multiple frequency points and the amplitude values corresponding to the multiple frequency points.
[0062] After acquiring the initial speech frame, the speech processing device first performs signal processing such as short-time Fourier transform on the initial speech frame to obtain an initial spectrum containing noise. The amplitude value of the k-th frequency point in the initial spectrum is |Y(k, l)|, where l represents the number of the initial speech frame in the speech signal.
[0063] S204, Based on the initial spectrum, obtain the initial power spectrum corresponding to the initial speech frame.
[0064] In one implementation, the initial power spectrum includes multiple frequency points and the power value of each frequency point.
[0065] The speech processing device calculates the power value of each frequency point based on the amplitude value of each frequency point in the initial power spectrum, and generates the initial power spectrum corresponding to the initial speech frame based on the power value of each frequency point.
[0066] For example, if the amplitude value of the k-th frequency point in the initial spectrum is Y(k, l), then the power value of the k-th frequency point in the initial power spectrum is |Y(k, l)|. 2 That is, to obtain the square of the amplitude value of the k-th frequency point, and use the obtained square value as the power value of the k-th frequency point.
[0067] In one implementation, the speech processing device acquires the initial power spectrum |Y(k,l)| 2 Afterwards, the initial power spectrum can be smoothed in time and dimension to obtain a new |Y(k,l)|. 2 That is, the temporal smoothing is shown in the following formula:
[0068]
[0069] Dimensional smoothing is shown in the following formula:
[0070] S(k, l) = α s (k, l)S(k, l-1)+S f (k, l)
[0071] It should be noted that b(i) is the standardized window function, 2w+1 is the window length, and |Y(ki,l)| 2 This represents the power value of the noisy speech under the short-time Fourier transform in the time-frequency domain.
[0072] By inputting the initial speech frame into the trained probability estimation model, the global speech presence probability of the initial speech frame can be directly obtained, reducing the tedious probability estimation process and improving speech processing efficiency.
[0073] S205, obtain a first frequency point from the initial power spectrum, wherein the first frequency point is any one of the frequency points.
[0074] In one implementation, since the initial power spectrum includes multiple frequency points, the speech processing device, when performing relevant speech processing on the initial speech frame based on the initial power spectrum, will sequentially acquire the first frequency point in the initial power spectrum until the initial power spectrum has been traversed completely. It should be noted that the first frequency point is any one of the frequency points in the initial power spectrum.
[0075] S206, if the audio phase corresponding to the first frequency point is a noise gradually increasing phase, then the first frequency point is determined to be the target frequency point that satisfies the noise gradually increasing phase. The audio phase corresponding to the first frequency point is determined by the speech frame position of the historical speech frame to which the historical frequency point with the minimum power value in the historical speech cycle belongs. The historical speech cycle is the previous speech cycle adjacent to the current speech cycle in which the initial speech frame is located. The historical frequency point has the same frequency as the first frequency point.
[0076] In one implementation, the speech processing device stores audio segments corresponding to each frequency point within the current speech cycle, and the audio segments corresponding to each frequency point are updated once per cycle. Specifically, the audio segments corresponding to each frequency point within the current speech cycle are determined based on the speech frame position of the historical speech frame to which the historical frequency point corresponding to the minimum power value belongs in a historical speech cycle. This historical speech cycle is the previous speech cycle adjacent to the current speech cycle in which the initial speech frame is located, and the historical frequency point has the same frequency as the first frequency point. It can be understood that the speech frame position refers to the position of the historical speech frame within the historical speech cycle.
[0077] The speech processing device directly acquires the audio stage corresponding to the first frequency point, and when the audio stage corresponding to the first frequency point is a noise gradation stage, it determines the first frequency point as the target frequency point that satisfies the noise gradation stage.
[0078] For example, when the voice processing device obtains a first frequency point in the initial power spectrum, it obtains the first frequency corresponding to the first frequency point, and then, based on the first frequency, determines a storage frequency point in the storage medium that is consistent with the first frequency point, and then uses the audio stage corresponding to the storage frequency point as the audio stage corresponding to the first frequency point.
[0079] By acquiring the audio stages corresponding to each frequency point stored in the speech processing device, the target frequency point that satisfies the noise gradation stage is determined in the initial speech power spectrum. Since the audio stages of each frequency point stored in the speech processing device are obtained based on the recognition of the previous historical speech cycle, the target frequency point that satisfies the noise gradation stage in the current speech cycle is determined by periodically updating the audio stages corresponding to each frequency point. The new noise gradation stage formed by the sudden noise in the speech signal is identified, thereby improving the accuracy of the audio stages corresponding to each frequency point.
[0080] S207, if the initial voice frame is the last frame of the current voice cycle, then obtain the first frequency of the first frequency point.
[0081] In one embodiment, the speech signal includes multiple speech cycles, and each speech cycle includes multiple speech frames arranged in the acquisition order. The speech cycle containing the initial speech frame is taken as the current speech cycle. The initial power spectrum includes multiple first frequency points with different frequencies, and the power values of the first frequency points.
[0082] When the initial speech frame is the last frame in the current speech cycle, the speech processing device acquires the first frequency of the first frequency point.
[0083] S208, in the current speech cycle, obtain the first frequency point sequence corresponding to the first frequency, and arrange all frequency points in the first frequency point sequence according to the acquisition order of the speech frames in the current speech cycle.
[0084] In one embodiment, the first frequency point sequence includes multiple frequency points arranged according to the acquisition order of speech frames in the current speech cycle, and the frequency corresponding to each frequency point in the first frequency point sequence is a first frequency.
[0085] For example, Figure 3 An example diagram of a current speech cycle is shown. For example... Figure 3As shown, the current speech cycle C1 includes multiple speech frames VF arranged in the acquisition order. Each speech frame VF has multiple corresponding frequency points FQ. The frequencies of each frequency point FQ in the multiple frequency points FQ corresponding to the speech frame VF are different. Frequency points FQ with the same frequency form the first frequency point sequence L1 corresponding to that frequency.
[0086] The speech processing device first acquires the initial power spectrum corresponding to each speech frame in the current speech cycle. Then, based on the initial power spectrum corresponding to each speech frame, it acquires multiple frequency points corresponding to the first frequency from multiple frequency points corresponding to each speech frame. Finally, based on the acquisition order of each speech frame and the frequency points corresponding to the multiple first frequencies, it acquires the first frequency point sequence corresponding to the first frequency.
[0087] S209, obtain the second frequency point corresponding to the first minimum power value in the first frequency point sequence, and obtain the voice frame position of the voice frame to which the second frequency point belongs in the current voice cycle.
[0088] In one embodiment, the speech processing device compares the power values of each frequency point in the first frequency point sequence one by one, obtains the minimum value among the power values of multiple frequency points, namely the first minimum power value, then determines the second frequency point corresponding to the first minimum power value in the first frequency point sequence, then determines the speech frame to which the second frequency point belongs, and finally obtains the speech frame position of the speech frame to which the second frequency point belongs in the current speech cycle.
[0089] For example, Figure 4 A schematic diagram illustrating an example of a first frequency point sequence is shown. For example... Figure 4 As shown, the first frequency point sequence L1 includes multiple frequency points FQ ( Figure 4 (The diagram shows 12 frequency points). Assume that the second frequency point FQ corresponding to the first minimum value is the frequency point FQ corresponding to frame2. min Then the second frequency point FQ min The position of the speech frame 2 in the current speech cycle C1 is 2.
[0090] S210, based on the voice frame position of the voice frame to which the second frequency point belongs and the set position range in the current voice cycle, determine the audio stage corresponding to the third frequency point indicated by the first frequency in the target voice cycle, wherein the target voice cycle is the next voice cycle adjacent to the current voice cycle.
[0091] In one implementation, the set position range is used to determine the position range within a target speech cycle that satisfies the noise crescendo phase. This set position range is determined based on the cycle length of the speech cycle. The target speech cycle is the next speech cycle adjacent to the current speech cycle. Optionally, if the speech cycle is defined by the speech processing device, then after determining the cycle length of the speech cycle, the set position range corresponding to the speech cycle is also obtained. For example, if the speech cycle length is A and the upper limit threshold of the set position range is X, then the upper limit of the set position range is AX, and the lower limit of the set position range is 1, i.e., the set position range is [1, AX]. Further, for example, if the speech cycle length is 12 and the upper limit threshold of the set position range is 0.25, then the set position range is [1, 0.25 * 12], i.e., [1, 3]. It should be noted that if the value of AX contains a decimal part, then the integer part of AX is used as the upper limit of the set position range. For example, if AX = 13 * 0.25 = 3.25, then the upper limit of the set position range is 3, i.e., the set position range is [1, 3]. It can be understood that the set location range is at the beginning of the speech cycle.
[0092] The speech processing device determines the audio stage corresponding to the third frequency point indicated by the first frequency in the target speech cycle based on the speech frame position of the speech frame to which the second frequency point belongs and the set position range in the current speech cycle. The third frequency point indicated by the first frequency in the target speech cycle refers to all frequency points in the target speech cycle with the first frequency.
[0093] By obtaining the second frequency point corresponding to the minimum power value in the current speech cycle, the position of the speech frame corresponding to the second frequency point in the current speech cycle is determined. Then, by judging whether the speech frame corresponding to the minimum power value is at the beginning of the speech cycle, it is determined whether there is a noise gradation phase in the speech cycle. Based on the current judgment result, the frequency point corresponding to the noise gradation phase is determined in the next speech cycle. That is, the frequency points that satisfy the noise gradation phase are dynamically updated, and then the probability of the frequency points that satisfy the noise gradation phase is dynamically corrected, which improves the recognition accuracy of the noise gradation phase and the probability accuracy of the target frequency point.
[0094] In one embodiment, when determining the audio stage corresponding to the third frequency point, the voice processing device may first determine whether the voice frame position is within a set position range in the current voice cycle. If the voice frame position is within the set position range in the current voice cycle, then the audio stage corresponding to the third frequency point corresponding to the first frequency in the target voice cycle is determined to be the noise gradually increasing stage.
[0095] The speech processing device first determines whether the speech frame position is within the set position range in the current speech cycle, that is, it compares the speech frame position with the upper limit and lower limit of the set position range. If the speech frame position is less than or equal to the upper limit and greater than or equal to the lower limit, it is determined that the speech frame position is within the set position range in the current speech cycle. Then, it is determined that the audio stages corresponding to the multiple third frequency points corresponding to the first frequency in the target speech cycle are all noise gradually increasing stages.
[0096] It is understandable that if the position of the speech frame is not within the set position range in the current speech cycle, it can be determined that the audio stages corresponding to the multiple third frequency points corresponding to the first frequency in the target speech cycle are all non-noise gradation stages.
[0097] By determining whether the position of the speech frame is within a set range, the audio phase of each frequency point in the target speech cycle is obtained. Based on the speech situation of the current speech cycle, the probability of speech presence in the next speech cycle is specifically corrected, thereby improving the accuracy of the local speech presence probability corresponding to each target frequency point in the next speech cycle.
[0098] S211, Obtain the second frequency point sequence corresponding to the first frequency of the first frequency point, the second frequency point sequence includes the fourth frequency point corresponding to the first frequency in each historical speech frame of the historical speech cycle, and the fifth frequency point corresponding to the first frequency in the speech frames already obtained in the current speech cycle.
[0099] In one implementation, the first frequency is the frequency of the first frequency point currently traversed. The historical speech period is the previous speech period adjacent to the current speech period to which the initial speech frame belongs. For example, Figure 5 A schematic diagram illustrating an example of a second frequency point sequence is shown. For example... Figure 5 As shown, the second frequency sequence L2 includes the fourth frequency FQ4 corresponding to the first frequency of each historical speech frame in the historical speech period C2, and the fifth frequency FQ5 corresponding to the first frequency of the speech frame acquired in the current speech period C1.
[0100] The speech processing device acquires the fourth frequency point corresponding to the first frequency in each historical speech frame of the historical speech cycle, and acquires the fifth frequency point corresponding to the first frequency in each acquired speech frame of the current speech cycle, and sorts them according to each historical speech frame and the acquisition order of each speech frame to obtain the second frequency point sequence corresponding to the first frequency.
[0101] S212, based on the power values of each frequency point in the second frequency point sequence, obtain the second minimum power value.
[0102] In one embodiment, the voice processing device compares the power values of each frequency point in the second frequency point sequence one by one to obtain the minimum value among the power values of multiple frequency points, namely the second minimum power value.
[0103] S213, obtain the local speech existence probability corresponding to the first frequency point based on the power value of the first frequency point and the second minimum power value.
[0104] In one embodiment, the speech processing device obtains the probability of local speech presence corresponding to the first frequency point based on the power value of the first frequency point and the second minimum power value.
[0105] By obtaining the minimum power value in the current speech cycle and the historical speech cycle, the minimum value within the two cycles is obtained, which improves the accuracy of the minimum value parameter in the calculation of the probability of local speech presence, and thus improves the accuracy of the probability of local speech presence corresponding to each frequency point in the initial speech frame.
[0106] In one embodiment, when the voice device acquires the probability of local speech presence, it can also acquire the probability of local speech absence corresponding to the first frequency point based on the power value of the first frequency point and the second minimum power value; and acquire the probability of local speech presence corresponding to the first frequency point based on the power value of the first frequency point and the probability of local speech absence.
[0107] The probability of local speech presence refers to the probability of speech presence in a speech segment corresponding to a frequency point in the initial power spectrum. The speech processing device first obtains the probability of local speech absence corresponding to the first frequency point based on the power value of the first frequency point and the second minimum power value, and then obtains the probability of local speech presence corresponding to the first frequency point based on the probability of local speech absence.
[0108] By sequentially acquiring the first frequency point in the initial power spectrum and then performing relevant processing on each acquired first frequency point, the initial power spectrum is processed through the processing of each frequency point. This allows the target frequency point in the initial power spectrum to be identified, and the local speech presence probability of the target frequency point to be obtained. This enables the speech processing device to perform probability correction on the local speech presence probability corresponding to the target frequency point, thereby improving the accuracy of the local speech presence probability corresponding to each frequency point.
[0109] In one embodiment, the speech processing device can obtain the probability of local speech absence corresponding to the first frequency point based on the power value of the first frequency point and the second minimum power value; obtain the sixth frequency point adjacent to the first frequency point in the second frequency point sequence, and obtain the historical noise power value, historical posterior signal-to-noise ratio and historical gain factor corresponding to the sixth frequency point; and obtain the probability of local speech presence corresponding to the first frequency point based on the historical noise power value, historical posterior signal-to-noise ratio, historical gain factor, power value of the first frequency point and the probability of local speech absence.
[0110] In one implementation, the speech processing device first obtains the probability that local speech corresponding to the first frequency point does not exist based on the power value of the first frequency point and the second minimum power value. For example, the probability of local speech not existing corresponding to the first frequency point is obtained as shown in the following formula:
[0111]
[0112]
[0113]
[0114] in, is the probability that the local speech corresponding to the k-th frequency point does not exist, and l is the number of frames of the initial speech frame in the speech signal; B is the second minimum power value corresponding to the k-th frequency point; min γ1 is the deviation of the minimum noise estimation, which is an empirical value of the speech processing device and can be adjusted according to the application scenario; γ1 is a threshold constant, usually set to 3.0.
[0115] The speech processing device first obtains the sixth frequency point adjacent to the first frequency point in the frequency point sequence. It should be noted that since the first frequency point is the frequency point corresponding to the latest acquired initial speech frame, the first frequency point must be located at the end of the corresponding second frequency point sequence. Therefore, there is only one sixth frequency point adjacent to the first frequency point.
[0116] The speech processing device retrieves the historical noise power value, historical posterior signal-to-noise ratio, and historical gain factor corresponding to the sixth frequency point from the storage module. Then, based on the historical noise power value, historical posterior signal-to-noise ratio, historical gain factor, power value, and the probability of local speech absence, it obtains the probability of local speech presence corresponding to the first frequency point. This means that the historical speech frame corresponding to the sixth frequency point is the previous initial speech frame, and the relevant parameters obtained by the speech processing device are the relevant parameters corresponding to the first frequency generated during the processing of the previous speech frame.
[0117] By obtaining the probability of the absence of local speech corresponding to the first frequency point, and then determining the relevant parameters corresponding to the first frequency in the previous initial speech frame, the probability of the presence of local speech corresponding to the first frequency point is calculated. This increases the correlation between each speech frame in the speech signal and avoids the error caused by separately calculating the probability of the presence of local speech corresponding to each frequency point in the initial speech frame, thereby improving the accuracy of the probability of the presence of local speech.
[0118] In one embodiment, the speech processing device can obtain the posterior signal-to-noise ratio (SNR) corresponding to the first frequency point based on historical noise power values and the power value of the first frequency point; obtain the prior SNR corresponding to the first frequency point based on historical posterior SNR and historical gain factors; and obtain the probability of local speech presence corresponding to the first frequency point based on prior SNR, posterior SNR, and the probability of local speech absence.
[0119] In one implementation, the posterior signal-to-noise ratio corresponding to the first frequency point is obtained as shown in the following formula:
[0120]
[0121] Where, λ d (k, l-1) represents the noise power value corresponding to the k-th frequency point in the (l-1)-th frame (i.e., the sixth frequency point), which is the historical noise power value; |Y(k, l)| 2 γ is the power value of the k-th frequency point corresponding to the initial speech frame; γ(k, l) is the posterior signal-to-noise ratio corresponding to the k-th frequency point.
[0122] The prior signal-to-noise ratio corresponding to the first frequency point is obtained as shown in the following formula:
[0123]
[0124] Where ξ(k, l) is the prior signal-to-noise ratio corresponding to the k-th frequency point; γ(k, l-1) is the gain factor corresponding to the k-th frequency point of the (l-1)-th frame (i.e., the sixth frequency point), which is the historical gain factor; γ(k, l-1) is the posterior signal-to-noise ratio corresponding to the k-th frequency point of the (l-1)-th frame (i.e., the sixth frequency point), which is the historical posterior signal-to-noise ratio.
[0125] The probability of the local speech corresponding to the first frequency point is obtained by the following formula:
[0126]
[0127] Where p(k, l) is the probability of the presence of local speech corresponding to the k-th frequency point, and q(k, l) is the probability of the absence of local speech corresponding to the k-th frequency point.
[0128] In one embodiment, the voice processing device can also generate a noise power value corresponding to the first frequency point, as shown in the following formula:
[0129] λ d (k, l) = α d (k, l)λ d (k, l-1)+[1-α] d [(k, l)]|Y(k, l)| 2
[0130] Where, α d (k, l) = +(1-α) d p(k, l), α d It is a smoothing constant, ranging from (0, 1), and is usually taken as 0.9.
[0131] By determining the relevant parameters corresponding to the first frequency in the previous initial speech frame, the probability of local speech presence at the first frequency point is calculated, which increases the correlation between each speech frame in the speech signal and avoids the error caused by separately calculating the probability of local speech presence at each frequency point in the initial speech frame, thereby improving the accuracy of the probability of local speech presence.
[0132] S214, based on the global speech existence probability, perform probability correction on the local speech existence probability corresponding to each target frequency point to obtain the target speech existence probability corresponding to each target frequency point.
[0133] In one embodiment, the speech processing device acquires a correction factor, which is a parameter corresponding to the speech processing device and is used to control the degree of probability correction of the probability of local speech presence. Optionally, the correction factor can be a preset parameter or an empirical value of the speech processing device, that is, an empirical value for probability correction of the probability of local speech presence. Further, the value range of the correction factor is (0, 1).
[0134] For example, suppose the probability of local speech presence at the k-th frequency point in the initial power spectrum is p(k, l), and the probability of global speech presence at the initial speech frame is p. global The correction factor is α, and the probability of local speech at the k-th frequency point in the initial power spectrum after correction, i.e., the probability of the target speech, is: It should be noted that l represents the number of frames in the speech signal for the initial speech frame. The correction process for the probability of local speech presence can then be represented by the following formula:
[0135]
[0136] S215, if the first frequency point is the target frequency point, then based on the probability of the existence of the target speech corresponding to the first frequency point, obtain the gain factor corresponding to the first frequency point.
[0137] In one implementation, the gain factor is a parameter used to amplify the power value of a frequency point. The speech processing device divides multiple frequency points in the initial power spectrum into a target frequency point and other frequency points, where the target frequency point is the frequency point in the noise gradation phase, and the other frequency points are the frequency points not in the noise gradation phase.
[0138] The speech processing device first determines the frequency type of the first frequency point. If the first frequency point is the target frequency point, it obtains the gain factor corresponding to the first frequency point based on the probability of the existence of the target speech, the prior signal-to-noise ratio, and the posterior signal-to-noise ratio corresponding to the first frequency point.
[0139] For example, the process of obtaining the gain factor corresponding to each frequency point can be shown in the following formula:
[0140]
[0141] in, The lower limit of the gain can be 0.02.
[0142] When the first frequency point is the target frequency point, p(k, l) in the formula is the probability of the local target speech corresponding to the first frequency point.
[0143] S216, if the first frequency point is another frequency point, then based on the local speech existence probability corresponding to the first frequency point, obtain the gain factor corresponding to the first frequency point.
[0144] In one implementation, the speech processing device first determines the frequency type of a first frequency point. If the first frequency point is another frequency point, it obtains the gain factor corresponding to the first frequency point based on the local speech presence probability, prior signal-to-noise ratio, and posterior signal-to-noise ratio. When the first frequency point is another frequency point, p(k, l) in the aforementioned formula represents the local speech presence probability corresponding to the first frequency point.
[0145] By determining the frequency type of the first frequency point, the calculation parameters corresponding to the gain factor are obtained based on the frequency type of the first frequency point. That is, the probability of the presence of the target speech is used for the target frequency point, and the probability of the presence of local speech is used for other frequency points. This improves the accuracy of the gain factor calculated based on the probability of the presence of local speech, thereby reducing the residual noise signal and speech distortion in the target speech frame obtained based on the initial speech frame and improving the clarity of the target speech frame.
[0146] S217, based on the gain factor corresponding to each frequency point, the initial spectrum is subjected to gain processing to obtain the target spectrum, and the target speech frame corresponding to the initial speech frame is generated based on the target spectrum.
[0147] In one implementation, the target speech frame is a low-noise and clear speech frame obtained by processing the initial speech frame.
[0148] The speech processing device performs amplitude gain on the amplitude values of each frequency point in the initial spectrum based on the gain factor corresponding to each frequency point, and obtains the target amplitude value corresponding to each frequency point after gain. Then, the target spectrum is obtained based on the target amplitude value corresponding to each frequency point, and the target speech frame is obtained based on the target spectrum.
[0149] For example, assuming the gain factor corresponding to the k-th frequency point in the initial power spectrum is G(k, l), then Y'(k, l) corresponding to the k-th frequency point in the initial power spectrum can be obtained based on the following formula:
[0150] Y'(k,l)=G(k,l)Y'(k,l)
[0151] Where l is the number of the initial speech frame corresponding to the initial spectrum in the speech signal.
[0152] The speech processing device obtains the target spectrum based on Y'(k, l) corresponding to each frequency point. Specifically, it can generate the target spectrum according to each frequency point and the amplitude value after gain at each frequency point, and finally directly process the target spectrum to generate the target speech frame corresponding to the target spectrum. Optionally, when generating the target speech frame based on the target spectrum, it can perform the inverse process relative to generating the initial spectrum based on the initial speech frame, or other speech frame generation methods, which are not limited here.
[0153] In this embodiment, by first identifying the target frequency point in the noise gradually increasing stage of the power spectrum, and then, based on the global speech presence probability of the initial speech frame, the probability of local speech presence corresponding to the target frequency point is specifically corrected, thereby improving the accuracy of the local speech presence probability corresponding to each frequency point. This improves the accuracy of the gain factor calculated based on the local speech presence probability, thereby reducing the residual noise signal and speech distortion in the target speech frame obtained based on the initial speech frame, and improving the clarity of the target speech frame.
[0154] Please see Figure 6 This is a flowchart illustrating a speech processing method provided in an embodiment of this application. Figure 6 As shown, the method may include the following steps.
[0155] S1, collects voice signals.
[0156] The voice processing device continuously collects ambient voice data from the environment in which the voice processing device is located through the voice acquisition device, so as to collect multiple voice frames. The voice frames are arranged in the acquisition order to form a voice signal, and S2 is executed when a voice frame is acquired.
[0157] S2, Obtain the initial audio frame.
[0158] The speech processing device acquires the currently collected speech frame and uses it as the initial speech frame, and then executes S3 and S4 respectively.
[0159] S3, obtain the global probability of speech presence.
[0160] The speech processing device obtains the global speech presence probability corresponding to the initial speech frame through a probability estimation model. Execute S9.
[0161] S4, obtain the initial spectrum.
[0162] The speech processing device performs signal processing procedures such as short-time Fourier transform (STFT) on the initial speech frame to obtain an initial spectrum containing noise. This initial spectrum includes multiple frequency points and the amplitude value of each frequency point. Then, S5 is executed.
[0163] S5, obtain the initial power spectrum.
[0164] The speech processing device obtains the power values of each frequency point in the initial power spectrum based on the amplitude values of each frequency point in the initial spectrum, and generates an initial power spectrum, which includes multiple frequency points and the power values of each frequency point. Execute S6.
[0165] S6, Determine the target frequency.
[0166] The speech processing device acquires the audio phase corresponding to each frequency point, and then uses the frequency point where the audio phase is the noise gradation phase as the target frequency point. Execute S9.
[0167] S7, obtain the probability that the local speech does not exist for each frequency point.
[0168] The speech processing device calculates the probability of the absence of local speech at each frequency point based on the power value and minimum power value corresponding to each frequency point, and then executes S8.
[0169] S8, obtain the probability of local speech presence corresponding to each frequency point.
[0170] The speech processing device obtains the historical noise power value, historical posterior signal-to-noise ratio and historical gain factor corresponding to each frequency point from the storage module. Then, based on the historical noise power value, historical posterior signal-to-noise ratio and historical gain factor, the initial power value and the probability of local speech non-existence corresponding to each frequency point, it calculates the probability of local speech existence corresponding to each frequency point and then executes S9.
[0171] S9, obtain the probability of the existence of the target speech corresponding to the target frequency point.
[0172] The speech processing device corrects the local speech existence probability corresponding to the target frequency point based on the global speech existence probability, obtains the target speech existence probability corresponding to the target frequency point, and executes S10.
[0173] S10, calculate the gain factor corresponding to each frequency point.
[0174] The speech processing device calculates the gain factor corresponding to each frequency point according to the frequency point type. Specifically, it calculates the gain factor corresponding to the target frequency point based on the probability of the target speech presence, the prior signal-to-noise ratio, and the posterior signal-to-noise ratio corresponding to the target frequency point; it calculates the gain factor corresponding to other frequency points based on the probability of the local speech presence, the prior signal-to-noise ratio, and the posterior signal-to-noise ratio corresponding to other frequency points, and executes S11.
[0175] S11, Generate the target spectrum.
[0176] The speech processing device performs gain processing on the initial amplitude value of each frequency point based on the gain factor corresponding to each frequency point to obtain the target amplitude value after gain. Then, based on the target amplitude value corresponding to each frequency point, it generates the target spectrum and executes S12.
[0177] S12, Generate the target speech frame.
[0178] The speech processing device generates a target speech frame based on the target spectrum.
[0179] In this embodiment, by first identifying the target frequency point in the noise gradually increasing stage of the power spectrum, and then, based on the global speech presence probability of the initial speech frame, the probability of local speech presence corresponding to the target frequency point is specifically corrected, thereby improving the accuracy of the local speech presence probability corresponding to each frequency point. This improves the accuracy of the gain factor calculated based on the local speech presence probability, thereby reducing the residual noise signal and speech distortion in the target speech frame obtained based on the initial speech frame, and improving the clarity of the target speech frame.
[0180] The following will be combined with the appendix Figure 7 -Appendix Figure 8 This application provides a detailed description of the voice processing device provided in its embodiments. It should be noted that the appendix... Figure 7 -Appendix Figure 8A voice processing device for executing the present application Figures 1-6 The methods shown in the embodiments are for illustrative purposes only, illustrating the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this application. Figures 1-6 The example shown.
[0181] Please see Figure 7 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application. Figure 7 As shown, the speech processing device 1 in this application embodiment may include: a power spectrum acquisition module 11, a frequency point determination module 12, a probability acquisition module 13, a probability correction module 14, a factor acquisition module 15, and a speech frame generation module 16.
[0182] The power spectrum acquisition module 11 is used to acquire an initial speech frame in a speech signal, acquire the initial spectrum and global speech presence probability corresponding to the initial speech frame, and acquire an initial power spectrum based on the initial spectrum. The initial power spectrum includes multiple frequency points and the power value of each frequency point among the multiple frequency points.
[0183] The frequency point determination module 12 is used to determine each target frequency point that satisfies the noise gradual increase stage among the plurality of frequency points based on the initial power spectrum;
[0184] The probability acquisition module 13 is used to acquire the probability of the presence of local speech corresponding to each frequency point;
[0185] The probability correction module 14 is used to perform probability correction on the local speech existence probability corresponding to each target frequency point based on the global speech existence probability, so as to obtain the target speech existence probability corresponding to each target frequency point.
[0186] The factor acquisition module 15 is used to acquire the gain factor corresponding to each frequency point based on the probability of the existence of target speech corresponding to each target frequency point and the probability of the existence of local speech corresponding to other frequency points, wherein the other frequency points are the frequency points among the multiple frequency points that are not in the noise gradually increasing stage.
[0187] The speech frame generation module 16 is used to perform gain processing on the initial spectrum based on the gain factor corresponding to each frequency point to obtain the target spectrum, and generate the target speech frame corresponding to the initial speech frame based on the target spectrum.
[0188] In this embodiment, by first identifying the target frequency point in the noise gradually increasing stage of the power spectrum, and then, based on the global speech presence probability of the initial speech frame, the probability of local speech presence corresponding to the target frequency point is specifically corrected, thereby improving the accuracy of the local speech presence probability corresponding to each frequency point. This improves the accuracy of the gain factor calculated based on the local speech presence probability, thereby reducing the residual noise signal and speech distortion in the target speech frame obtained based on the initial speech frame, and improving the clarity of the target speech frame.
[0189] In one embodiment, the frequency point determination module 12 is specifically used for:
[0190] A first frequency point is obtained from the initial power spectrum, where the first frequency point is any one of the frequency points;
[0191] If the audio phase corresponding to the first frequency point is a noise gradually increasing phase, then the first frequency point is determined to be the target frequency point that satisfies the noise gradually increasing phase. The audio phase corresponding to the first frequency point is determined by the speech frame position of the historical speech frame to which the historical frequency point with the minimum power value in the historical speech cycle belongs. The historical speech cycle is the previous speech cycle adjacent to the current speech cycle in which the initial speech frame is located. The historical frequency point has the same frequency as the first frequency point.
[0192] In this embodiment, the audio stages corresponding to each frequency point are obtained by storing them in the speech processing device, so as to determine the target frequency point that satisfies the noise gradation stage in the initial speech power spectrum. The audio stages of each frequency point stored in the speech processing device are obtained based on the recognition of the previous historical speech cycle. By periodically updating the audio stages corresponding to each frequency point, the target frequency point that satisfies the noise gradation stage in the current speech cycle is determined, and the new noise gradation stage formed by the sudden noise in the speech signal is identified, thereby improving the accuracy of the audio stages corresponding to each frequency point.
[0193] In another implementation, please refer to Figure 8 This document provides a schematic diagram of the structure of a voice processing device as described in an embodiment of this specification. Figure 8 As shown, the voice processing device 1 described in the embodiments of this specification may further include: a stage determination module 17.
[0194] The stage determination module 17 is specifically used for:
[0195] If the initial speech frame is the last frame of the current speech cycle, then the first frequency of the first frequency point is obtained;
[0196] In the current speech cycle, a first frequency point sequence corresponding to the first frequency is obtained, and all frequency points in the first frequency point sequence are arranged according to the acquisition order of the speech frames in the current speech cycle.
[0197] In the first frequency point sequence, obtain the second frequency point corresponding to the first minimum power value, and obtain the voice frame position of the voice frame to which the second frequency point belongs in the current voice period;
[0198] Based on the voice frame position of the voice frame to which the second frequency point belongs and the set position range in the current voice cycle, the audio stage corresponding to the third frequency point indicated by the first frequency in the target voice cycle is determined, and the target voice cycle is the next voice cycle adjacent to the current voice cycle.
[0199] In this embodiment, by obtaining the second frequency point corresponding to the minimum power value in the current speech cycle, the position of the speech frame corresponding to the second frequency point in the current speech cycle is determined. Then, by judging whether the speech frame corresponding to the minimum power value is at the beginning of the speech cycle, it is determined whether there is a noise gradation phase in the speech cycle. Based on the current judgment result, the frequency point corresponding to the noise gradation phase is determined in the next speech cycle, that is, the frequency point that satisfies the noise gradation phase is dynamically updated. Then, the frequency point that satisfies the noise gradation phase is dynamically probabilistically corrected, thereby improving the recognition accuracy of the noise gradation phase and the probability accuracy of the target frequency point.
[0200] In one implementation, the stage determination module 17 is specifically used for:
[0201] If the position of the speech frame is within a set position range in the current speech cycle, then the audio phase corresponding to the third frequency point corresponding to the first frequency in the target speech cycle is determined to be the noise gradually increasing phase.
[0202] In this embodiment of the application, by determining whether the position of the speech frame is within a set position range, the audio stage of each frequency point in the target speech cycle is obtained. Based on the speech situation of the current speech cycle, the speech existence probability in the next speech cycle is specifically corrected, thereby improving the accuracy of the local speech existence probability corresponding to each target frequency point in the next speech cycle.
[0203] In one embodiment, the probability acquisition module 12 is specifically used for:
[0204] Obtain the second frequency point sequence corresponding to the first frequency of the first frequency point. The second frequency point sequence includes the fourth frequency point corresponding to the first frequency in each historical speech frame of the historical speech cycle and the fifth frequency point corresponding to the first frequency in the speech frames already acquired in the current speech cycle.
[0205] Based on the power values of each frequency point in the second frequency point sequence, obtain the second minimum power value;
[0206] The probability of local speech presence corresponding to the first frequency point is obtained based on the power value of the first frequency point and the second minimum power value.
[0207] In this embodiment of the application, by obtaining the minimum power value in the current speech cycle and the historical speech cycle, the minimum value within the two cycles is obtained, thereby improving the accuracy of the minimum value parameter in the calculation of the probability of local speech presence, and thus improving the accuracy of the probability of local speech presence corresponding to each frequency point in the initial speech frame.
[0208] In one embodiment, the probability acquisition module 12 is specifically used for:
[0209] Based on the power value of the first frequency point and the second minimum power value, the probability that the local speech corresponding to the first frequency point does not exist is obtained;
[0210] Based on the power value of the first frequency point and the probability of the absence of local speech, the probability of the presence of local speech corresponding to the first frequency point is obtained.
[0211] In this embodiment, by sequentially acquiring first frequency points in the initial power spectrum and then performing relevant processing on each acquired first frequency point, the processing of each frequency point is used to process the initial power spectrum, thereby identifying the target frequency point in the initial power spectrum and obtaining the local speech existence probability of the target frequency point. This allows the speech processing device to perform probability correction on the local speech existence probability corresponding to the target frequency point, thereby improving the accuracy of the local speech existence probability corresponding to each frequency point.
[0212] In one embodiment, the factor acquisition module 14 is specifically used for:
[0213] If the first frequency point is the target frequency point, then the gain factor corresponding to the first frequency point is obtained based on the probability of the existence of the target speech, the prior signal-to-noise ratio, and the posterior signal-to-noise ratio corresponding to the first frequency point.
[0214] If the first frequency point is another frequency point, then the gain factor corresponding to the first frequency point is obtained based on the local speech existence probability, prior signal-to-noise ratio and posterior signal-to-noise ratio corresponding to the first frequency point.
[0215] In one implementation, by determining the frequency type of the first frequency point, the calculation parameters corresponding to the gain factor are obtained based on the frequency type of the first frequency point. That is, the probability of the presence of target speech is used for the target frequency point, and the probability of the presence of local speech is used for other frequency points. This improves the accuracy of the gain factor calculated based on the probability of the presence of local speech, thereby reducing the residual noise signal and speech distortion in the target speech frame obtained based on the initial speech frame and improving the clarity of the target speech frame.
[0216] In one embodiment, the power spectrum acquisition module 11 is specifically used for:
[0217] Acquire the initial speech frame from the speech signal;
[0218] A probability estimation model is used to obtain the global speech presence probability of the initial speech frame;
[0219] Perform a Fourier transform on the initial speech frame to obtain the initial spectrum corresponding to the initial speech frame. The initial spectrum includes multiple frequency points and the amplitude values corresponding to the multiple frequency points.
[0220] Based on the initial spectrum, the initial power spectrum corresponding to the initial speech frame is obtained.
[0221] In this embodiment, by inputting the initial speech frame into the trained probability estimation model, the global speech presence probability of the initial speech frame can be directly obtained, reducing the tedious probability estimation process and improving speech processing efficiency.
[0222] This application embodiment also provides a storage medium that can store multiple program instructions, which are adapted to be loaded and executed by a processor as described above. Figures 1-6 The method steps of the illustrated embodiment can be found in the following documentation for detailed execution. Figures 1-6 The specific details of the illustrated embodiments will not be elaborated here.
[0223] Please see Figure 9 This document provides a schematic diagram of the structure of a computer device according to an embodiment of this application. Figure 9As shown, the computer device 1000 may include: at least one processor 1001, at least one communication bus 1002, at least one input / output interface 1003, at least one network interface 1004, and at least one memory 1005. The processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the computer device 1000 using various interfaces and lines, and performs various functions of the terminal 1000 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling data stored in the memory 1005. The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the processor 1001. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The communication bus 1002 is used to implement communication between these components. Figure 9 As shown, the memory 1005, which serves as a storage medium for a terminal device, may include an operating system, a network communication module, an input / output interface module, and a voice processing program.
[0224] exist Figure 9 In the computer device 1000 shown, the input / output interface 1003 is mainly used to provide an input interface for users and access devices to obtain data input by users and access devices.
[0225] In one embodiment.
[0226] The processor 1001 can be used to call the voice processing program stored in the memory 1005 and specifically perform the following operations:
[0227] An initial speech frame is acquired from the speech signal, and the initial spectrum and global speech presence probability corresponding to the initial speech frame are obtained. An initial power spectrum is obtained based on the initial spectrum, and the initial power spectrum includes multiple frequency points and the power value of each frequency point among the multiple frequency points.
[0228] Based on the initial power spectrum, each target frequency point that satisfies the noise gradual increase stage is determined among the multiple frequency points, and the local speech existence probability corresponding to each frequency point is obtained.
[0229] Based on the global speech existence probability, the local speech existence probability corresponding to each target frequency point is corrected to obtain the target speech existence probability corresponding to each target frequency point.
[0230] Based on the probability of the presence of target speech corresponding to each target frequency point and the probability of the presence of local speech corresponding to other frequency points, the gain factor corresponding to each frequency point is obtained, wherein the other frequency points are the frequency points among the multiple frequency points that are not in the noise gradually increasing stage.
[0231] Based on the gain factor corresponding to each frequency point, the initial spectrum is processed to obtain the target spectrum, and the target speech frame corresponding to the initial speech frame is generated based on the target spectrum.
[0232] Optionally, when the processor 1001 performs the step of determining each target frequency point that satisfies the noise gradual increase phase among the plurality of frequency points based on the initial power spectrum, it specifically performs the following operations:
[0233] A first frequency point is obtained from the initial power spectrum, where the first frequency point is any one of the frequency points;
[0234] If the audio phase corresponding to the first frequency point is a noise gradually increasing phase, then the first frequency point is determined to be the target frequency point that satisfies the noise gradually increasing phase. The audio phase corresponding to the first frequency point is determined by the speech frame position of the historical speech frame to which the historical frequency point with the minimum power value in the historical speech cycle belongs. The historical speech cycle is the previous speech cycle adjacent to the current speech cycle in which the initial speech frame is located. The historical frequency point has the same frequency as the first frequency point.
[0235] Optionally, the processor 1001 can also be used to call the voice processing program stored in the memory 1005 and specifically perform the following operations:
[0236] If the initial speech frame is the last frame of the current speech cycle, then the first frequency of the first frequency point is obtained;
[0237] In the current speech cycle, a first frequency point sequence corresponding to the first frequency is obtained, and all frequency points in the first frequency point sequence are arranged according to the acquisition order of the speech frames in the current speech cycle.
[0238] In the first frequency point sequence, obtain the second frequency point corresponding to the first minimum power value, and obtain the voice frame position of the voice frame to which the second frequency point belongs in the current voice period;
[0239] Based on the voice frame position of the voice frame to which the second frequency point belongs and the set position range in the current voice cycle, the audio stage corresponding to the third frequency point indicated by the first frequency in the target voice cycle is determined, and the target voice cycle is the next voice cycle adjacent to the current voice cycle.
[0240] Optionally, when the processor 1001 determines the audio stage corresponding to the third frequency point indicated by the first frequency in the target speech cycle based on the speech frame position of the speech frame to which the second frequency point belongs and the set position range in the current speech cycle, it specifically performs the following operations:
[0241] If the position of the speech frame is within a set position range in the current speech cycle, then the audio phase corresponding to the third frequency point corresponding to the first frequency in the target speech cycle is determined to be the noise gradually increasing phase.
[0242] Optionally, when the processor 1001 performs the step of obtaining the local speech existence probability corresponding to each frequency point, it specifically performs the following operations:
[0243] Obtain the second frequency point sequence corresponding to the first frequency of the first frequency point. The second frequency point sequence includes the fourth frequency point corresponding to the first frequency in each historical speech frame of the historical speech cycle and the fifth frequency point corresponding to the first frequency in the speech frames already acquired in the current speech cycle.
[0244] Based on the power values of each frequency point in the second frequency point sequence, obtain the second minimum power value;
[0245] The probability of local speech presence corresponding to the first frequency point is obtained based on the power value of the first frequency point and the second minimum power value.
[0246] Optionally, when the processor 1001 executes the process of obtaining the local speech presence probability corresponding to the first frequency point based on the power value of the first frequency point and the second minimum power value, it specifically performs the following operations:
[0247] Based on the power value of the first frequency point and the second minimum power value, the probability that the local speech corresponding to the first frequency point does not exist is obtained;
[0248] Based on the power value of the first frequency point and the probability of the absence of local speech, the probability of the presence of local speech corresponding to the first frequency point is obtained.
[0249] Optionally, when the processor 1001 executes the step of obtaining the gain factor corresponding to each frequency point based on the probability of the presence of target speech corresponding to each target frequency point and the probability of the presence of local speech corresponding to other frequency points, it specifically performs the following operations:
[0250] If the first frequency point is the target frequency point, then the gain factor corresponding to the first frequency point is obtained based on the probability of the existence of the target speech corresponding to the first frequency point.
[0251] If the first frequency point is another frequency point, then the gain factor corresponding to the first frequency point is obtained based on the local speech existence probability corresponding to the first frequency point.
[0252] Optionally, when the processor 1001 performs the steps of acquiring an initial speech frame from the speech signal, obtaining the initial spectrum and global speech presence probability corresponding to the initial speech frame, and obtaining an initial power spectrum based on the initial spectrum, wherein the initial power spectrum includes multiple frequency points and the power value of each of the multiple frequency points, the processor 1001 specifically performs the following operations:
[0253] Acquire the initial speech frame from the speech signal;
[0254] A probability estimation model is used to obtain the global speech presence probability of the initial speech frame;
[0255] Perform a Fourier transform on the initial speech frame to obtain the initial spectrum corresponding to the initial speech frame. The initial spectrum includes multiple frequency points and the amplitude values corresponding to the multiple frequency points.
[0256] Based on the initial spectrum, the initial power spectrum corresponding to the initial speech frame is obtained.
[0257] In this embodiment, by first identifying the target frequency point in the noise gradually increasing stage of the power spectrum, and then, based on the global speech presence probability of the initial speech frame, the probability of local speech presence corresponding to the target frequency point is specifically corrected, thereby improving the accuracy of the local speech presence probability corresponding to each frequency point. This improves the accuracy of the gain factor calculated based on the local speech presence probability, thereby reducing the residual noise signal and speech distortion in the target speech frame obtained based on the initial speech frame, and improving the clarity of the target speech frame.
[0258] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0259] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0260] The above is a description of a voice processing method, apparatus, storage medium, and device provided in this application. For those skilled in the art, based on the ideas of the embodiments of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A speech processing method, characterized in that, The method includes: An initial speech frame is acquired from the speech signal, and the initial spectrum and global speech presence probability corresponding to the initial speech frame are obtained. An initial power spectrum is obtained based on the initial spectrum, and the initial power spectrum includes multiple frequency points and the power value of each frequency point among the multiple frequency points. Based on the initial power spectrum, target frequency points that satisfy the noise gradually increasing stage are determined among the multiple frequency points, and the local speech existence probability corresponding to each frequency point is obtained; the noise gradually increasing stage refers to the stage in which the noise energy gradually increases. Based on the correction factor, the global speech existence probability and the local speech existence probability corresponding to each target frequency point are weighted and calculated to obtain the target speech existence probability corresponding to each target frequency point after the probability correction. If the first frequency point is the target frequency point, then based on the probability of the existence of the target speech corresponding to the first frequency point, the gain factor corresponding to the first frequency point is obtained; the first frequency point is any one of the frequency points. If the first frequency point is another frequency point, then the gain factor corresponding to the first frequency point is obtained based on the local speech existence probability corresponding to the first frequency point; the other frequency points are the frequency points that are not in the noise gradually increasing stage among the multiple frequency points. Based on the gain factor corresponding to each frequency point, the initial spectrum is processed to obtain the target spectrum, and the target speech frame corresponding to the initial speech frame is generated based on the target spectrum.
2. The method according to claim 1, characterized in that, The step of determining each target frequency point that satisfies the noise gradual increase phase among the plurality of frequency points based on the initial power spectrum includes: The first frequency point is obtained from the initial power spectrum; If the audio phase corresponding to the first frequency point is a noise gradually increasing phase, then the first frequency point is determined to be the target frequency point that satisfies the noise gradually increasing phase. The audio phase corresponding to the first frequency point is determined by the speech frame position of the historical speech frame to which the historical frequency point with the minimum power value in the historical speech cycle belongs. The historical speech cycle is the previous speech cycle adjacent to the current speech cycle in which the initial speech frame is located. The historical frequency point has the same frequency as the first frequency point.
3. The method according to claim 2, characterized in that, Also includes: If the initial speech frame is the last frame of the current speech cycle, then the first frequency of the first frequency point is obtained; In the current speech cycle, a first frequency point sequence corresponding to the first frequency is obtained, and all frequency points in the first frequency point sequence are arranged according to the acquisition order of the speech frames in the current speech cycle. In the first frequency point sequence, obtain the second frequency point corresponding to the first minimum power value, and obtain the voice frame position of the voice frame to which the second frequency point belongs in the current voice period; Based on the voice frame position of the voice frame to which the second frequency point belongs and the set position range in the current voice cycle, the audio stage corresponding to the third frequency point indicated by the first frequency in the target voice cycle is determined, and the target voice cycle is the next voice cycle adjacent to the current voice cycle.
4. The method according to claim 3, characterized in that, The step of determining the audio phase corresponding to the third frequency point indicated by the first frequency in the target speech cycle based on the speech frame position of the speech frame to which the second frequency point belongs and the set position range in the current speech cycle includes: If the position of the speech frame is within a set position range in the current speech cycle, then the audio phase corresponding to the third frequency point corresponding to the first frequency in the target speech cycle is determined to be the noise gradually increasing phase.
5. The method according to claim 2, characterized in that, The step of obtaining the local speech existence probability corresponding to each frequency point includes: Obtain the second frequency point sequence corresponding to the first frequency of the first frequency point. The second frequency point sequence includes the fourth frequency point corresponding to the first frequency in each historical speech frame of the historical speech cycle and the fifth frequency point corresponding to the first frequency in the speech frames already acquired in the current speech cycle. Based on the power values of each frequency point in the second frequency point sequence, obtain the second minimum power value; The probability of local speech presence corresponding to the first frequency point is obtained based on the power value of the first frequency point and the second minimum power value.
6. The method according to claim 5, characterized in that, The step of obtaining the local speech presence probability corresponding to the first frequency point based on the power value of the first frequency point and the second minimum power value includes: Based on the power value of the first frequency point and the second minimum power value, the probability that the local speech corresponding to the first frequency point does not exist is obtained; Based on the power value of the first frequency point and the probability of the absence of local speech, the probability of the presence of local speech corresponding to the first frequency point is obtained.
7. The method according to claim 1, characterized in that, The process involves acquiring an initial speech frame from the speech signal, obtaining the initial spectrum and global speech presence probability corresponding to the initial speech frame, and obtaining an initial power spectrum based on the initial spectrum. The initial power spectrum includes multiple frequency points and the power value of each of the multiple frequency points, including: Acquire the initial speech frame from the speech signal; A probability estimation model is used to obtain the global speech presence probability of the initial speech frame; Perform a Fourier transform on the initial speech frame to obtain the initial spectrum corresponding to the initial speech frame. The initial spectrum includes multiple frequency points and the amplitude values corresponding to the multiple frequency points. Based on the initial spectrum, the initial power spectrum corresponding to the initial speech frame is obtained.
8. A voice processing device, characterized in that, include: The power spectrum acquisition module is used to acquire an initial speech frame in a speech signal, acquire the initial spectrum and global speech presence probability corresponding to the initial speech frame, and acquire an initial power spectrum based on the initial spectrum. The initial power spectrum includes multiple frequency points and the power value of each frequency point among the multiple frequency points. The frequency point determination module is used to determine each target frequency point that satisfies the noise gradual increase stage among the plurality of frequency points based on the initial power spectrum; The noise gradually increasing stage refers to the stage in which the noise energy gradually increases. The probability acquisition module is used to acquire the probability of local speech presence corresponding to each frequency point; The probability correction module is used to perform a weighted calculation on the global speech existence probability and the local speech existence probability corresponding to each target frequency point based on the correction factor, so as to obtain the target speech existence probability corresponding to each target frequency point after probability correction. The factor acquisition module is used to acquire a gain factor corresponding to the first frequency point based on the probability of the presence of target speech corresponding to the first frequency point if the first frequency point is the target frequency point; the first frequency point is any one of the frequency points; if the first frequency point is another frequency point, the gain factor corresponding to the first frequency point is acquired based on the probability of the presence of local speech corresponding to the first frequency point; the other frequency points are frequency points that are not in the noise gradation stage among the multiple frequency points. The speech frame generation module is used to perform gain processing on the initial spectrum based on the gain factor corresponding to each frequency point to obtain the target spectrum, and generate the target speech frame corresponding to the initial speech frame based on the target spectrum.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech processing method according to any one of claims 1-7.
10. A computer device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the steps of the speech processing method as described in any one of claims 1-7.