Voice processing methods, devices, storage media and electronic devices

By performing energy level estimation, abrupt change detection, and jitter detection on speech frames and adaptively calculating gain, the problem of unstable volume in multi-person online real-time voice scenarios is solved, achieving stability and real-time performance of speech gain and improving user experience.

CN116206619BActive Publication Date: 2026-07-31GUANGZHOU BOGUAN TELECOMM TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU BOGUAN TELECOMM TECH LTD
Filing Date
2023-03-02
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In multi-person online real-time voice scenarios, due to differences in hardware devices and changes in speaker position causing volume instability, traditional automatic gain methods are susceptible to noise interference and have a delayed response, affecting the user experience.

Method used

By performing energy level estimation, abrupt change detection, and jitter detection on real-time acquired speech frames, the initial and final gains are adaptively calculated to achieve both stability and real-time performance of the speech gain.

Benefits of technology

It improves the stability and real-time performance of voice gain, reduces volume differences between different speakers and fluctuates volume for the same speaker, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206619B_ABST
    Figure CN116206619B_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of voice control technology, specifically to voice processing methods, voice processing devices, storage media, and electronic devices. The voice processing method includes: estimating the energy level of a real-time acquired current voice frame to obtain a current energy value, and determining an initial gain for the current voice frame based on the current energy value; performing abrupt change detection on the current voice frame based on the current energy value to obtain abrupt change detection result, and performing jitter detection on the current voice frame based on the initial gain and the abrupt change detection result to obtain a jitter detection result; determining a final gain based on the abrupt change detection result and the jitter detection result, and applying the final gain to the current voice frame. The voice processing method provided by this disclosure can guarantee the stability and real-time performance of the voice gain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of voice control technology, specifically to voice processing methods, voice processing devices, storage media, and electronic devices. Background Technology

[0002] In multi-user online real-time voice scenarios such as gaming sessions and live streaming, two problems exist. First, due to differences in the hardware devices of various users, the volume of different speakers received at the same receiving end may vary significantly. Second, due to changes in the relative position of the speaker and the microphone, the volume of the same speaker may also fluctuate. Simply changing the volume gain at the receiving end cannot resolve these issues, thus impacting the user experience.

[0003] The two scenarios mentioned above face the same problem: after obtaining the speech to be processed, find an appropriate gain to scale it so that the volume fluctuation of the processed speech is reduced and tends to be stable, or in other words, close to the target volume value.

[0004] Therefore, it is necessary to introduce automatic gain control to dynamically adjust the volume. Traditional automatic gain control methods are based on comparing the peak and target values ​​of speech, which are easily affected by noise. At the same time, the gain response time is long, and due to its lag, it tends to amplify large volumes and reduce small volumes when the volume fluctuates.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this disclosure is to provide a speech processing method, speech processing device, storage medium, and electronic device, which aim to solve the problems of improving the stability and real-time performance of speech gain.

[0007] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.

[0008] According to one aspect of the embodiments of this disclosure, a voice processing method is provided, including:

[0009] The energy level of the current speech frame is estimated in real time to obtain the current energy value, and the initial gain of the current speech frame is determined based on the current energy value.

[0010] Based on the current energy value, abrupt change detection is performed on the current speech frame to obtain abrupt change detection result; and based on the initial gain and the abrupt change detection result, jitter detection is performed on the current speech frame to obtain a jitter detection result.

[0011] The final gain is determined based on the mutation detection result and the jitter detection result, and then the final gain is applied to the current speech frame.

[0012] According to a second aspect of the present disclosure, a voice processing apparatus is provided, comprising:

[0013] The estimation module is used to estimate the energy level of the current speech frame acquired in real time to obtain the current energy value, and to determine the initial gain of the current speech frame based on the current energy value.

[0014] The detection module is used to perform abrupt change detection on the current speech frame based on the current energy value to obtain abrupt change detection result, and to perform jitter detection on the current speech frame based on the initial gain and the abrupt change detection result to obtain a jitter detection result;

[0015] A gain module is used to determine a final gain based on the mutation detection result and the jitter detection result, so as to apply the final gain to the current speech frame.

[0016] According to a third aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the speech processing method as described in the above embodiments.

[0017] According to a fourth aspect of the present disclosure, an electronic device is provided, characterized in that it includes: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the voice processing method as described in the above embodiments.

[0018] The exemplary embodiments disclosed herein may have some or all of the following beneficial effects:

[0019] In the technical solutions provided by some embodiments of this disclosure, on the one hand, jitter detection and sudden change detection are specifically added during speech processing to ensure the stability of the gain; on the other hand, energy level estimation is performed after real-time acquisition of speech frames, and the initial gain and final gain are adaptively calculated and applied, thereby improving the real-time performance of speech gain.

[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0022] Figure 1 The illustration shows a flowchart of a speech processing method according to an exemplary embodiment of the present disclosure.

[0023] Figure 2 The schematic diagram illustrates a flowchart of an energy level estimation method according to an exemplary embodiment of the present disclosure;

[0024] Figure 3 The schematic diagram illustrates a flowchart of a VAD detection method according to an exemplary embodiment of the present disclosure;

[0025] Figure 4 The illustration schematically shows a flowchart of a method for determining an initial gain in an exemplary embodiment of the present disclosure;

[0026] Figure 5 The illustration schematically shows a flowchart of a method for determining a current jitter count result in an exemplary embodiment of the present disclosure;

[0027] Figure 6 The illustration schematically shows a flowchart of a method for determining the final gain in an exemplary embodiment of the present disclosure;

[0028] Figure 7 This illustration schematically depicts a speech diagram of an initial input in an exemplary embodiment of the present disclosure;

[0029] Figure 8 A schematic diagram illustrating the result of an existing gain control is shown in the speech diagram.

[0030] Figure 9 A schematic diagram illustrating a gain control result in an exemplary embodiment of the present disclosure is provided.

[0031] Figure 10 This schematic diagram illustrates another initial input voice representation in an exemplary embodiment of the present disclosure.

[0032] Figure 11 This schematic diagram illustrates a speech representation of another existing gain control result. Figure 11 The results shown are those processed using the traditional AGC method;

[0033] Figure 12 A schematic diagram illustrating another gain control result in an exemplary embodiment of the present disclosure is shown in the speech diagram.

[0034] Figure 13 This schematically illustrates a flowchart of controlling voice gain in an exemplary embodiment of the present disclosure;

[0035] Figure 14 This schematic diagram illustrates the composition of a voice processing apparatus according to an exemplary embodiment of the present disclosure;

[0036] Figure 15 This schematic diagram illustrates a computer-readable storage medium according to an exemplary embodiment of the present disclosure;

[0037] Figure 16 The schematic diagram illustrates the structure of a computer system of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0038] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.

[0039] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0040] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0041] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0042] In multi-user online real-time voice scenarios such as gaming sessions and live streaming, two problems exist. First, due to differences in the hardware devices of various users, the volume of different speakers received at the same receiving end may vary significantly. Second, due to changes in the relative position of the speaker and the microphone, the volume of the same speaker may also fluctuate. Simply changing the volume gain at the receiving end cannot resolve these issues, thus impacting the user experience.

[0043] The two scenarios mentioned above face the same problem: after obtaining the speech to be processed, find an appropriate gain to scale it so that the volume fluctuation of the processed speech is reduced and tends to be stable, or in other words, close to the target volume value.

[0044] In existing technologies, some applications address problem one by providing individual volume settings for each user in the voice stream, but users need to adjust them manually; while for problem two, most applications only provide a fixed gain option, which cannot solve the problem of fluctuating volume.

[0045] Therefore, it is necessary to introduce automatic gain control to dynamically adjust the volume. Traditional automatic gain control methods are based on comparing the peak and target values ​​of speech, which are easily affected by noise. At the same time, the gain response time is long, and due to its lag, it tends to amplify large volumes and reduce small volumes when the volume fluctuates.

[0046] Based on this, this disclosure provides a speech processing method with short gain response time, avoidance of noise interference, and stable gain, aiming to solve one or more problems existing in the prior art.

[0047] The implementation details of the technical solutions of the embodiments of this disclosure are described in detail below.

[0048] The voice processing method in one embodiment of this disclosure can be applied in application scenarios such as games, live streaming, or other audio and video processing fields. In specific execution, the voice processing method can run on a local terminal device or a server. When running on a server, the method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and client devices.

[0049] Figure 1 This illustration schematically shows a flowchart of a speech processing method according to an exemplary embodiment of the present disclosure. Figure 1 As shown, the speech processing method includes steps S101 to S103:

[0050] Step S101: Estimate the energy level of the current speech frame acquired in real time to obtain the current energy value, and determine the initial gain of the current speech frame based on the current energy value;

[0051] Step S102: Perform abrupt change detection on the current speech frame based on the current energy value to obtain abrupt change detection result, and perform jitter detection on the current speech frame based on the initial gain to obtain a jitter detection result;

[0052] Step S103: Determine the final gain based on the mutation detection result and the jitter detection result, and apply the final gain to the current speech frame.

[0053] In the technical solutions provided by some embodiments of this disclosure, on the one hand, jitter detection and sudden change detection are specifically added during speech processing to ensure the stability of the gain; on the other hand, energy level estimation is performed after real-time acquisition of speech frames, and the initial gain and final gain are adaptively calculated and applied, thereby improving the real-time performance of speech gain.

[0054] The following will describe in more detail the various steps of the speech processing method in this exemplary embodiment with reference to the accompanying drawings and embodiments.

[0055] In step S101, the energy level of the real-time acquired current speech frame is estimated to obtain the current energy value, and the initial gain of the current speech frame is determined based on the current energy value.

[0056] In one embodiment of this disclosure, when performing voice processing, voice messages are acquired and analyzed frame by frame.

[0057] For a given speech frame, the energy level can first be estimated to obtain the current energy value. Figure 2 This schematic diagram illustrates a flowchart of an energy level estimation method according to an exemplary embodiment of the present disclosure, with reference to... Figure 2 As shown, the energy level estimation method in step S101 specifically includes the following steps:

[0058] Step S201: Determine the current initial state evaluation value, the current steady state evaluation value, and the current long-term estimation value based on the speech segment detection results of the current speech frame;

[0059] Step S202: Select the current initial state evaluation value or the current steady state evaluation value as the current short-time estimate value based on the speech segment detection result of the current speech frame;

[0060] Step S203: Determine the current energy value based on the most recently updated energy value, the current short-term estimate, and the current long-term estimate.

[0061] The following is a detailed explanation of steps S201 and S203.

[0062] In step S201, the current initial state evaluation value, the current steady state evaluation value, and the current long-term estimation value are determined based on the speech segment detection results of the current speech frame.

[0063] When analyzing speech information, short-time estimation can be performed by quickly tracking real-time changes in volume using a short-time estimation window. The result of short-time estimation includes determining the current initial and steady-state evaluation values ​​for each new speech frame. Simultaneously, long-time estimation can be performed by recording the overall average volume level of the speech segment using a long-time estimation window. The result of long-time estimation is the current long-time estimate.

[0064] The voice segment detection result is the result obtained after performing VAD (Voice Activity Detection) detection on the speech frame. In one embodiment of this disclosure, VAD detection can be performed on the current speech frame based on deep learning; in particular, the specific method of VAD detection is not limited here.

[0065] Figure 3 This schematic diagram illustrates a flowchart of a VAD detection method according to an exemplary embodiment of the present disclosure, with reference to... Figure 3 As shown, firstly, the frequency domain amplitude spectrum is calculated using STFT (short-time Fourier transform) on the current frame. A CNN layer is then used for feature compression. Next, an RNN layer is used to learn the VAD (Voice Amplitude Difference) characteristics temporally. Finally, a fully connected layer outputs the VAD probability. The closer the probability is to 1, the higher the probability that the current speech frame contains speech. A probability higher than the VAD threshold is considered a speech frame; otherwise, it is considered a no-speech frame. The VAD threshold can be set as needed.

[0066] Therefore, after obtaining the current speech frame, the update process of step S201 is as follows: obtain the most recently updated initial state evaluation value, steady state evaluation value, and long-term estimation value; update one or more of the initial state evaluation value, steady state evaluation value, and long-term estimation value based on the speech segment detection result of the current speech frame to obtain the current initial state evaluation value, current steady state evaluation value, and current long-term estimation value.

[0067] Among them, the most recently updated initial state evaluation value, steady state evaluation value, and long-term estimate value are the same as the initial state evaluation value, steady state evaluation value, and long-term estimate value recorded in the previous speech frame.

[0068] Updating one or more of the initial state evaluation value, steady state evaluation value, and long-term estimate value can be divided into two parts: first, updating the initial state evaluation value and steady state evaluation value based on the short-term estimate to obtain the current initial state evaluation value and the current steady state evaluation value; second, updating the long-term estimate value to obtain the current long-term estimate value.

[0069] For short-time estimation, the initial state evaluation value e can be maintained. i and steady-state evaluation value e s These two parameters. When acquiring each speech frame, the initial state evaluation value needs to be updated. Simultaneously, the recorded initial state evaluation value and steady-state evaluation value are converted to each other to determine the final initial state evaluation value e. i and steady-state evaluation value e s This information is then used to update subsequent short-term estimation results.

[0070] Therefore, after acquiring each current speech frame, the update strategy is determined based on the speech segment detection results of the current speech frame, and the initial state evaluation value e is adjusted accordingly. i Update first.

[0071] When the speech frame detection result indicates that the current speech frame is a speech frame, the initial state evaluation value e is updated based on the dynamic averaging method. i The dynamic averaging method is shown in formula (1):

[0072] k i+1 =αk i +(1-α)b (1)

[0073] In the formula, k i Let represent the dynamic average value of the i-th time, α be the dynamic average coefficient, and b be the mean square value of the current speech frame.

[0074] When the speech segment detection result indicates that the current speech frame is a no-speech frame, then the initial state evaluation value e i No change for now.

[0075] Then determine the initial state evaluation value e. i and steady-state evaluation value e s Whether a conversion is needed, specifically, when transitioning from a segmented to a segmentless state, and the number of consecutive segments with voices exceeds a threshold, the initial state evaluation value e should be copied. i to steady-state evaluation value e s Otherwise, discard the initial state evaluation value and copy the steady-state evaluation value e when transitioning from no voice segment to a voice segment. s To the initial state evaluation value e i .

[0076] After the update is complete, based on the initial state evaluation value e maintained after the update. i and steady-state evaluation value e s Obtain the current initial state evaluation value e' s and the current steady-state assessment value e' s .

[0077] For long-term estimation, the main thing to maintain is the long-term estimated value e. lOne parameter. Specifically, after acquiring each current speech frame, the update strategy is determined based on the speech segment detection results of the current speech frame.

[0078] When the speech frame detected is determined to be a speech frame, different dynamic averaging coefficients are selected to update e based on the relative magnitudes of the mean square value and the long-term estimate. l When the speech detection result indicates that the current speech frame is a no-speech frame, no updates are performed.

[0079] After the update is complete, the long-term estimate e maintained after the update is completed. l This can be used as the current long-term estimate E' l .

[0080] Step S202: Select the current initial state evaluation value or the current steady state evaluation value as the current short-time estimate value based on the speech segment detection result of the current speech frame.

[0081] The process of determining the current short-time estimate differs depending on the different speech segment detection results for the current speech frame. The specific steps are as follows: When the current speech frame is a speech frame and consecutive speech frames exceed a consecutive threshold, the current initial state evaluation value is selected as the current short-time estimate value; when the current speech frame is a speech frame and consecutive speech frames do not exceed the consecutive threshold, or when the speech segment detection result indicates that the current speech frame is a no-speech frame, the current steady-state evaluation value is selected as the current short-time estimate value.

[0082] Specifically, if the current speech frame is a spoken frame, and the accumulated spoken frames exceed the continuous threshold, it indicates that the speech state is relatively stable. Therefore, the updated initial state evaluation value can represent the energy level of the speech in a short period of time, so the current initial state evaluation value can be selected as the current short-time estimate E'. s The continuous threshold can be a set value.

[0083] However, even if the current speech frame is a spoken frame, the cumulative number of spoken frames has not exceeded the continuous threshold. Therefore, the current initial state evaluation value is inaccurate, and the current steady state evaluation value is needed to represent the energy level of speech in a short period of time.

[0084] If the current speech frame is a silent frame, the current steady-state evaluation value should also be used to represent the energy level of the speech over a short period of time.

[0085] Step S203: Determine the current energy value based on the most recently updated energy value, the current short-term estimate, and the current long-term estimate.

[0086] The most recently updated energy value E is the energy value recorded in the speech information of the previous speech frame. After acquiring the current speech frame, it is calculated based on the current short-time estimate E'. s and the current long-term estimate E' l The current energy value E' corresponding to the current speech frame is obtained by updating the most recently updated energy value E.

[0087] In one embodiment of this disclosure, the process of determining the current energy value in step S203 includes: determining whether the current short-term estimate is valid based on the current long-term estimate and the current short-term estimate; if the current short-term estimate is invalid, using the most recently updated energy value as the current energy value, otherwise using the current short-term estimate as the current energy value.

[0088] Specifically, using the current long-term estimate E' l Determine the current short-time estimate E' s Is it valid if E' s Much smaller than E' l If the misjudgment of a call segment is considered to be due to noise or other reasons, then it is considered to be a short-time estimate E'. s Invalid. The most recently updated energy value E should be avoided. Instead, the original energy value E should be retained as the current energy value E'.

[0089] If E' does not appear s Much smaller than E' l We can assume that the current short-term estimate E' is arbitrarily considered. s If it is valid, then the current short-term estimate E' is... s This is the current energy value E'.

[0090] Figure 4 This schematically illustrates a flowchart of a method for determining an initial gain in an exemplary embodiment of this disclosure, with reference to... Figure 4 As shown, determining the initial gain of the current speech frame in step S101 specifically includes the following steps:

[0091] Step S401: Based on the short-time estimation window, the parameters are updated using the corresponding update strategies when the current speech frame has a voice segment and when it has no voice segment, respectively.

[0092] Step S402: Based on the long-term estimation window, the parameters are updated using the corresponding update strategies when the current speech frame contains a speech segment and when it does not contain a speech segment.

[0093] Step S403, extract the current short-term estimate E' s and the current long-term estimate E' l And determine the current energy value E'.

[0094] After obtaining the current energy value E' in step S101, the initial gain of the current speech frame can be determined based on the current energy value;

[0095] Specifically, determining the initial gain of the current speech frame based on the current energy value includes: obtaining a preset target energy value; and using the difference between the target energy value and the current energy value as the initial gain.

[0096] The target energy value is a preset value that can be configured according to actual needs. The initial gain is the difference between the target energy value and the current energy value.

[0097] The initial gain is calculated based on the target energy value for subsequent gain control. The adaptive gain calculation adjusts speech at different volumes to near the target setting, reduces the volume difference between different speakers, corrects the problem of inconsistent volume for the same speaker, and improves the listening experience for the receiving user.

[0098] In step S102, abrupt change detection is performed on the current speech frame based on the current energy value to obtain abrupt change detection result, and jitter detection is performed on the current speech frame based on the initial gain and the abrupt change detection result to obtain a jitter detection result.

[0099] Specifically, whether the current speech frame experiences energy mutations and jitter will affect the final speech gain, so mutation detection and jitter detection are required for the current speech frame.

[0100] In one embodiment of this disclosure, step S102, which involves detecting a mutation in the current speech frame, specifically includes the following steps: obtaining the mean square value of the current speech frame; when the difference between the current energy value and the mean square value exceeds a mutation threshold, obtaining a mutation detection result indicating that the current speech frame has undergone a mutation; otherwise, obtaining a mutation detection result indicating that the current speech frame has not undergone a mutation.

[0101] Specifically, the mean square value E of the current speech frame rms It can be obtained by feature extraction from the current speech frame. The formula for calculating the mean square value is shown in formula (2):

[0102]

[0103] In the formula, N is the number of sampling points contained in one frame of the current speech frame, and p i This represents the sampled value at the i-th sampling point.

[0104] If the current energy value E' and the mean square value E of the current speech frame are... rms The difference exceeds the mutation threshold T sIf the difference does not exceed the mutation threshold T, an energy mutation is considered to have occurred, and the mutation detection result is recorded as a mutation in the current speech frame. Possible causes include the user modifying the hardware device's built-in volume settings, a significant change in the relative distance between the user and the microphone, or the sound of someone else speaking from a distance. s If the mutation detection result is negative, it is recorded as no mutation has occurred in the current speech frame.

[0105] In one embodiment of this disclosure, step S102, which involves detecting jitter in the current speech frame, specifically includes the following steps: determining the current jitter count result based on the speech segment detection result, the abrupt change detection result, and the initial gain of the current speech frame; when the current jitter count result exceeds the jitter threshold, obtaining the jitter detection result that the current speech frame is jittered, otherwise obtaining the jitter detection result that the current speech frame is not jittered.

[0106] Specifically, it is necessary to determine whether jitter has occurred in the current speech frame based on the jitter count result. Therefore, it is necessary to determine the current jitter count result based on the speech segment detection result, the sudden change detection result, and the initial gain of the current speech frame.

[0107] In one embodiment of this disclosure, the specific process of determining the current jitter count result includes:

[0108] When the current voice frame is a spoken frame and there is no sudden change in the current voice frame, the current jitter count result is calculated based on the initial gain.

[0109] Specifically, the initial gain of the most recent preset number of talk frames is accumulated to obtain the current jitter count result. Taking the calculation of the initial gain of the most recent 100 talk frames as an example, the calculation formula for the current jitter count result is shown in formula (3).

[0110]

[0111] In the formula, M represents a voice frame, and g i The initial gain is the value corresponding to each voice frame in frame M.

[0112] When the current voice frame is a spoken frame and the current voice frame undergoes a sudden change, the initial gain is recorded and the current jitter count result is configured to 0;

[0113] Specifically, at this point, the jitter counter is reset, the current jitter count is recorded as 0, and the process resumes after the sudden change ends. As mentioned above, it is necessary to accumulate the initial gain of the spoken frames when necessary, so the initial gain of the current speech frame can also be recorded.

[0114] When the current speech frame is a silent frame, the most recently updated jitter count result is used as the current jitter count result. The most recently updated jitter count result is the jitter count result corresponding to the previous speech frame.

[0115] Since the current voice frame is a silent frame, there is no need for jitter detection, and the original jitter count results can be retained.

[0116] Then, based on the current jitter count, it is determined whether the current speech frame has jittered. Specifically: when the current jitter count exceeds the jitter threshold, the jitter detection result is that the current speech frame has jittered; otherwise, the jitter detection result is that the current speech frame has not jittered.

[0117] Specifically, if C' exceeds the jitter threshold S s If C' does not exceed the jitter threshold S, then the current speech frame is considered to be free of jitter; s If the fluctuation is normal, it is considered a normal tone fluctuation, and the jitter detection result is that the current speech frame is jittering, so no gain change is needed subsequently.

[0118] Figure 5 This schematically illustrates a flowchart of a method for determining a current jitter count result in an exemplary embodiment of this disclosure, with reference to... Figure 5 As shown, the method for determining the current jitter count result specifically includes the following steps:

[0119] Step S501: Determine whether the current voice frame contains a speech segment based on the VAD detection result. If it does not contain a speech segment, proceed directly to step S504 and return the jitter count result. The jitter count result returned at this time is the most recently updated jitter count result, which is the jitter count result of the previous voice frame.

[0120] If the current voice frame contains a speech segment, then proceed to step S502 to determine if the current voice frame has a sudden change. If no sudden change occurs, proceed to step S503 to update the jitter counter by +1, and then proceed to step S505 to return the jitter count result. If a sudden change occurs, proceed to step S504 to reset the jitter counter to zero, and then proceed to step S505 to return the jitter count result.

[0121] Finally, step S506 is executed to obtain the jitter detection result based on the jitter counter.

[0122] Step S103: Determine the final gain based on the mutation detection result and the jitter detection result, and apply the final gain to the current speech frame.

[0123] In one embodiment of this disclosure, the specific process for determining the final gain in step S103 is as follows:

[0124] When a sudden change occurs in the current speech frame, the final gain is calculated based on the mean square value of the current speech frame.

[0125] Specifically, if a sudden change occurs in the current speech frame, the gain is recalculated based on the mean square value of the current speech frame, i.e., target energy - current mean square value = gain, thereby avoiding the use of short-time estimates with lag when a sudden change occurs, which would lead to popping sounds.

[0126] When there are no sudden changes and no jitter in the current speech frame, the final gain is calculated based on the initial gain and the most recently updated gain.

[0127] Specifically, calculating the final gain based on the initial gain and the most recently updated gain includes: calculating an initial change based on the initial gain and the most recently updated gain; correcting the initial change according to the magnitude of the gain change to obtain a target change; and using the sum of the most recently updated gain and the target change as the final gain.

[0128] First, the difference between the initial gain and the most recently updated gain can be used as the initial change value, where the most recently updated gain is the gain of the previous frame.

[0129] Then, the initial change is corrected according to the gain change amplitude to ensure that the change does not exceed the gain change amplitude. The gain change amplitude can also be configured as needed. If the initial change does not exceed the gain change amplitude, then the initial change is used as the target change. If it exceeds the gain change amplitude, then the initial change is trimmed according to the gain change amplitude to obtain the target change.

[0130] Finally, the most recently updated gain and the target change are added together to obtain the final gain of the current speech frame.

[0131] The following example, using the initial gain of the current speech frame calculated as 30, the most recently updated gain as 20, and the preset gain change amplitude as 5, will be used to explain in detail the calculation of the final gain.

[0132] First, the initial change is calculated as 30-20=10. Since 10>5, the change is truncated from 10 to 5. Finally, the final gain is calculated as 20+5=25.

[0133] Based on the above method, under non-abrupt conditions, the maximum gain change value per frame is set to the gain change amplitude, which can ensure that the gain change between adjacent frames is as smooth as possible, thereby improving the user's listening experience. For example, if the desired gain of the current audio frame is 100, but the maximum allowable gain change value per frame is only 5, then without clipping, the gain value curve might be 0, 100, 100, while with clipping, the gain value curve would be 0, 5, 10, 15.

[0134] When the current speech frame does not experience a sudden change but jitter occurs, the most recently updated gain is used as the final gain.

[0135] If the current speech frame does not experience a sudden change but jitter occurs, it is considered a normal fluctuation in tone, and no gain change is made; the most recently updated gain is retained.

[0136] Figure 6 This schematically illustrates a flowchart of a method for determining the final gain in an exemplary embodiment of this disclosure, with reference to... Figure 6 As shown, the method for determining the final gain specifically includes the following steps:

[0137] Step S601: Determine whether the current speech frame has a sudden change. If a sudden change occurs, proceed to step S604 to recalculate the final gain based on the mean square value. If no sudden change occurs, proceed to step S602 to calculate the initial gain change value.

[0138] Then, step S603 is executed to determine whether the current speech frame is jittery. If no jitter occurs, step S605 is executed to trim the initial gain change value. If jitter occurs, step S606 is executed to set the initial gain change value to 0.

[0139] Finally, step S607 is executed to calculate the final gain based on the changed target gain value.

[0140] In step S103, after obtaining the final gain, the final gain is applied to the current speech frame. Specifically, applying the final gain to the current speech frame includes: acquiring all sampling points of the current speech frame; and performing linear interpolation calculation on the final gain to apply it to each sampling point.

[0141] Specifically, a speech frame includes multiple sampling points. After the final gain interpolation calculation is performed, it is extended to each sampling point, which can make the gain change smoothly between each sampling point and ensure the stability of the speech gain.

[0142] Based on the above method, the advantage of adding abrupt change detection to the calculation of the final gain is that when energy changes abruptly, traditional methods will produce a more obvious popping sound due to their lag, while abrupt change detection can effectively avoid this phenomenon and ensure the stability of speech.

[0143] Figure 7 This illustration schematically depicts a speech diagram of an initial input in an exemplary embodiment of the present disclosure. Figure 8 This schematic diagram illustrates a speech representation of an existing gain control result. Figure 8 The results shown are those processed by the traditional AGC method, such as... Figure 8 As shown, a noticeable popping sound appeared at the site of the abrupt change. Figure 9 This schematic diagram illustrates a speech representation of a gain control result in an exemplary embodiment of the present disclosure, such as... Figure 9 As shown, after adding mutation protection, the amplitude and duration of the popping sound were significantly reduced.

[0144] The advantage of adding jitter detection to the final gain calculation is that traditional automatic gain algorithms without jitter detection depend on the length of the evaluation window, resulting in too much evaluation lag. The gain of each frame is affected by the previous frame, which amplifies the local energy differences. The speech processing method in this application, after adding jitter detection, maintains its original normal fluctuations, and the same gain is applied to each frame.

[0145] For example, suppose that due to differences in normal tone and pronunciation, there are five adjacent frames with an energy of 13213 and a target energy of 5. After adding jitter detection, the original normal fluctuation will be maintained, and the result after gain will be 57657, with the same gain applied to each frame. However, the traditional automatic gain algorithm without jitter detection may produce 17447, with too much evaluation lag, which amplifies the local energy differences.

[0146] Figure 10 This schematically illustrates a speech diagram of another initial input in an exemplary embodiment of the present disclosure. Figure 11 This schematic diagram illustrates a speech representation of another existing gain control result. Figure 11 The results shown are those processed by the traditional AGC method, such as... Figure 11 As shown, in many places, values ​​that were originally large have been excessively amplified. Figure 12 A schematic diagram illustrating another gain control result in an exemplary embodiment of this disclosure is shown, such as... Figure 12 As shown, after adding debouncing, the envelope of the normal fluctuations of the original input is preserved to the greatest extent, and the overall size is reduced.

[0147] In one embodiment of this disclosure, after determining the final gain based on the mutation detection result and the jitter detection result, the method further includes amplitude limiting the final gain, including: when the final gain is in a first signal amplitude range, configuring a linear clipping factor based on the final gain and a clipping threshold, and updating the final gain based on the clipping threshold and the linear clipping factor; when the final gain is in a second signal amplitude range, modifying the final gain to the maximum amplitude value, so as to apply the maximum amplitude value to the current speech frame.

[0148] Specifically, the portion of the amplified signal that exceeds the legal range is clipped to further prevent popping noises. Since clipping may introduce high-frequency harmonics, soft clipping is used to minimize the impact of high-frequency harmonics introduced by clipping.

[0149] The specific idea of ​​soft clipping is as follows: x is the signal amplitude before clipping. According to the magnitude of x, it is divided into three regions, as shown in formula (4):

[0150]

[0151] First, in the constant region, for amplitudes that do not reach the clipping threshold k1, the output is not changed;

[0152] Second, the soft clipping region selects the clipping coefficient based on the portion of the input amplitude that exceeds the clipping threshold k1. The example here uses linear soft clipping, where α is the linear clipping factor. A dynamic clipping factor can also be used.

[0153] Thirdly, in the strong restriction region, for the portion of the input amplitude that, even with soft clipping, still exceeds the legal maximum amplitude, the maximum output amplitude is forced. Therefore, k² + α(x - k²) = k max .

[0154] Figure 13 This schematically illustrates a flowchart of controlling voice gain according to an exemplary embodiment of the present disclosure. (Reference) Figure 13 As shown, the steps for controlling speech gain are as follows:

[0155] Step S1301: Perform feature extraction on the current speech frame to determine the mean square value;

[0156] Step S1302: Perform VAD detection on the current speech frame;

[0157] Step S1303: Estimate the energy level of the current speech frame;

[0158] Step S1304: Calculate the initial gain of the current speech frame;

[0159] Step S1305: Perform mutation detection;

[0160] Step S1306: Perform jitter detection;

[0161] Step S1307: Calculate the final gain;

[0162] Step S1308: Limit the amplitude.

[0163] The processes of steps S1301 to S1308 described above have been detailed in previous speech processing methods, so they will not be repeated here.

[0164] In summary, the speech processing method provided in this application can adaptively calculate the final gain of speech frames to adjust speech at different volumes to near the target set value, reduce the range of volume differences between different speakers, correct the problem of fluctuating volume for the same speaker, and improve the listening experience for the receiving user. Simultaneously, targeted jitter detection and abrupt change detection functions are incorporated to ensure the stability and real-time performance of the gain.

[0165] Figure 14 This schematic diagram illustrates the composition of a voice processing apparatus according to an exemplary embodiment of the present disclosure, such as... Figure 14 As shown, the speech processing device 1400 may include an estimation module 1401, a detection module 1402, and a gain module 1403. Wherein:

[0166] The estimation module 1401 is used to estimate the energy level of the current speech frame acquired in real time to obtain the current energy value, and to determine the initial gain of the current speech frame based on the current energy value.

[0167] The detection module 1402 is used to perform abrupt change detection on the current speech frame based on the current energy value to obtain abrupt change detection result, and to perform jitter detection on the current speech frame based on the initial gain and the abrupt change detection result to obtain a jitter detection result;

[0168] Gain module 1403 is used to determine a final gain based on the mutation detection result and the jitter detection result, so as to apply the final gain to the current speech frame.

[0169] According to an exemplary embodiment of this disclosure, the estimation module 1401 is further configured to determine a current initial state evaluation value, a current steady state evaluation value, and a current long-term estimation value based on the speech segment detection result of the current speech frame; select the current initial state evaluation value or the current steady state evaluation value as the current short-term estimation value based on the speech segment detection result of the current speech frame; and determine the current energy value based on the most recently updated energy value, the current short-term estimation value, and the current long-term estimation value.

[0170] According to an exemplary embodiment of the present disclosure, the estimation module 1401 is further configured to obtain the most recently updated initial state evaluation value, steady state evaluation value, and long-term estimation value; and update one or more of the initial state evaluation value, steady state evaluation value, and long-term estimation value based on the speech segment detection result of the current speech frame to obtain the current initial state evaluation value, current steady state evaluation value, and current long-term estimation value.

[0171] According to an exemplary embodiment of this disclosure, the estimation module 1401 is further configured to select the current initial state evaluation value as the current short-time estimation value when the current voice frame is a voice frame and consecutive voice frames exceed a consecutive threshold; and to select the current steady-state evaluation value as the current short-time estimation value when the current voice frame is a voice frame and consecutive voice frames do not exceed the consecutive threshold, or when the voice segment detection result indicates that the current voice frame is a no-voice frame.

[0172] According to an exemplary embodiment of the present disclosure, the estimation module 1401 is further configured to determine whether the current short-term estimate is valid based on the current long-term estimate and the current short-term estimate; if the current short-term estimate is invalid, the most recently updated energy value is used as the current energy value, otherwise the current short-term estimate is used as the current energy value.

[0173] According to an exemplary embodiment of this disclosure, the estimation module 1401 is further configured to obtain a preset target energy value; and use the difference between the target energy value and the current energy value as the initial gain.

[0174] According to an exemplary embodiment of this disclosure, the detection module 1402 is further configured to obtain the mean square value of the current speech frame; when the difference between the current energy value and the mean square value exceeds the mutation threshold, a mutation detection result is obtained indicating that the current speech frame has undergone a mutation; otherwise, a mutation detection result is obtained indicating that the current speech frame has not undergone a mutation.

[0175] According to an exemplary embodiment of this disclosure, the detection module 1402 is further configured to determine a current jitter count result based on the speech segment detection result, the abrupt change detection result, and the initial gain of the current speech frame; when the current jitter count result exceeds the jitter threshold, a jitter detection result is obtained that the current speech frame is jittered, otherwise a jitter detection result is obtained that the current speech frame is not jittered.

[0176] According to an exemplary embodiment of this disclosure, the detection module 1402 is further configured to: calculate the current jitter count result based on the initial gain when the current voice frame is a voice frame and the current voice frame does not undergo a sudden change; record the initial gain and configure the current jitter count result to 0 when the current voice frame is a voice frame and the current voice frame undergoes a sudden change; and use the most recently updated jitter count result as the current jitter count result when the current voice frame is a voiceless frame.

[0177] According to an exemplary embodiment of the present disclosure, the gain module 1403 is further configured to calculate the final gain based on the mean square value of the current speech frame when the current speech frame undergoes a sudden change; calculate the final gain based on the initial gain and the most recently updated gain when the current speech frame does not undergo a sudden change and does not experience jitter; and use the most recently updated gain as the final gain when the current speech frame does not undergo a sudden change but experiences jitter.

[0178] According to an exemplary embodiment of this disclosure, the gain module 1403 is further configured to calculate an initial change based on the initial gain and the most recently updated gain; correct the initial change according to the gain change magnitude to obtain a target change; and use the sum of the most recently updated gain and the target change as the final gain.

[0179] According to an exemplary embodiment of this disclosure, the gain module 1403 is further configured to acquire all sampling points of the current speech frame; and to perform linear interpolation calculation on the final gain to apply it to each of the sampling points.

[0180] According to an exemplary embodiment of the present disclosure, the gain module 1403 is further configured to configure a linear clipping factor based on the final gain and a clipping threshold when the final gain is in a first signal amplitude range, and update the final gain based on the clipping threshold and the linear clipping factor; and to modify the final gain to the maximum amplitude value when the final gain is in a second signal amplitude range, so as to apply the maximum amplitude value to the current speech frame.

[0181] The specific details of each module in the aforementioned speech processing device 1400 have been described in detail in the corresponding speech processing methods, so they will not be repeated here.

[0182] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0183] In an exemplary embodiment of this disclosure, a storage medium capable of implementing the above-described method is also provided. Figure 15 This schematic diagram illustrates a computer-readable storage medium according to an exemplary embodiment of the present disclosure, such as... Figure 15As shown, a program product 1500 for implementing the above-described method according to an embodiment of the present disclosure is described. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a mobile phone. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0184] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided. Figure 16 The schematic diagram illustrates the structure of a computer system of an electronic device according to an exemplary embodiment of the present disclosure.

[0185] It should be noted that, Figure 16 The computer system 1600 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0186] like Figure 16 As shown, the computer system 1600 includes a Central Processing Unit (CPU) 1601, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1602 or programs loaded from storage section 1608 into Random Access Memory (RAM) 1603. The RAM 1603 also stores various programs and data required for system operation. The CPU 1601, ROM 1602, and RAM 1603 are interconnected via a bus 1604. An Input / Output (I / O) interface 1605 is also connected to the bus 1604.

[0187] The following components are connected to I / O interface 1605: an input section 1606 including a keyboard, mouse, etc.; an output section 1607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1608 including a hard disk, etc.; and a communication section 1609 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1609 performs communication processing via a network such as the Internet. A drive 1610 is also connected to I / O interface 1605 as needed. Removable media 1611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1610 as needed so that computer programs read from them can be installed into storage section 1608 as needed.

[0188] In particular, according to embodiments of this disclosure, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1609, and / or installed from removable medium 1611. When the computer program is executed by central processing unit (CPU) 1601, it performs various functions defined in the system of this disclosure.

[0189] It should be noted that the computer-readable medium shown in the embodiments of this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0190] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0191] The units described in the embodiments of this disclosure can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the unit itself.

[0192] In another aspect, this disclosure also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0193] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0194] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this disclosure.

[0195] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0196] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A voice processing method, characterized by, include: The energy level of the current speech frame is estimated in real time to obtain the current energy value, and the initial gain of the current speech frame is determined based on the current energy value. Based on the current energy value, a mutation detection result is obtained by performing mutation detection on the current speech frame; when the current speech frame is a speech frame and no mutation occurs, the current jitter count result is calculated based on the initial gain; when the current speech frame is a speech frame and a mutation occurs, the initial gain is recorded and the current jitter count result is configured to 0; when the current speech frame is a no-speech frame, the most recently updated jitter count result is used as the current jitter count result. And based on the current jitter count, a jitter detection result is obtained; The final gain is determined based on the mutation detection result and the jitter detection result, and then the final gain is applied to the current speech frame.

2. The voice processing method of claim 1, wherein, The step of estimating the energy level of the current speech frame to obtain the current energy value includes: Based on the speech segment detection results of the current speech frame, determine the current initial state evaluation value, the current steady state evaluation value, and the current long-term estimate value; The current initial state evaluation value or the current steady state evaluation value is selected as the current short-time estimate value based on the speech segment detection result of the current speech frame; The current energy value is determined based on the most recently updated energy value, the current short-term estimate, and the current long-term estimate.

3. The speech processing method according to claim 2, characterized in that, The determination of the current initial state evaluation value, the current steady state evaluation value, and the current long-term estimate value based on the speech segment detection results of the current speech frame includes: Obtain the most recently updated initial state assessment, steady state assessment, and long-term estimate; Based on the speech segment detection results of the current speech frame, update one or more of the initial state evaluation value, steady state evaluation value, and long-term estimate value to obtain the current initial state evaluation value, current steady state evaluation value, and current long-term estimate value.

4. The speech processing method according to claim 2, characterized in that, The step of selecting the current initial state evaluation value or the current steady state evaluation value as the current short-time estimate value based on the speech segment detection result of the current speech frame includes: When the current voice frame is a talk frame and the number of consecutive talk frames exceeds the consecutive threshold, the current initial state evaluation value is selected as the current short-term estimate value. When the current voice frame is a voice frame and the number of consecutive voice frames does not exceed the consecutive threshold, or when the voice segment detection result indicates that the current voice frame is a no-voice frame, the current steady-state evaluation value is selected as the current short-term estimate value.

5. The speech processing method according to claim 2, characterized in that, Determining the current energy value based on the most recently updated energy value, the current short-term estimate, and the current long-term estimate includes: Determine whether the current short-term estimate is valid based on the current long-term estimate and the current short-term estimate; If the current short-term estimate is invalid, the most recently updated energy value is used as the current energy value; otherwise, the current short-term estimate is used as the current energy value.

6. The speech processing method according to claim 1, characterized in that, Determining the initial gain of the current speech frame based on the current energy value includes: Obtain the preset target energy value; The difference between the target energy value and the current energy value is used as the initial gain.

7. The speech processing method according to claim 1, characterized in that, The step of performing abrupt change detection on the current speech frame based on the current energy value to obtain abrupt change detection result includes: Obtain the mean square value of the current speech frame; When the difference between the current energy value and the mean square value exceeds the mutation threshold, the mutation detection result is that the current speech frame has undergone a mutation; otherwise, the mutation detection result is that the current speech frame has not undergone a mutation.

8. The speech processing method according to claim 1, characterized in that, The process of obtaining the jitter detection result based on the current jitter count includes: When the current jitter count exceeds the jitter threshold, the jitter detection result is that the current speech frame is jittered; otherwise, the jitter detection result is that the current speech frame is not jittered.

9. The speech processing method according to claim 1, characterized in that, The step of determining the final gain based on the mutation detection result and the jitter detection result includes: When a sudden change occurs in the current speech frame, the final gain is calculated based on the mean square value of the current speech frame; When there is no abrupt change and no jitter in the current speech frame, the final gain is calculated based on the initial gain and the most recently updated gain; When the current speech frame does not experience a sudden change but jitter occurs, the most recently updated gain is used as the final gain.

10. The speech processing method according to claim 9, characterized in that, The calculation of the final gain based on the initial gain and the most recently updated gain includes: The initial change is calculated based on the initial gain and the most recently updated gain; The initial change is corrected based on the magnitude of the gain change to obtain the target change. The sum of the most recently updated gain and the target change is taken as the final gain.

11. The speech processing method according to claim 1, characterized in that, Applying the final gain to the current speech frame includes: Obtain all sampling points of the current speech frame; The final gain is calculated using linear interpolation and applied to each of the sampling points.

12. The speech processing method according to claim 1, characterized in that, After determining the final gain based on the mutation detection result and the jitter detection result, the method further includes: When the final gain is within the first signal amplitude range, a linear clipping factor is configured based on the final gain and the clipping threshold, and the final gain is updated based on the clipping threshold and the linear clipping factor. When the final gain is within the second signal amplitude range, the final gain is modified to the maximum amplitude value so that the maximum amplitude value is applied to the current speech frame.

13. A voice processing device, characterized in that, include: The estimation module is used to estimate the energy level of the current speech frame acquired in real time to obtain the current energy value, and to determine the initial gain of the current speech frame based on the current energy value. The detection module is used to perform abrupt change detection on the current speech frame based on the current energy value to obtain abrupt change detection result; when the current speech frame is a speech frame and no abrupt change occurs in the current speech frame, it calculates the current jitter count result based on the initial gain; when the current speech frame is a speech frame and abrupt change occurs in the current speech frame, it records the initial gain and configures the current jitter count result to 0; when the current speech frame is a no-speech frame, it uses the most recently updated jitter count result as the current jitter count result. And based on the current jitter count, a jitter detection result is obtained; A gain module is used to determine a final gain based on the mutation detection result and the jitter detection result, so as to apply the final gain to the current speech frame.

14. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the speech processing method as described in any one of claims 1 to 12.

15. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the speech processing method as described in any one of claims 1 to 12.