Voice gain control method, device, voice control equipment and storage medium

By using volume parameters from historical statistical periods to determine the target gain level in automatic gain control of voice, the problem of volume instability caused by inaccurate single gain calculation is solved, achieving stable volume control of voice signals and improving user experience and call quality.

CN116721671BActive Publication Date: 2026-05-26MAIPU COMM TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MAIPU COMM TECH CO LTD
Filing Date
2023-07-25
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing automatic gain control methods for speech signals fail to provide stable volume due to inaccurate single-step gain calculations, resulting in the failure of speech signal volume equalization.

Method used

The target gain level for the current statistical period is determined based on the volume parameters of multiple frames of speech signals within the historical statistical period, and the volume parameters of each frame of speech signal are adjusted according to the target gain level to achieve smooth adjustment.

Benefits of technology

It ensures that the volume of the voice signal remains stable over multiple cycles, improving the user's listening experience, especially in multi-person voice calls and voice conferences, keeping the volume of all speakers consistent and protecting the listener's hearing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116721671B_ABST
    Figure CN116721671B_ABST
Patent Text Reader

Abstract

This application provides a voice gain control method, apparatus, voice control device, and storage medium, relating to the field of voice signal processing technology. The method includes: determining a target gain level for the current statistical period based on a first volume parameter of multiple frames of voice signals within a historical statistical period; and adjusting the second volume parameter of each frame of voice signal within the current statistical period based on a second volume parameter and the target gain level. This application can improve the stability of voice signal volume gain control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech signal processing technology, and more specifically, to a speech gain control method, apparatus, speech control device, and storage medium. Background Technology

[0002] With the increasing prevalence of real-time voice and video calls, audio and video technologies are receiving more and more attention.

[0003] Automatic Gain Control (AGC) is an important part of voice signal processing in audio and video technology. It is used to deal with the problem of fluctuating volume of voice signals, so as to make the volume of voice signals relatively stable and improve the user's listening experience.

[0004] Existing automatic gain control methods for speech amplify speech signals through single-step gain calculations. However, inaccurate single-step gain calculations can lead to the failure of volume equalization in speech signals, resulting in the inability to provide stable speech signals. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of the prior art by providing a voice gain control method, apparatus, voice control device, and storage medium to improve the stability of voice signal volume gain control.

[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0007] In a first aspect, embodiments of this application provide a voice gain control method, the method comprising:

[0008] The target gain level for the current statistical period is determined based on the first volume parameter of multiple frames of speech signals within the historical statistical period.

[0009] The second volume parameter of each frame of speech signal is adjusted based on the second volume parameter of each frame of speech signal in the current statistical period and the target gain level.

[0010] Optionally, determining the target gain level for the current statistical period based on the first volume parameter of multiple frames of speech signals within a historical statistical period includes:

[0011] Determine whether the first volume parameter meets the preset gain level adjustment conditions;

[0012] If the first volume parameter satisfies the preset gain level adjustment condition, then the initial gain level corresponding to the current statistical period is adjusted according to the gain adjustment method corresponding to the preset gain level adjustment condition to obtain the target gain level.

[0013] Optionally, determining whether the first volume parameter meets the preset gain level adjustment conditions includes:

[0014] Determine whether the first volume parameter of the multi-frame speech signal corresponding to the previous historical statistical period of the current statistical period is less than the first minimum volume threshold.

[0015] If the first volume parameter satisfies the preset gain level adjustment condition, then according to the gain adjustment method corresponding to the preset gain level adjustment condition, the initial gain level corresponding to the current statistical period is adjusted to obtain the target gain level, including:

[0016] If the first volume parameter of the previous historical statistical period is less than the first minimum volume threshold, then the initial gain level is increased according to the first gain adjustment method corresponding to the first minimum volume threshold to obtain the target gain level.

[0017] Optionally, determining whether the first volume parameter meets the preset gain level adjustment conditions includes:

[0018] Determine whether the first volume parameter of the corresponding multi-frame speech signal within multiple consecutive historical statistical periods is within the optimal volume threshold range;

[0019] If the first volume parameter satisfies the preset gain level adjustment condition, then according to the gain adjustment method corresponding to the preset gain level adjustment condition, the initial gain level corresponding to the current statistical period is adjusted to obtain the target gain level, including:

[0020] If the first volume parameter is not within the optimal volume threshold range, then the initial gain level is adjusted according to the gain adjustment method corresponding to the optimal volume threshold range to obtain the target gain level;

[0021] Wherein, if the first volume parameter is less than the minimum value of the optimal volume threshold range, the initial gain level is increased according to the second gain adjustment method corresponding to the minimum value of the optimal volume threshold range to obtain the target gain level; or...

[0022] If the first volume parameter is greater than the maximum value of the optimal volume threshold range, the initial gain level is lowered according to the third gain adjustment method corresponding to the maximum value of the optimal volume threshold range to obtain the target gain level.

[0023] Optionally, adjusting the volume parameter of each frame of speech signal based on the second volume parameter of each frame of speech signal within the current statistical period and the target gain level includes:

[0024] If the second volume parameter is greater than or equal to the second minimum volume threshold, the second volume parameter is adjusted according to the second volume parameter and the target gain level.

[0025] Optionally, the method further includes:

[0026] If the second volume parameter is less than the second minimum volume threshold, the corresponding voice signal is determined to be noise.

[0027] Optionally, each historical statistical period includes a preset number of voice signals, wherein the first volume parameter of the preset number of voice signals is greater than or equal to the second minimum volume threshold.

[0028] Secondly, embodiments of this application provide a voice gain control device, the device comprising:

[0029] The gain level determination module is used to determine the target gain level for the current statistical period based on the first volume parameter of multiple frames of speech signals within the historical statistical period.

[0030] The volume parameter adjustment module is used to adjust the second volume parameter of each frame of speech signal according to the second volume parameter of each frame of speech signal in the current statistical period and the target gain level.

[0031] Optionally, the gain level determination module includes:

[0032] The adjustment condition judgment unit is used to determine whether the first volume parameter meets the preset gain level adjustment condition;

[0033] The gain level determination unit is used to adjust the initial gain level corresponding to the current statistical period according to the gain adjustment method corresponding to the preset gain level adjustment condition if the first volume parameter meets the preset gain level adjustment condition, so as to obtain the target gain level.

[0034] Optionally, the adjustment condition judgment unit is specifically used to determine whether the first volume parameter of the multi-frame voice signal corresponding to the previous historical statistical period of the current statistical period is less than the first minimum volume threshold.

[0035] The gain level determination unit is specifically used to increase the initial gain level according to the first gain adjustment method corresponding to the first minimum volume threshold if the first volume parameter of the previous historical statistical period is less than the first minimum volume threshold, so as to obtain the target gain level.

[0036] Optionally, the adjustment condition judgment unit is specifically used to determine whether the first volume parameter of the corresponding multi-frame speech signal within multiple consecutive historical statistical periods is within the optimal volume threshold range.

[0037] The gain level determination unit is specifically used to adjust the initial gain level according to the gain adjustment method corresponding to the optimal volume threshold range if the first volume parameter is not within the optimal volume threshold range, so as to obtain the target gain level.

[0038] Specifically, the gain level determination unit is used to: if the first volume parameter is less than the minimum value of the optimal volume threshold range, increase the initial gain level according to the second gain adjustment method corresponding to the minimum value of the optimal volume threshold range to obtain the target gain level; or if the first volume parameter is greater than the maximum value of the optimal volume threshold range, decrease the initial gain level according to the third gain adjustment method corresponding to the maximum value of the optimal volume threshold range to obtain the target gain level.

[0039] Optionally, the volume parameter adjustment module is specifically used to adjust the second volume parameter according to the second volume parameter and the target gain level if the second volume parameter is greater than or equal to the second minimum volume threshold.

[0040] Optionally, the device further includes:

[0041] The noise determination module is used to determine that the corresponding speech signal is noise if the second volume parameter is less than the second minimum volume threshold.

[0042] Optionally, each historical statistical period includes a preset number of voice signals, wherein the first volume parameter of the preset number of voice signals is greater than or equal to the second minimum volume threshold.

[0043] Thirdly, embodiments of this application also provide a voice control device, including: a processor, a storage medium, and a bus, wherein the storage medium stores program instructions executable by the processor, and when the voice control device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the voice gain control method as described in any of the first aspects.

[0044] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the voice gain control method as described in any of the first aspects.

[0045] The beneficial effects of this application are:

[0046] The voice gain control method, apparatus, voice control device, and storage medium provided in this application determine the target gain level for the current statistical period based on the volume parameters of multiple frames of voice signals within a historical statistical period. This allows for adjustment of the volume parameters of each frame of voice signal within the current statistical period according to the target gain level, achieving smooth adjustment of the voice signal's decibel value. This ensures that the volume of the voice signal remains stable across multiple periods, preventing listeners from experiencing abrupt or unstable volume fluctuations. Furthermore, in voice call scenarios, the speaker's volume can be adjusted to the optimal call volume without requiring the speaker to consciously adjust their speaking volume. Adjusting the volume of a speaker with a high volume to the optimal call volume also protects the listener's hearing and improves the user's listening experience. Attached Figure Description

[0047] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 An architecture diagram of the voice control system provided in an embodiment of this application;

[0049] Figure 2 A flowchart illustrating the speech gain control method provided in this application embodiment. Figure 1 ;

[0050] Figure 3 A flowchart illustrating the speech gain control method provided in this application embodiment. Figure 2 ;

[0051] Figure 4 A flowchart illustrating the speech gain control method provided in this application embodiment. Figure 3 ;

[0052] Figure 5 A flowchart illustrating the speech gain control method provided in this application embodiment. Figure 4 ;

[0053] Figure 6 A flowchart illustrating the speech gain control method provided in an embodiment of this application;

[0054] Figure 7 This is a schematic diagram of the structure of the voice gain control device provided in the embodiments of this application;

[0055] Figure 8 A schematic diagram of a voice control device provided in an embodiment of this application. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0057] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0058] Furthermore, the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Additionally, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0059] It should be noted that, where there is no conflict, the features in the embodiments of this application can be combined with each other.

[0060] Existing speech gain control methods generally employ single-step gain amplification of the speech signal, calculating the gain based on the current volume of the speech signal and adjusting the volume accordingly. However, the calculation of single-step gain may be inaccurate, leading to volume equalization failure. Another method uses the envelope value of the speech signal as the standard for gain adjustment, but this also struggles to achieve volume equalization. Therefore, it is evident that existing speech gain control methods lack stability and fail to achieve stable volume gain control for speech signals.

[0061] Before describing the voice gain control method, apparatus, voice control device and storage medium provided in the embodiments of this application, the voice control system involved in this application will be explained in order to better understand the solution of this application.

[0062] Please refer to Figure 1 Here is an architecture diagram of the voice control system provided in the embodiments of this application, as shown below. Figure 1As shown, the voice control system includes: a voice receiving module, a gain control module, and a voice transmitting module. The gain control module includes: a volume calculation unit, a gain calculation unit, and a gain control unit. The voice receiving module can be, for example, a microphone device, and the voice transmitting module can be, for example, a power amplifier module.

[0063] Specifically, the voice receiving module is used to receive voice signals. During a voice call, the voice signal is input to the voice receiving module as a signal stream using the Real-time Transport Protocol. The voice receiving module then sends the real-time voice signal to the gain control module.

[0064] The volume calculation unit in the gain control module calculates the volume of the real-time voice signal, obtains the volume parameters of the real-time voice signal, and sends the volume parameters of the real-time voice signal to the gain control unit; the gain calculation unit is used to determine the target gain level of the real-time voice signal based on the volume parameters of the historical volume signal, and sends the calculated target gain level to the gain control unit.

[0065] The gain control unit adjusts the volume parameters of the real-time voice signal according to the volume parameters and target gain level, and sends the real-time voice signal with adjusted volume parameters to the voice transmission module so that the voice transmission module can play the real-time voice signal with adjusted volume parameters.

[0066] Based on the above-described voice control system, the following description, in conjunction with embodiments, illustrates the voice gain control method, apparatus, voice control device, and storage medium provided in this application.

[0067] Please refer to Figure 2 The following is a flowchart illustrating the voice gain control method provided in the embodiments of this application. Figure 1 ,like Figure 2 As shown, the method may include:

[0068] S10: Determine the target gain level for the current statistical period based on the first volume parameter of multiple frames of speech signals within the historical statistical period.

[0069] In this embodiment, the statistical period of the voice signal can be divided according to the time of acquiring the voice signal or according to the number of frames of the acquired voice signal. For example, a statistical period can be limited by a preset duration or by a preset number of voice frames.

[0070] Each statistical period includes multiple frames of speech signals. The volume of the multiple frames of speech signals in the historical statistical period is calculated by the volume calculation unit to obtain the first volume parameter of the multiple frames of speech signals in the historical statistical period. The first volume parameter can be calculated based on the volume parameter of the multiple frames of speech signals in the historical statistical period after gain adjustment. The volume parameter can be the sound intensity of the speech signal.

[0071] In some embodiments, since the voice signal is continuous over a period of time, in order to ensure that the volume of the voice signal remains in a smooth and stable state, the volume parameters of the voice signal in the current statistical period can be adjusted according to the volume parameters of the voice signal in the historical statistical period.

[0072] Specifically, the historical statistical period is the statistical period preceding the current statistical period. The gain calculation unit determines the target gain level of the current statistical period based on the first volume parameter of multiple frames of speech signals within the historical statistical period in the following way: the target gain level of the current statistical period is determined based on the average volume parameter corresponding to the first volume parameter of multiple frames of speech signals within the historical statistical period.

[0073] In some examples, the target gain level can be determined based on the average volume parameter by calculating the difference between the average volume parameter and the optimal volume threshold range. The optimal volume threshold range corresponds to the optimal decibel level for human hearing; for example, the optimal volume threshold range could be set to 60 to 80 decibels.

[0074] The target gain level corresponds to the volume adjustment coefficient. If the average volume parameter is less than the minimum value of the optimal volume threshold range, the larger the difference, the higher the target gain level, indicating a larger volume adjustment coefficient. If the average volume parameter is greater than the maximum value of the optimal volume threshold range, the larger the difference, the lower the target gain level, indicating a smaller volume adjustment coefficient. The volume adjustment coefficient is set to 1 as the middle value, which means no adjustment, greater than 1 means increasing the volume, and less than 1 means decreasing the volume.

[0075] For example, please refer to Table 1, which shows the relationship between the volume adjustment coefficients corresponding to the target gain level. As shown in Table 1, if the target gain level is zero gain level, the volume adjustment coefficient is 1; if the target gain level is a gain level greater than zero gain level, the volume adjustment coefficient is greater than 1, and the larger the target gain level, the larger the volume adjustment coefficient; if the target gain level is a gain level less than zero gain level, the volume adjustment coefficient is less than 1, and the smaller the target gain level, the smaller the volume adjustment coefficient.

[0076] Table 1. Relationship between volume adjustment coefficients corresponding to target gain levels.

[0077] Target gain level -4 -3 -2 -1 0 1 2 3 4 Volume adjustment coefficient 0.2 0.4 0.6 0.8 1 1.3 1.5 2.0 3.0

[0078] In other examples, the target gain level can be determined based on the average volume parameter by determining the relationship between the average volume parameter and the optimal volume threshold range.

[0079] The system is divided into multiple gain levels, with zero gain as the default. These gain levels include: zero gain, multiple negative gain levels below zero, and multiple positive gain levels above zero. Positive gain levels are used to increase the volume of the speech signal, while negative gain levels are used to decrease it. Using zero gain or the gain level from the previous historical statistical period as a standard, the gain level is gradually decreased over multiple statistical periods when the average volume parameter exceeds the maximum value of the optimal volume threshold range; conversely, the gain level is gradually increased over multiple statistical periods when the average volume parameter is below the minimum value of the optimal volume threshold range.

[0080] It should be noted that adjusting the gain level step by step may require adjustments over multiple consecutive statistical cycles to bring the volume of the voice signal to the optimal level.

[0081] S20: Adjust the second volume parameter of each frame of speech signal according to the second volume parameter and target gain level of each frame of speech signal in the current statistical period.

[0082] In this embodiment, within the current statistical period, for each received frame of voice signal, the volume calculation unit calculates the volume of each frame of voice signal within the current statistical period to obtain a second volume parameter for each frame of voice signal within the current statistical period. The gain control unit adjusts the second volume parameter of each frame of voice signal within the current statistical period based on the second volume parameter and the target gain level. If the target gain level is used to indicate a decrease in volume, the gain control unit performs a gain operation on the second volume parameter to decrease the volume of the frame of voice signal; if the target gain level is used to indicate an increase in volume, the gain control unit performs a gain operation on the second volume parameter to increase the volume of the frame of voice signal.

[0083] After adjusting the second volume parameter of each frame of the voice signal in the current statistical period, the gain control unit plays the adjusted voice signal through the voice transmission module.

[0084] In some embodiments, the second volume parameter after adjustment of each frame of speech signal in the current statistical period is also included in the calculation of the target gain level in the next statistical period.

[0085] It should be noted that the target gain level can be calculated after the historical statistical period ends. After the voice signal of the current statistical period arrives, the calculated target gain level can be used directly to adjust the volume parameter of the voice signal of the current statistical period. There is no need to wait until the voice signal of the current statistical period is received before calculating the target gain level. This separates the calculation logic of the target gain level from the adjustment logic of the volume parameter, thereby improving the real-time performance of the volume parameter adjustment and playback of the voice signal.

[0086] The voice gain control method provided in the above embodiments determines the target gain level for the current statistical period based on the volume parameters of multiple frames of voice signals within a historical statistical period. It then adjusts the volume parameters of each frame of voice signal within the current statistical period according to the target gain level, achieving smooth adjustment of the decibel value of the voice signal. This ensures that the volume of the voice signal remains stable across multiple periods, preventing listeners from experiencing abrupt or unstable volume fluctuations. Furthermore, in voice call scenarios, the speaker's volume can be adjusted to the optimal call volume without requiring the speaker to consciously adjust their voice. Adjusting the volume of a speaker with a high volume to the optimal call volume also protects the listener's hearing and improves the user's listening experience.

[0087] Using the voice gain control method provided in the above embodiments, in scenarios of multi-person voice calls and voice conferences, since the volume parameter of the voice signal in the current statistical period needs to be adjusted according to the target gain level corresponding to the volume parameter of the voice signal in the historical statistical period, it can be ensured that the volume of all speakers in the scenario of multi-person voice calls and voice conferences remains basically consistent after adjustment, thus ensuring the call quality during the voice call or voice conference process.

[0088] The following describes one possible implementation of determining the target gain level, in conjunction with an embodiment.

[0089] Please refer to Figure 3 The following is a flowchart illustrating the voice gain control method provided in the embodiments of this application. Figure 2 ,like Figure 3 As shown, the process of determining the target gain level for the current statistical period based on the first volume parameter of multiple frames of speech signals within a historical statistical period in S10 may include:

[0090] S101: Determine whether the first volume parameter meets the preset gain level adjustment conditions.

[0091] S102: If the first volume parameter meets the preset gain level adjustment conditions, then adjust the initial gain level according to the gain adjustment method corresponding to the preset gain level adjustment conditions to obtain the target gain level.

[0092] In this embodiment, the preset gain level adjustment condition is used to indicate the volume conditions that need to be met to adjust the initial gain level to the target gain level. The preset gain level adjustment condition has a corresponding gain adjustment method, which is used to adjust the initial gain level corresponding to the current statistical period according to the corresponding gain adjustment method when the first volume parameter meets the preset gain level adjustment condition, so as to obtain the target gain level.

[0093] The initial gain level can be the default gain level of the current statistical period or the target gain level of the previous historical statistical period. The default gain level can be, for example, zero gain level, which means that the volume parameter of the speech signal is not increased or decreased. Adjusting the initial gain level to obtain the target gain level can be, for example, increasing the initial gain level or decreasing the initial gain level.

[0094] In some embodiments, the preset gain level adjustment conditions may include multiple adjustment conditions, each with a corresponding gain adjustment method. Based on the target adjustment conditions satisfied by the first volume parameter, the initial gain level is adjusted using the gain adjustment method corresponding to the target adjustment condition to obtain the target gain level. If the first volume parameter does not satisfy any of the multiple adjustment conditions, the initial gain level is not adjusted.

[0095] The voice gain control method provided in the above embodiments adjusts the initial gain level corresponding to the current statistical period based on the gain adjustment method corresponding to the preset gain level adjustment conditions satisfied by the first volume parameter, and obtains the target gain level. This ensures that the voice signal in the current statistical period is adjusted based on the target gain level, thereby ensuring that the volume of the voice signal in the historical statistical period and the current statistical period is smooth and stable. This prevents the listener from feeling the abruptness or instability of the voice signal volume and improves the user's listening experience.

[0096] In one possible implementation, please refer to Figure 4 The following is a flowchart illustrating the voice gain control method provided in the embodiments of this application. Figure 3 ,like Figure 4 As shown, the process of determining whether the first volume parameter meets the preset gain level adjustment conditions in S101 can include:

[0097] S111: Determine whether the first volume parameter of the multi-frame speech signal corresponding to the previous historical statistical period of the current statistical period is less than the first minimum volume threshold.

[0098] If the first volume parameter in step S102 meets the preset gain level adjustment conditions, then according to the gain adjustment method corresponding to the preset gain level adjustment conditions, the initial gain level corresponding to the current statistical period is adjusted to obtain the target gain level, which may include:

[0099] S121: If the first volume parameter of the previous historical statistical period is less than the first minimum volume threshold, then the initial gain level is increased according to the first gain adjustment method corresponding to the first minimum volume threshold to obtain the target gain level.

[0100] In this embodiment, the historical statistical period includes multiple consecutive historical statistical periods. The previous historical statistical period corresponding to the current statistical period is determined from multiple consecutive historical statistical periods. The first minimum volume threshold is a preset minimum volume threshold that ensures the speech signal can be heard clearly, also known as the minimum tolerance volume. Speech signals below this minimum tolerance volume are difficult to hear clearly. For example, the audio decibel value corresponding to the first minimum volume threshold can be 40 decibels.

[0101] Determine whether the average volume parameter of multiple frames of speech signals in the previous historical statistical period is less than the first minimum volume threshold. The first gain adjustment method corresponding to the first minimum volume threshold is to increase the gain level. If the average volume parameter of multiple frames of speech signals in the previous historical statistical period is less than the first minimum volume threshold, then increase the initial gain level to obtain the target gain level.

[0102] In this embodiment, if the average volume parameter of the multi-frame speech signal in the previous historical statistical period is less than the first minimum volume threshold, it means that the multi-frame speech signal in the previous historical statistical period is difficult for the listener to hear clearly. In this case, it is necessary to increase the gain level so that the volume parameter of the speech signal in the current statistical period is increased according to the target gain level in the next statistical period, i.e., the current statistical period, so that the volume parameter of the speech signal in the current statistical period is greater than the first minimum volume threshold, ensuring that the speech signal of the current statistical period being played can be heard clearly by the listener.

[0103] It should be noted that in a sustained audio signal, since the first statistical period lacks a gain level for adjusting volume parameters, only the first statistical period might be inaudible. In each subsequent statistical period, a target gain level can be determined based on the volume parameters of the previous historical statistical period. This adjustment ensures that the volume parameters of the audio signal in each subsequent period exceed a first minimum volume threshold, guaranteeing that the volume parameters of the audio signal in each subsequent statistical period are audible. Because each statistical period is very short, such as half a minute or one minute, the impact of the inaudible audio signal from the first statistical period is minimal, thus ensuring that the audio signal sustained for a sustained period is audible to the listener.

[0104] The voice gain control method provided in the above embodiments increases the initial gain level when the volume parameter of the multi-frame voice signal in the previous historical statistical period is less than the first minimum volume threshold, thereby obtaining the target gain level. Based on the adjustment of the target gain level, the volume of the voice signal in the current statistical period meets the requirement of the first minimum volume threshold, ensuring that the listener can hear the voice signal in the current statistical period clearly and improving the user's listening experience.

[0105] In another possible implementation, please refer to Figure 5 The following is a flowchart illustrating the voice gain control method provided in the embodiments of this application. Figure 4 ,like Figure 5 As shown, the process of determining whether the first volume parameter meets the preset gain level adjustment conditions in S101 can include:

[0106] S112: Determine whether the first volume parameter of the corresponding multi-frame speech signal within multiple consecutive historical statistical periods is within the optimal volume threshold range.

[0107] If the first volume parameter in step S102 meets the preset gain level adjustment conditions, then according to the gain adjustment method corresponding to the preset gain level adjustment conditions, the initial gain level corresponding to the current statistical period is adjusted to obtain the target gain level, which may include:

[0108] S122: If the first volume parameter is not within the optimal volume threshold range, adjust the initial gain level according to the gain adjustment method corresponding to the optimal volume threshold range to obtain the target gain level.

[0109] In this embodiment, the optimal volume threshold range can be a pre-set volume threshold range with the most ideal listening effect, or it can be called the tolerance range. The minimum value of the optimal volume threshold range is greater than the first minimum volume threshold.

[0110] To ensure that the volume of the speech signal is within the optimal volume threshold range, the average volume parameter of the speech signal of multiple frames in multiple consecutive historical statistical periods can be judged. If the average volume parameter of the speech signal of multiple frames in multiple consecutive historical statistical periods is not within the optimal volume threshold range, it is determined that the initial gain level needs to be adjusted. Specifically, the initial gain level is adjusted according to the gain adjustment method corresponding to the optimal volume threshold range to obtain the target gain level.

[0111] In some embodiments, if the first volume parameter is not within the optimal volume threshold range, the process of adjusting the initial gain level according to the gain adjustment method corresponding to the optimal volume threshold range in step S122 to obtain the target gain level may include:

[0112] If the first volume parameter is less than the minimum value of the optimal volume threshold range, the initial gain level is increased according to the second gain adjustment method corresponding to the minimum value of the optimal volume threshold range to obtain the target gain level.

[0113] In other embodiments, if the first volume parameter is not within the optimal volume threshold range, the process of adjusting the initial gain level according to the gain adjustment method corresponding to the optimal volume threshold range in step S122 to obtain the target gain level may include:

[0114] If the first volume parameter is greater than the maximum value of the optimal volume threshold range, the initial gain level is lowered according to the third gain adjustment method corresponding to the maximum value of the optimal volume threshold range to obtain the target gain level.

[0115] In this embodiment, the optimal volume threshold range is defined by a minimum value and a maximum value. The average volume parameter of the multi-frame speech signal corresponding to multiple consecutive historical statistical periods not being within the optimal volume threshold range includes: the average volume parameter of the multi-frame speech signal corresponding to multiple consecutive historical statistical periods being less than the minimum value of the optimal volume threshold range, or the average volume parameter of the multi-frame speech signal corresponding to multiple consecutive historical statistical periods being greater than the maximum value of the optimal volume threshold range.

[0116] When the volume parameter is less than the minimum value of the optimal volume threshold range, or when the volume parameter is greater than the maximum value of the optimal volume threshold range, the volume of the voice signal is not an ideal listening volume. In this case, the initial gain level needs to be adjusted according to the second gain adjustment method corresponding to the minimum value of the optimal volume threshold range, or the third gain adjustment method corresponding to the maximum value of the optimal volume threshold range, to obtain the target gain level.

[0117] For example, if the average volume parameter of the corresponding multi-frame speech signal in multiple consecutive historical statistical periods is less than the minimum value of the optimal volume threshold range, then the initial gain level is increased according to the second gain adjustment method corresponding to the minimum value of the optimal volume threshold range to obtain the target gain level; if the average volume parameter of the corresponding multi-frame speech signal in multiple consecutive historical statistical periods is greater than the maximum value of the optimal volume threshold range, then the initial gain level is decreased according to the third gain adjustment method corresponding to the maximum value of the optimal volume threshold range to obtain the target gain level; so that after adjustment based on the target gain level, the volume parameter of the speech signal in the current statistical period is within the optimal volume threshold range.

[0118] The voice gain control method provided in the above embodiments adjusts the initial gain level to obtain a target gain level when the first volume parameter of a multi-frame voice signal corresponding to multiple consecutive historical statistical periods is not within the optimal volume threshold range. This adjustment, based on the target gain level, ensures that the volume parameter of the voice signal within the current statistical period is within the optimal volume threshold range, thereby improving voice call quality. In voice call scenarios, the speaker's volume can be adjusted to the optimal call volume without the speaker consciously adjusting their speaking volume. Adjusting the volume of a speaker with a high volume to the optimal call volume can also protect the listener's hearing and improve the user's listening experience.

[0119] It should be noted that the judgment conditions for the first minimum volume threshold, the minimum value of the optimal volume threshold range, and the maximum value of the optimal volume threshold range can be independent of each other or related to each other. If they are related, when the first volume parameter of the previous historical statistical period is greater than or equal to the first minimum volume threshold, it is judged whether the first volume parameter of the corresponding multi-frame speech signal in multiple consecutive historical statistical periods is less than the minimum value of the optimal volume threshold range; if not, it is judged whether the first volume parameter of the corresponding multi-frame speech signal in multiple consecutive historical statistical periods is greater than the maximum value of the optimal volume threshold range; if none of the three conditions are met, the initial gain level is not adjusted.

[0120] Furthermore, calculating the target gain level based on the volume parameters of speech signals from multiple consecutive historical statistical periods can also avoid the impact of sudden speech signals such as coughing, sneezing, or door opening / closing sounds on the target gain level calculation during a continuous period of speech signals, thus ensuring the accuracy of the target gain level calculation.

[0121] In one possible implementation, the process of adjusting the second volume parameter of each frame of speech signal according to the second volume parameter and the target gain level of each frame of speech signal in the current statistical period, as described in S20, may include:

[0122] If the second volume parameter is greater than or equal to the second minimum volume threshold, the second volume parameter is adjusted according to the second volume parameter and the target gain level.

[0123] In this embodiment, the second minimum volume threshold is lower than the first minimum volume threshold. During voice calls, there may be some minor ambient noise. To avoid affecting the user's listening experience after amplification of the ambient noise, and to avoid affecting the calculation of the target gain level for the next statistical period, the noise needs to be filtered out. Only voice signals with a second volume parameter greater than or equal to the second minimum volume threshold are adjusted based on the target gain level. Voice signals with a second volume parameter greater than or equal to the second minimum volume threshold are considered valid voice signals. For example, the audio decibel value corresponding to the second minimum volume threshold can be 30 decibels.

[0124] In some embodiments, the method may further include:

[0125] If the second volume parameter is less than the second minimum volume threshold, the corresponding speech signal is determined to be noise.

[0126] In this embodiment, for a voice signal whose second volume parameter is less than the second minimum volume threshold, the voice signal is determined to be noise. It is possible to filter out the noise and not play it, or replace the noise with a preset comfortable noise that does not affect the listener's listening to other valid voice signals, and choose to play the comfortable noise.

[0127] In one possible implementation, each historical statistical period includes a preset number of voice signals, and the first volume parameter of the preset number of voice signals is greater than or equal to a second minimum volume threshold.

[0128] In this embodiment, based on the filtering of the above noise, it is known that the voice signal in each historical statistical period is a valid voice signal with the first volume parameter greater than or equal to the second minimum volume threshold. The number of frames of the valid voice signal in each historical statistical period is a preset number of frames, that is, the cumulative number of valid voice signals received is determined as one statistical period.

[0129] The voice gain control method provided in the above embodiments filters the voice signal based on the second minimum volume threshold to determine the effective voice signal. On the one hand, it can avoid the impact of environmental noise on the voice call quality, and on the other hand, it will not affect the calculation of the target gain level in the next statistical period, thus ensuring the accuracy of the target gain level calculation.

[0130] The flow of the voice gain control method provided in the above embodiments is described below with reference to the accompanying drawings.

[0131] Please refer to Figure 6 Here is a flowchart of the speech gain control method provided in the embodiments of this application, as shown below. Figure 6 As shown, the voice gain control process includes:

[0132] S31: The voice receiving module acquires the voice signal generated by the speaker's speech.

[0133] S32: The gain control module receives each frame of the voice signal for the current statistical period sent by the voice receiving module.

[0134] S33: Calculate the volume parameter of each frame of the speech signal in the current statistical period.

[0135] S34: Determine whether the volume parameter of each frame of the speech signal in the current statistical period is less than the second minimum volume threshold. If yes, jump to S35; otherwise, jump to S36.

[0136] S35: Determine that the audio signal in this frame is noise, and send comfortable noise to the audio transmission module.

[0137] S36: Determine that the audio signal in this frame is a valid audio signal, and adjust the volume parameter of the valid audio signal according to the initial gain level.

[0138] S37: Send the valid voice signal after gain / loss to the voice transmission module.

[0139] S38: The voice transmission module plays the effective voice signal after gain and loss and / or comfortable noise.

[0140] S39: Calculate the volume parameters of the effective speech signal after gain and loss.

[0141] S40: Determine whether the number of valid voice signals in the current statistical period has reached the preset number of frames. If not, end the process and determine that the current statistical period has not ended. If yes, determine that the current statistical period has ended and jump to S41 to calculate the target gain level for the next statistical period.

[0142] S41: Calculate the average volume parameter of the effective speech signal within the current statistical period based on the volume parameters after gain and loss.

[0143] S42: Determine whether the average volume parameter of the valid voice signal in the current statistical period is less than the first minimum volume threshold. If yes, proceed to S43; otherwise, proceed to S44.

[0144] S43: Increase the initial gain level to obtain the target gain level for the next statistical period.

[0145] S44: Determine whether the average volume parameter of the effective speech signal in multiple consecutive statistical periods is less than the minimum value of the optimal volume threshold range. If yes, proceed to S45; otherwise, proceed to S46. The multiple consecutive statistical periods include at least one historical statistical period and the current statistical period.

[0146] S45: Increase the initial gain level to obtain the target gain level for the next statistical period.

[0147] S46: Determine whether the average volume parameter of the effective speech signal in multiple consecutive statistical cycles is greater than the maximum value of the optimal volume threshold range. If yes, proceed to S47; otherwise, end the process and determine not to adjust the gain level of the next statistical cycle.

[0148] S47: Reduce the initial gain level to obtain the target gain level for the next statistical period.

[0149] As can be seen from the above speech gain control method, the initial gain level in the current statistical period is calculated based on the volume parameters of the speech signal in the historical statistical period. The calculation of the target gain level for the next statistical period is independent of the gain adjustment of the volume parameters of the speech signal in the current statistical period. That is, the gain control of the volume parameters in the current statistical period is independent of the calculation logic of the target gain level in the next statistical period. There is no need to calculate the gain level first and then adjust the volume parameters of the speech signal, thus improving the real-time performance of the gain adjustment of the speech signal in the current statistical period.

[0150] Based on the above method embodiments, this application provides a voice gain control device. Please refer to... Figure 7 This is a schematic diagram of the structure of the voice gain control device provided in the embodiments of this application, as shown below. Figure 7 As shown, the device may include:

[0151] The gain level determination module 10 is used to determine the target gain level of the current statistical period based on the first volume parameter of multiple frames of speech signals within the historical statistical period, wherein the historical statistical period is the period before the current statistical period.

[0152] The volume parameter adjustment module 20 is used to adjust the second volume parameter of each frame of speech signal according to the second volume parameter and target gain level of each frame of speech signal in the current statistical period.

[0153] Optionally, the gain level determination module 10 includes:

[0154] The adjustment condition judgment unit is used to determine whether the first volume parameter meets the preset gain level adjustment condition;

[0155] The gain level determination unit is used to adjust the initial gain level corresponding to the current statistical period according to the gain adjustment method corresponding to the preset gain level adjustment conditions if the first volume parameter meets the preset gain level adjustment conditions, so as to obtain the target gain level.

[0156] Optionally, the condition judgment unit is adjusted to determine whether the first volume parameter of the multi-frame speech signal corresponding to the previous historical statistical period of the current statistical period is less than the first minimum volume threshold.

[0157] The gain level determination unit is specifically used to increase the initial gain level according to the first gain adjustment method corresponding to the first minimum volume threshold if the first volume parameter of the previous historical statistical period is less than the first minimum volume threshold, so as to obtain the target gain level.

[0158] Optionally, the condition judgment unit is adjusted to determine whether the first volume parameter of the corresponding multi-frame speech signal within multiple consecutive historical statistical periods is within the optimal volume threshold range.

[0159] The gain level determination unit is specifically used to adjust the initial gain level according to the gain adjustment method corresponding to the optimal volume threshold range if the first volume parameter is not within the optimal volume threshold range, so as to obtain the target gain level.

[0160] Optionally, the gain level determination unit is specifically used to increase the initial gain level according to the second gain adjustment method corresponding to the minimum value of the optimal volume threshold range if the first volume parameter is less than the minimum value of the optimal volume threshold range, so as to obtain the target gain level.

[0161] Optionally, the gain level determination unit is specifically used to reduce the initial gain level and obtain the target gain level if the first volume parameter is greater than the maximum value of the optimal volume threshold range, according to the third gain adjustment method corresponding to the maximum value of the optimal volume threshold range.

[0162] Optionally, the volume parameter adjustment module 20 is specifically used to adjust the second volume parameter according to the second volume parameter and the target gain level if the second volume parameter is greater than or equal to the second minimum volume threshold.

[0163] Optionally, the device may also include:

[0164] The noise determination module is used to determine that the corresponding speech signal is noise if the second volume parameter is less than the second minimum volume threshold.

[0165] Optionally, each historical statistical period includes a preset number of voice signals, wherein the first volume parameter of the preset number of voice signals is greater than or equal to the second minimum volume threshold.

[0166] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.

[0167] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more microprocessors, or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).

[0168] This application also provides a voice control device, please refer to... Figure 8 This is a schematic diagram of the voice control device provided in the embodiments of this application, such as... Figure 8 As shown, the voice control device 100 includes a processor 101, a storage medium 102, and a bus. The storage medium 102 stores program instructions executable by the processor 101. When the voice control device 100 is running, the processor 101 communicates with the storage medium 102 via the bus, and the processor 101 executes the program instructions to perform the above-described method embodiment. The specific implementation and technical effects are similar and will not be described in detail here.

[0169] Optionally, the present invention also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the above-described method embodiments.

[0170] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0171] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0172] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0173] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0174] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A voice gain control method, characterized in that, The method includes: After the historical statistical period ends, the target gain level of the current statistical period is determined based on the first volume parameter of the multiple frames of speech signals within the historical statistical period. The first volume parameter is calculated based on the volume parameter of the multiple frames of speech signals within the historical statistical period after gain adjustment. After each frame of speech signal in the current statistical period is reached, the second volume parameter of each frame of speech signal is adjusted according to the second volume parameter of each frame of speech signal in the current statistical period and the target gain level; The step of determining the target gain level for the current statistical period based on the first volume parameter of multiple frames of speech signals within the historical statistical period includes: Determine whether the first volume parameter meets the preset gain level adjustment conditions; If the first volume parameter satisfies the preset gain level adjustment condition, then the initial gain level is adjusted according to the gain adjustment method corresponding to the preset gain level adjustment condition to obtain the target gain level, wherein the initial gain level is the target gain level of the previous historical statistical period.

2. The method as described in claim 1, characterized in that, The step of determining whether the first volume parameter meets the preset gain level adjustment conditions includes: Determine whether the first volume parameter of the multi-frame speech signal corresponding to the previous historical statistical period of the current statistical period is less than the first minimum volume threshold. If the first volume parameter satisfies the preset gain level adjustment condition, then according to the gain adjustment method corresponding to the preset gain level adjustment condition, the initial gain level is adjusted to obtain the target gain level, including: If the first volume parameter of the previous historical statistical period is less than the first minimum volume threshold, then the initial gain level is increased according to the first gain adjustment method corresponding to the first minimum volume threshold to obtain the target gain level.

3. The method as described in claim 1, characterized in that, The step of determining whether the first volume parameter meets the preset gain level adjustment conditions includes: Determine whether the first volume parameter of the corresponding multi-frame speech signal within multiple consecutive historical statistical periods is within the optimal volume threshold range; If the first volume parameter satisfies the preset gain level adjustment condition, then according to the gain adjustment method corresponding to the preset gain level adjustment condition, the initial gain level is adjusted to obtain the target gain level, including: If the first volume parameter is not within the optimal volume threshold range, then the initial gain level is adjusted according to the gain adjustment method corresponding to the optimal volume threshold range to obtain the target gain level; Wherein, if the first volume parameter is less than the minimum value of the optimal volume threshold range, the initial gain level is increased according to the second gain adjustment method corresponding to the minimum value of the optimal volume threshold range to obtain the target gain level; or... If the first volume parameter is greater than the maximum value of the optimal volume threshold range, the initial gain level is lowered according to the third gain adjustment method corresponding to the maximum value of the optimal volume threshold range to obtain the target gain level.

4. The method as described in claim 1, characterized in that, The step of adjusting the volume parameter of each frame of speech signal based on the second volume parameter of each frame of speech signal within the current statistical period and the target gain level includes: If the second volume parameter is greater than or equal to the second minimum volume threshold, the second volume parameter is adjusted according to the second volume parameter and the target gain level.

5. The method as described in claim 4, characterized in that, The method further includes: If the second volume parameter is less than the second minimum volume threshold, the corresponding voice signal is determined to be noise.

6. The method as described in claim 4, characterized in that, Each historical statistical period includes a preset number of voice signals, wherein the first volume parameter of the preset number of voice signals is greater than or equal to the second minimum volume threshold.

7. A voice gain control device, characterized in that, The device includes: The gain level determination module is used to determine the target gain level of the current statistical period after the end of the historical statistical period, based on the first volume parameter of the multiple frames of speech signals in the historical statistical period. The first volume parameter is calculated based on the volume parameter of the multiple frames of speech signals in the historical statistical period after gain adjustment. The volume parameter adjustment module is used to adjust the second volume parameter of each frame of speech signal after each frame of speech signal in the current statistical period is reached, based on the second volume parameter of each frame of speech signal in the current statistical period and the target gain level. The gain level determination module includes: The adjustment condition judgment unit is used to determine whether the first volume parameter meets the preset gain level adjustment condition; The gain level determination unit is used to adjust the initial gain level corresponding to the current statistical period according to the gain adjustment method corresponding to the preset gain level adjustment condition if the first volume parameter meets the preset gain level adjustment condition, so as to obtain the target gain level, wherein the initial gain level is the target gain level of the previous historical statistical period.

8. A voice control device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores program instructions executable by the processor, and when the voice control device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the voice gain control method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, performs the steps of the speech gain control method as described in any one of claims 1 to 6.