Audio processing method based on dynamic gain adjustment, electronic equipment and storage medium

Through the sliding window mechanism and multi-stage gain adjustment method, the microphone gain of the Bluetooth voice remote control is dynamically adjusted, solving the volume adaptation problem of the remote control with limited resources in complex environments, and improving audio quality and user experience.

CN120390180APending Publication Date: 2025-07-29AISPEECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510428455.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The microphone gain configuration of the Bluetooth voice remote control is static and has limited resources, which makes the volume not suitable for complex usage environments, affecting the voice recognition rate and user experience.

Method used

The sliding window mechanism is used to combine sampling point energy comparison and multi-point statistics to perform high-volume trigger detection. The hardware analog gain is reduced or restored in stages through multi-stage gain adjustment, avoiding the sharp jump in sampling point energy, and introducing a holding phase to reduce frequent detection.

Benefits of technology

It realizes the smoothness and nature of the microphone gain adjustment process, reduces system resource usage, and improves audio quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390180A_ABST
    Figure CN120390180A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing method based on dynamic gain adjustment, electronic equipment and a storage medium. The method comprises the following steps: carrying out high-volume trigger detection through a sliding window mechanism in combination with sampling point energy comparison and multi-point statistics; when a large-volume scene is detected, hardware simulation gain is reduced in stages by adopting multi-stage gain adjustment; entering a holding stage after the reduction of the hardware analog gain is completed, re-starting the sliding window mechanism to detect whether the large volume disappears or not after the holding stage is ended, and pausing the detection and maintaining the current gain during the holding stage; and when it is confirmed that the large volume disappears, recovering the hardware analog gain by stages by adopting multi-stage gain adjustment. According to the embodiment of the invention, the multi-stage gain adjustment process is introduced, so that the energy of the sampling point is prevented from abruptly jumping along with the gain, and the gain adjustment process of the microphone is more stable and natural, and is free of time delay, low in calculation amount and low in system resource occupation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of audio processing, and particularly relates to an audio processing method based on dynamic gain adjustment, an electronic device and a storage medium. Background Art

[0002] In the current field of voice remote controls, the microphone gain is statically configured, and this configuration is often determined during the R & D process of software and hardware. Conventional automatic gain control (AGC) schemes rely on additional circuits, or complex algorithms and models. However, the single-microphone structure of Bluetooth voice remote controls and the shortage of MCU resources restrict the deployment of traditional AGC algorithm schemes. For a small number of remote control schemes, they support the voice remote control on the mobile terminal (mobile phone or TV). Through the Bluetooth link channel, the microphone gain setting is sent to modify, which belongs to the category of dynamic update and is not automatic gain adjustment. In addition, some manufacturers improve the recognition rate of audio of medium and low-quality remote controls by optimizing the cloud voice recognition model; or add some user prompts on the terminal side such as TVs and projectors to guide users to use the remote control in a standardized manner.

[0003] Infrared remote controls have the advantages of low cost, mature technology, high reliability, etc. Therefore, for a long time, the remote controls of household electrical appliances have mainly been infrared remote controls. However, the infrared transmission angle and distance are limited, and it must be within a relatively close range and aligned with the device direction to be effectively used. With the development of smart home appliances and the increasing requirements for their use experience, the low-power Bluetooth remote control technology has attracted more and more attention from manufacturers and consumers because of its characteristics such as "no need to align", "can control around corners", "long control distance", and "higher transmission bandwidth". The progress of human-computer interaction technology has promoted the replacement of conventional infrared and Bluetooth button controls by voice control, which can complete more abundant intention controls and provide a better interaction and use experience.

[0004] An audio front-end module usually includes acoustic echo cancellation (AEC), automatic noise suppression (ANS), and automatic gain control (AGC), commonly known as the 3A algorithm. AEC and ANS respectively solve the problems of echo and noise in the call and audio processing scenarios, greatly improving the user experience. In the scenario of volume problems, if the volume is too small, we cannot clearly hear the specific voice information. If the volume is too large, we not only cannot hear clearly, but our ears will also be damaged. If the volume fluctuates, rising and falling, being erratic, it will undoubtedly be a "torture" for users, and the experience is extremely poor. Although a voice remote control can preset a moderate microphone gain to make the output audio volume of the remote control at a suitable level, it is not convenient enough and it is difficult to satisfy everyone. In actual scenarios, the usage environment of the remote control is complex and changeable. The original volume of the voice varies with different speakers. For example, the volume of the elderly and children is slightly lower, and the volume of middle-aged people is slightly higher. The attenuation of sound propagation varies with the distance between the speaker and the microphone. If the microphone sensor of the same remote control hardware and software solution is replaced, the collected audio volume will also be different. In some usage scenarios, when the user's initial voice interaction fails to obtain the expected recognition and semantic results for various reasons, the user will increase the volume and subconsciously speak the command words more clearly and precisely. If the cloud still cannot return the recognition result due to poor audio collection, resulting in no execution response on the device side, it will seriously reduce the user's enthusiasm for use. The above-mentioned differentiated usage scenarios may all lead to deviations in the voice recognition results and affect the experience. If only relying on the preset gain during the R & D process or dynamic setting during the use process, it will inevitably bring a burden to users and is not "elegant" enough for product design.

[0005] The audio collected and output by the voice remote control is mainly uploaded to the cloud, and the cloud's language processing model performs text recognition. Therefore, poor-quality collected audio will affect the recognition rate of the cloud recognition engine. Further, in the production test scenario, the current verification of the microphone and voice channel of the remote control mainly relies on the tester manually pressing the voice button, speaking the command word, and listening with the human ear to judge whether the audio broadcast by the companion board is consistent with the test command word, so as to judge whether the current hardware under test passes; therefore, "clipping" or distorted sound caused by improper acquisition gain will affect the subjective judgment of the tester.

[0006] The inventors found that: Conventional automatic gain control solutions rely on additional circuits, or complex algorithms and models. Bluetooth voice remote controls usually need to control costs, and have a low main frequency, limited memory and Flash. Therefore, a method for automatically adjusting the microphone gain suitable for a remote control system with limited resources and low cost is designed to optimize the audio quality in the product and production test stages and improve the user experience. Summary of the Invention

[0007] The embodiments of the present invention aim to solve at least one of the above technical problems.

[0008] In a first aspect, an embodiment of the present invention provides an audio processing method based on dynamic gain adjustment, including: detecting large volume triggering through a sliding window mechanism combined with sampling point energy comparison and multi-point statistics; when a large volume scenario is detected, adopting multi-stage gain adjustment to gradually reduce the hardware analog gain; entering a holding stage after the reduction of the hardware analog gain is completed, and re-enabling the sliding window mechanism to detect whether the large volume has disappeared after the end of the holding stage, wherein during the holding stage, the detection is paused and the current gain is maintained; when it is confirmed that the large volume has disappeared, adopting the multi-stage gain adjustment to gradually restore the hardware analog gain.

[0009] In a second aspect, an embodiment of the present invention provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any one of the above-mentioned audio processing methods based on dynamic gain adjustment of the present invention.

[0010] In a third aspect, an embodiment of the present invention provides a storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) for executing any one of the above-mentioned audio processing methods based on dynamic gain adjustment of the present invention.

[0011] In a fourth aspect, an embodiment of the present invention further provides a computer program product, which includes a computer program stored on a storage medium, and the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute any one of the above-mentioned audio processing methods based on dynamic gain adjustment.

[0012] In the embodiment of the present invention, by using a sliding window mechanism, combined with audio frame sampling point energy comparison and multi-point statistics, large volume detection is realized; then, in the sound reduction stage, a multi-stage gain adjustment process is introduced to avoid sharp jumps in sampling point energy caused by the gain, resulting in noise. By introducing a holding stage, overly frequent energy detection and gain adjustment are avoided; subsequently, based on the same sliding window and energy detection mechanism, it is determined whether the large volume has disappeared; finally, if there is no large volume, with the help of the multi-stage gain adjustment process, gain callback and sound increase are completed. This method makes the microphone gain adjustment process smoother, more natural, with no delay, low computational complexity, and low system resource occupancy. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0014] Figure 1 It is a flowchart of an embodiment of the audio processing method based on dynamic gain adjustment of the present invention; Figure 2 It is a flowchart of another embodiment of the audio processing method based on dynamic gain adjustment of the present invention; Figure 3 It is a schematic diagram of the sound reduction process of the audio processing method based on dynamic gain adjustment of the present invention; Figure 4 It is a schematic diagram of the holding stage of the audio processing method based on dynamic gain adjustment of the present invention; Figure 5 It is a schematic diagram of the sound increase process of the audio processing method based on dynamic gain adjustment of the present invention; Figure 6 It is a flowchart of an audio processing process based on dynamic gain adjustment provided by an embodiment of the present invention; Figure 7 It is a schematic diagram of the structure of an embodiment of the electronic device of the present invention. Detailed implementation manners

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0016] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0017] The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0018] In the present invention, "module", "device", "system", etc. refer to relevant entities applied to a computer, such as hardware, a combination of hardware and software, software, or software in execution, etc. Specifically, for example, an element can be, but is not limited to, a process running on a processor, a processor, an object, an executable element, an execution thread, a program, and / or a computer. Also, an application program or a script program running on a server, and the server can both be elements. One or more elements can be in the process and / or thread of execution, and the elements can be localized on one computer and / or distributed between two or more computers, and can be run by various computer-readable media. The elements can also communicate through local and / or remote processes according to a signal having one or more data packets, for example, a signal from data that interacts with another element in a local system, a distributed system, and / or interacts with other systems through a signal on a network of the Internet.

[0019] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also include other elements not explicitly listed, or also include elements inherent to such a process, method, article, or device. Without more limitations, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article, or device including the said elements.

[0020] An embodiment of the present invention provides an audio processing method based on dynamic gain adjustment, and this method can be applied to an electronic device. The electronic device can be a computer, a server, or other electronic products, etc., and the present invention does not make any limitations in this regard.

[0021] Please refer to Figure 1 , which shows an audio processing method based on dynamic gain adjustment provided by an embodiment of the present invention.

[0022] As Figure 1 shown, in step 101, large-volume trigger detection is performed through a sliding window mechanism combined with sampling point energy comparison and multi-point statistics; In step 102, when a large-volume scene is detected, multi-stage gain adjustment is adopted to gradually reduce the hardware analog gain; In step 103, after the reduction of the hardware analog gain is completed, it enters a holding stage, and after the end of the holding stage, the sliding window mechanism is re-enabled to detect whether the large volume has disappeared, where detection is paused and the current gain is maintained during the holding stage; In step 104, when it is confirmed that the large volume disappears, the multi-level gain adjustment is adopted to gradually restore the hardware analog gain.

[0023] In this embodiment, for step 101, the large volume trigger detection is performed by combining the sliding window mechanism with the sampling point energy comparison and multi-point statistics. For the large volume detection based on the sampling point energy, no additional delay can be introduced in the voice remote control, and not too much system resources can be occupied; in addition, the large volume scenario needs to be rigorously identified to avoid the intrusion and misjudgment of random signals. By mainly performing arithmetic comparisons on each sampling point and the statistical values of each frame within the window, and a division operation once per frame, it is ensured that no additional delay is introduced in the detection; only the statistical values are temporarily stored within the sliding window, and the additional system overhead is controllable; the sliding window and multi-point judgment strategies ensure the rigor of the judgment.

[0024] The large volume trigger detection of this application adopts a normalized judgment method based on the sampling point energy. Conventional recorders segment the streaming audio in units of frames and periodically notify the system of the completion of each frame of recording. Common frame lengths are 8, 16, 32 milliseconds, etc., and each frame contains a specified number of sampling points N sp 。

[0025] After that, for step 102, after detecting the large volume scenario, the multi-level gain adjustment gradually reduces the hardware analog gain. In the volume reduction stage, if the microphone gain is adjusted to the target gear at one time, it is easy to cause the sampling point energy to jump sharply with the gain and induce noise; for this reason, inspired by the attack period and release period of the traditional AGC scheme, this application introduces a multi-level gain adjustment process. The object of the microphone gain adjustment is the hardware analog gain. Although it cannot be strictly smoothed, when the number of steps increases, the attenuation process of the audio waveform will be smoother. Define the gain adjustment interval in the volume reduction stage as INTV dec , which is an integer multiple of the frame duration, that is, the adjustment node is also at the frame recording completion callback; define the number of steps as N steps , the gain reduction is G delta decibels (dB), and the gain adjusted each time is G step decibels, that is . When the large volume is detected, the first gain adjustment is performed; as in the verification example, INTV dec is 96 milliseconds, the frame length is 8 milliseconds, assuming the first adjustment time point is T0 , then the second adjustment time point T1 is after milliseconds, that is, after 12 frames. Further, the large volume detection is not performed during the adjustment process.

[0026] Then, for step 103, after reducing the hardware simulation gain is completed, enter the hold stage. After the hold stage ends, re-enable the sliding window mechanism to detect whether the loud volume has disappeared. During the hold stage, pause the detection and maintain the current gain. To avoid overly frequent energy detection and gain adjustment, the hold stage is introduced. The hold stage does not perform loud volume detection nor adjust the gain; its duration is defined as INTV hold , starting from the last down-tone gain adjustment, and the duration is also an integer multiple of the frame duration. When each hold interval ends, start performing loud volume detection; if detected, it means the loud volume state continues and has not disappeared, and it is necessary to stay in the hold stage and re-time; if not detected, it means the loud volume state has disappeared and it is necessary to enter the up-tone stage of the volume callback. Specifically, each time loud volume detection starts, clear the sliding window queue. Further, since the hardware gain of the current microphone is no longer the default value, the absolute value upper limit threshold of the sampling point energy E sp (defined as U limit_hold ) needs to be updated according to the formula , converting the change amount in the DB domain of the input microphone signal into the change amount in the linear domain; the other settings related to loud volume detection remain unchanged.

[0027] Finally, for step 104, when it is confirmed that the loud volume has disappeared, use multi-stage gain adjustment to gradually restore the hardware simulation gain. The stepped up-tone process of this application is similar to the down-tone process. If the microphone gain is adjusted to the target gear at once, it is easy to cause the sampling point energy to jump sharply with the gain and induce noise. Therefore, multi-stage gain adjustment is adopted to make the up-tone process transition smoothly. The up-tone gain adjustment interval is INTV inc , the same as INTV dec ; the number of steps follows N steps , the gain increase amplitude is G delta decibels, and each time the gain is adjusted by G step decibels. Further, the up-tone adjustment process does not perform loud volume detection. After the up-tone is completed, re-enter the loud volume detection stage.

[0028] The method of the embodiment of the present application realizes high - volume detection by using a sliding window mechanism in combination with audio frame sampling point energy comparison and multi - point statistics. Then, in the volume reduction stage, a multi - level gain adjustment process is introduced to avoid sharp jumps in sampling point energy due to gain, which may cause noise. By introducing a holding stage, overly frequent energy detection and gain adjustment are avoided. Subsequently, based on the same sliding window and energy detection mechanism, it is determined whether the high volume has disappeared. Finally, if there is no high volume, the gain callback and volume increase are completed through the multi - level gain adjustment process. This method makes the microphone gain adjustment process smoother, more natural, with no delay, low computational complexity, and low system resource occupancy.

[0029] Please refer to Figure 2 , which shows another audio processing method based on dynamic gain adjustment provided by an embodiment of the present invention. This flowchart mainly further defines the steps of "performing high - volume trigger detection through a sliding window mechanism combined with sampling point energy comparison and multi - point statistics" in step 101 of the flowchart Figure 1 .

[0030] As Figure 2 shown, in step 201, for multiple sampling points in each audio frame, the energy of the multiple sampling points is compared with a preset energy threshold to obtain the number of super - energy sampling points of the sampling points whose energy is greater than or equal to the preset energy threshold among the multiple sampling points; In step 202, a sliding window mechanism is introduced based on a queue data structure to record the number of super - energy sampling points of audio frames in the most recent preset number of frames. The sliding window can accommodate the number of super - energy sampling points of the preset number of frames; In step 203, the number of super - energy sampling points is compared with a preset point threshold to obtain the number of audio frames in the sliding window whose number of super - energy sampling points is greater than or equal to the preset point threshold; In step 204, when the ratio of the number of audio frames to the preset number of frames is greater than or equal to a preset ratio, it is determined that a high - volume trigger occurs.

[0031] In this embodiment, for step 201, for multiple sampling points in each audio frame, the energy of the multiple sampling points is compared with a preset energy threshold to obtain the number of super - energy sampling points of the sampling points whose energy is greater than or equal to the preset energy threshold among the multiple sampling points. First, the preset sampling point energy E sp absolute value upper - limit threshold (defined as U limit ) is a certain ratio of the sampling point energy range. For example, when the sampling bit width is 16 bits, its sampling point energy range is [-32768, 32768]. If the preset upper - limit threshold is 80%, then the sampling point energy E spThe absolute upper limit threshold is detected as 26214. Secondly, when each frame of recording is completed, the energy of all sampling points in the current frame is arithmetically compared with U limit and the number of points with energy greater than U limit is recorded N bv .

[0032] After that, for step 202, a sliding window mechanism is introduced based on the queue data structure. Introducing the sliding window mechanism according to the queue data structure is used to record the number of super-energy sampling points of the audio frames in the most recent preset number of frames. Among them, the sliding window can accommodate the number of super-energy sampling points of the preset number of frames; for example, introducing the sliding window mechanism based on the queue data structure is used to record the N bv of the most recent audio frames, and the initial values are all 0; the window width represents the N bv of recording up to several frames of audio, which is defined as L win , for example, preset to 6 frames; the N bv of each frame of audio is stored at the end of the window queue. If the window is full, the oldest record is removed and the subsequent records are moved forward one grid as a whole.

[0033] Then, for step 203, compare the number of super-energy sampling points with the preset point threshold to obtain the number of audio frames in the sliding window where the number of super-energy sampling points is greater than or equal to the preset point threshold; finally, for step 204, when the ratio of the number of the audio frames to the preset number of frames is greater than or equal to the preset ratio, it is determined that a high-volume trigger occurs. For example, in the callback processing when each frame of recording is completed, judge N bv is greater than the preset number of points SP bv of the number of frames F bv ; if F bv the proportion is greater than the preset ratio R bv , that is , then it is determined that a high-volume scene occurs; for example, R bv is set to two-thirds, L win is 6, F bv it is required that no less than 4 frames are judged as a high-volume trigger.

[0034] The method of the embodiment of the present application mainly performs arithmetic comparison on each sampling point and each frame statistical value within the window, and a division operation once per frame, ensuring that the detection does not introduce additional delay; only the statistical values are temporarily stored within the sliding window, and the additional system overhead is controllable; the sliding window and multi-point judgment strategy ensure the rigor of the judgment.

[0035] In some optional embodiments, when comparing the number of super-energy sampling points with a preset point threshold, normalization processing is adopted; the energy range of the super-energy sampling points is set according to the bit width of the number of super-energy sampling points; arithmetic comparison is performed on all sampling points of each frame and the number of super-energy sampling points is counted. A normalized judgment method based on the energy of sampling points is adopted. Conventional recorders segment streaming audio in frames and periodically notify the system of the completion of each frame of recording. Common frame lengths are 8, 16, 32 milliseconds, etc., and each frame contains a specified number of sampling points. N sp 。

[0036] In some optional embodiments, the multi-level gain adjustment is adopted, including: defining the gain adjustment interval as an integer multiple of the frame duration; setting the total gain change amount and the number of adjustment steps; each time the gain adjustment amount is the total gain change amount divided by the number of adjustment steps. Whether in the volume reduction stage or the volume increase stage, multi-level gain adjustment is adopted. The multi-level gain adjustment is to adjust the microphone gain, and the object is the analog gain of the hardware. Whether in the volume reduction stage or the volume increase stage, it is impossible to achieve strict smoothness, but when the number of steps increases or decreases, the attenuation process of the audio waveform will be smoother.

[0037] Please refer to Figure 3 , which shows a schematic diagram of the volume reduction process of the audio processing method based on dynamic gain adjustment of the present invention.

[0038] As Figure 3 shown, when detecting a large volume for the first time, the first gain reduction is immediately executed. In the subsequent volume reduction stage, each adjustment interval is set according to the duration of the first detection of a large volume. During the gain adjustment period in the volume reduction stage, the detection of a large volume is paused. When detecting a large volume, the first gain adjustment is executed; as in the verification example, INTV dec is 96 milliseconds, the frame length is 8 milliseconds, assuming the first adjustment time point is T0 , then the second adjustment time point T1 is after milliseconds, that is, after 12 frames. Further, the large volume detection is not performed during the adjustment process.

[0039] Please refer to Figure 4 , which shows a schematic diagram of the holding stage of the audio processing method based on dynamic gain adjustment of the present invention.

[0040] As Figure 4As shown, after the hardware simulation gain is reduced, it enters the holding phase. Starting from the last volume reduction adjustment, the sliding window queue is reset during the holding period, and the preset energy threshold of the sampling point is updated. The holding duration of the holding phase is an integer multiple of the frame duration. To avoid overly frequent energy detection and gain adjustment, the holding phase is introduced. The large volume detection is not performed during the holding phase, nor is the gain adjusted; its duration is defined as INTV hold , starting from the last volume reduction gain adjustment, and the duration is also an integer multiple of the frame duration.

[0041] After the holding phase ends, the sliding window mechanism is re-enabled to detect whether the large volume has disappeared. If the large volume is detected, it is necessary to stay in the holding phase and restart the timing; if the large volume is not detected, it enters the volume increase phase. When each holding interval ends, the large volume detection is started; if detected, it means that the large volume state continues and has not disappeared, and it is necessary to stay in the holding phase and restart the timing; if not detected, it means that the large volume state has disappeared and it is necessary to enter the volume increase phase of the volume callback.

[0042] Please refer to Figure 5 , which shows the schematic diagram of the volume increase process of the audio processing method based on dynamic gain adjustment of the present invention.

[0043] As Figure 5 shown, the hardware simulation gain is restored in stages by using multi-level gain adjustment. Restoring the hardware simulation gain for the volume increase phase also includes that the adjustment interval of each time in the volume increase phase is set according to the adjustment interval of the volume reduction phase, and the same number of steps as the volume reduction phase is maintained for gain restoration; after the gain adjustment in the volume increase phase is completed, the large volume detection is reactivated. For example, the volume increase gain adjustment interval is INTV inc , the same as INTV dec ; the number of steps follows N steps , the gain increase amplitude is G delta decibels, and the gain is adjusted by G step decibels each time. The large volume detection is not performed during the volume increase adjustment process. After the volume increase is completed, it re-enters the large volume detection stage.

[0044] Please refer to Figure 6 , which shows the implementation flowchart of the audio processing method based on dynamic gain adjustment of the present invention.

[0045] As Figure 6As shown, first, the sliding window mechanism is used, combined with the comparison of the energy of audio frame sampling points and multi-point statistics, to achieve high-volume detection; then, in the volume reduction stage, a multi-level gain adjustment process is introduced to avoid the sharp jump of the sampling point energy with the gain, which may cause noise; further, by introducing a holding stage, overly frequent energy detection and gain adjustment are avoided; subsequently, based on the same sliding window and energy detection mechanism, it is determined whether the high volume has disappeared; finally, if there is no high volume, the gain callback and volume increase are completed through the multi-level gain adjustment process.

[0046] It should be noted that traditional AGC algorithm solutions involve processes such as linear domain conversion to logarithmic domain, gain calculation, gain smoothing, and logarithmic domain conversion to linear domain, corresponding to operators such as normalization, level calculation, DRC static curve calculation, and FFT. The algorithm occupies a relatively large amount of system CPU and memory resources and also introduces a certain amount of additional delay to the audio stream. Taking the comparison of Soc platform 1 as an example, it has a heterogeneous dual-core architecture including an ARM Cortex M4F MCU and a Cadence HiFi3 voice-specific DSP. The former executes application and business logic, and the latter executes voice algorithms; the digital AGC algorithm on its DSP requires 11 MCPS of computing power. Coincidentally, the comparison of Soc platform 2 has a heterogeneous dual-core architecture including an ARM Cortex M4 MCU and a CEVA enhanced TL4x voice-specific DSP, and the digital AGC algorithm on its DSP requires 10 MCPS of computing power. The MCUs of Bluetooth voice remote controls are mainly Cortex M0 and M3, with common upper frequency limits of 48MHz and 64MHz, and limited memory and Flash, which restricts the deployment of traditional AGC algorithm solutions. In this application, the high-volume trigger detection and holding stage and high-volume disappearance detection mainly perform arithmetic comparisons on each sampling point and the statistical values of each frame within the sliding window, and a division operation once per frame, ensuring that the detection does not introduce additional calculations and computing power requirements, and does not require caching audio frames, thus not introducing additional delay; only statistical values are temporarily stored within the sliding window, and the additional system memory overhead is controllable; the sliding window and multi-point judgment strategy also ensure the rigor of the judgment.

[0047] The object of the microphone gain adjustment in this design is the analog gain of the hardware; the change in the analog gain has an instantaneous effect on the amplitude of the voice signal. The stepped volume reduction process and stepped volume increase process in this application introduce stepped multi-level gain adjustment to avoid inducing noise caused by the sharp jump of the sampling point energy with the gain if the microphone gain is adjusted to the target gear at one time, enabling a smooth transition between the volume increase and volume reduction processes. In addition, by adding a callback link, the remote control can adapt to both high-volume and low-volume scenarios, and has better compatibility with complex and changing usage environments.

[0048] The audio processing method based on dynamic gain adjustment of the present application relates to a low-power Bluetooth voice remote control (BLE VOICE Remote Control Unit, hereinafter referred to as RCU); the RCU generally includes a microphone, an infrared transceiver, a keyboard, an MCU, a memory, a Flash, a low-power Bluetooth BLE, audio codec technology, etc., and belongs to an embedded system in the intelligent Internet of Things voice application scenario.

[0049] The inventors of the present application have noticed that for the increasingly popular smart home appliances and Bluetooth voice remote controls, due to the inability of the volume to automatically adapt to the usage scenario, it is easy to cause obstacles to the intelligent promotion of products and poor user experience. A method for automatically adjusting the microphone gain of a remote control micro-system suitable for resource-limited and low-cost is designed to optimize the audio quality in the product stage and the production test stage, improve the user experience, enhance the product competitiveness, and solve the long-existing but easily overlooked technical problems in this field. Those skilled in the art should clearly understand that the invention is not limited to a specific Bluetooth version, not limited to the specific brands and models of the remote control and the host, not limited to the hosting platform (ARM or RISC-V, etc.) implemented by means of the present invention, or other hardware forms and embedded systems.

[0050] The present application also provides an alternative solution: in the holding stage, each time the detection of the disappearance of the maximum volume is performed, a new U limit_hold . is calculated and updated. To reduce the calculation amount, U limit_hold it can be calculated in advance on a computer and directly used and compared in the holding stage. Because in the verification example U limit and U limit_hold are preset, and the relationship between the two based on the formula is also determined. This strategy can further simplify the calculation amount of the MCU system. For example, in the verification example, this calculation consumes about 400 microseconds.

[0051] The present application also provides a Beta version: during the process of increasing and decreasing the sound, the microphone gain was once adjusted to the final target gear; however, the energy of the sampling point, that is, the volume, jumps sharply with the gain, and there is a problem of low recognition rate for the words near the adjustment time point, and it also sounds quite abrupt to the human ear.

[0052] It should be noted that the inventor has found similar technologies to this application: a. CN218957055 U Clothes Drying Machine Volume Control System and Clothes Drying Machine; b. CN116184855 A Method and System for Human-Machine Interaction Volume Control Adaptive to Users; c. CN118972726 A Method and Device for Controlling the Gain of an Intelligent Sound-Picking Device Based on the User's Usage Status; d. CN112333606 B Method and Device for Adjusting Abnormal Microphone Gain; e. CN110650410 A Method, Device and Storage Medium for Automatic Microphone Gain Control; f. CN112242147 A Method for Voice Gain Control and Computer Storage Medium.

[0053] Regarding the prior art, in which the volume adjustment of intelligent clothes drying racks is improper and lacks humanity, Technology a provides a clothes drying machine volume control system that adaptively controls the voice volume of the voice module of the clothes drying machine according to the user's identity, facilitating the use of various family members and effectively improving the user experience of using the clothes drying machine.

[0054] Regarding the volume adjustment technology of existing intelligent household appliances, in which external environmental factors during use are not considered, resulting in sudden device announcements or unclear sounds due to fixed volumes, Technology b invents a method for automatically adjusting the broadcast volume of intelligent household appliances based on the user's voice volume and the external environmental volume.

[0055] Regarding the existing technology for suppressing howling, which mainly adjusts the gain based on the volume value of the sound signal played by the speaker and attempts to avoid howling by reducing the gain, Technology c invents a method for controlling the gain of an intelligent sound-picking device based on the usage status.

[0056] Regarding softphone products, due to the variety of input devices, it is difficult to effectively and automatically keep the sound sizes collected by different input devices within a reasonable range. Technology d proposes an automatic gain control method that can effectively solve the adaptive adjustment of the microphone input volume gain.

[0057] Regarding the online meeting scenario, when the microphone is separated from the terminal device, the sound collected by the microphone may be clipped and cause terminal echo. Technology e invents a method and device that can timely determine the unreasonable gain setting of the microphone based on the output signal of the microphone and timely adjust the gain of the microphone according to the preset method.

[0058] Regarding the conference call scenario, the emergence of instantaneous noise distracts the audience's attention and results in a poor experience. Moreover, there are defects in existing AGC technologies and algorithms. Therefore, an invention is made for a speech gain control method based on a deep learning neural network. The amplitude of instantaneous noise is greatly reduced, far lower than the gain amplification threshold in AGC. The gain correction process can also reduce the risk of instantaneous noise being misamplified by the AGC algorithm, improving the quality of voice calls.

[0059] In practical applications, when applying the automatic gain control method, one is to perform gain control on the current signal collected by the microphone, and the other is to perform gain control on the output voice through the speaker.

[0060] Both Technology a and Technology b are for adjusting the volume of the voice output by the speaker and do not involve microphone gain control.

[0061] Technology c is for intelligent sound collection device gain control based on the usage status and relies on sensors or cameras (preferably) on the microphone to obtain environmental data within the sound collection area of the microphone, and then relies on a judgment module to judge whether there is preset target feature data in the environmental data to identify the usage status of the microphone.

[0062] Technology d needs to rely on a Voice Activity Detection (VAD) module, and combines the microphone and the echo channel to detect the magnitude of voice energy and subsequent automatic adjustment, without paying attention to the single microphone scenario.

[0063] In the scenario where the microphone is separated from the terminal device, Technology e detects the microphone energy on the terminal and adjusts the pre-stage microphone in a timely manner. It only focuses on how to reduce the voice energy and does not pay attention to increasing it again; and it only executes gain control based on whether a single point exceeds the threshold, with low rigor.

[0064] Technology f uses technologies such as Fourier transform and neural network to improve the call quality in the instantaneous noise scenario, but these two technologies consume a lot of resources and are not suitable for small and micro systems.

[0065] Technology a and Technology b do not pay attention to the application scenario of the Bluetooth voice remote control and do not notice the automatic gain control requirements of intelligent devices with only a microphone and no speaker.

[0066] Technology c relies on external sensors, cameras and other auxiliary tools for sensing data and image collection, and then compares with the preset usage status, without noticing the automatic gain control requirements without the assistance of external tools.

[0067] The detection method of Technology d relies on VAD, which requires additional system resources, and is not suitable for the single microphone scenario of the Bluetooth voice remote control for intelligent conference devices.

[0068] Although the detection means of Technique e is more efficient and has low system resource occupancy, it only performs gain control based on whether a single point exceeds the threshold, without considering dynamic gain callback, so its applicable scenarios are limited.

[0069] Technique e is also targeted at intelligent conference devices, which have rich resources and can deploy complex AI algorithm models, but it is not suitable for small and micro systems with compact resources.

[0070] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a combination of a series of actions. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention. In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0071] In some embodiments, the embodiments of the present invention provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to be used for executing any one of the above audio processing methods based on dynamic gain adjustment of the present invention.

[0072] In some embodiments, the embodiments of the present invention further provide a computer program product, the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute any one of the above audio processing methods based on dynamic gain adjustment.

[0073] In some embodiments, the embodiments of the present invention further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute an audio processing method based on dynamic gain adjustment.

[0074] Figure 7 It is a schematic hardware structure diagram of an electronic device for executing an audio processing method based on dynamic gain adjustment provided by another embodiment of the present application. As Figure 7 shown, the device includes: One or more processors 710 and a memory 720, Figure 7Take a processor 710 as an example.

[0075] The device that executes the audio processing method based on dynamic gain adjustment may further include: an input device 730 and an output device 740.

[0076] The processor 710, the memory 720, the input device 730, and the output device 740 may be connected through a bus or other means. Figure 7 Take the connection through the bus as an example.

[0077] As a non-volatile computer-readable storage medium, the memory 720 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the audio processing method based on dynamic gain adjustment in the embodiments of the present application. By running the non-volatile software programs, instructions, and modules stored in the memory 720, the processor 710 executes various functional applications and data processing of the server, that is, implements the audio processing method based on dynamic gain adjustment in the above method embodiments.

[0078] The memory 720 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the audio processing device based on dynamic gain adjustment, etc. In addition, the memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 720 may optionally include a memory remotely set relative to the processor 710, and these remote memories can be connected to the audio processing device based on dynamic gain adjustment through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0079] The input device 730 can receive input digital or character information, and generate signals related to the user settings and function controls of the audio processing device based on dynamic gain adjustment. The output device 740 may include a display device such as a display screen.

[0080] The one or more modules are stored in the memory 720 and, when executed by the one or more processors 710, execute the audio processing method based on dynamic gain adjustment in any of the above method embodiments.

[0081] The above product can execute the method provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects of the executed method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.

[0082] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0083] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc.

[0084] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and intelligent toys and portable vehicle navigation devices.

[0085] (4) Other on-board electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.

[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0088] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. An audio processing method based on dynamic gain adjustment, comprising: Performing large-volume trigger detection through a sliding window mechanism combined with sampling point energy comparison and multi-point statistics; When a large-volume scene is detected, adopting multi-level gain adjustment to gradually reduce the hardware analog gain in stages; Entering a holding stage after the reduction of the hardware analog gain is completed. After the end of the holding stage, re-enabling the sliding window mechanism to detect whether the large volume disappears, wherein during the holding stage, the detection is paused and the current gain is maintained; When it is confirmed that the large volume has disappeared, adopting the multi-level gain adjustment to gradually restore the hardware analog gain in stages.

2. The method according to claim 1, wherein: The performing large-volume trigger detection through a sliding window mechanism combined with sampling point energy comparison and multi-point statistics includes: For multiple sampling points in each audio frame, comparing the energy of the multiple sampling points with a preset energy threshold to obtain the number of super-energy sampling points of the sampling points whose energy is greater than or equal to the preset energy threshold among the multiple sampling points; Introducing a sliding window mechanism based on a queue data structure for recording the number of super-energy sampling points of audio frames in the most recent preset number of frames. The sliding window can accommodate the number of super-energy sampling points of the preset number of frames; Comparing the number of super-energy sampling points with a preset point threshold to obtain the number of audio frames in the sliding window whose number of super-energy sampling points is greater than or equal to the preset point threshold; When the ratio of the number of the audio frames to the preset number of frames is greater than or equal to a preset ratio, it is determined that a large-volume trigger occurs.

3. The method according to claim 2, wherein: The comparing the number of super-energy sampling points with a preset point threshold adopts normalization processing, which includes: Setting the energy range of the number of super-energy sampling points according to the bit width of the number of super-energy sampling points; Performing arithmetic comparison on all sampling points of each frame and counting the number of super-energy sampling points.

4. The method according to claim 1, wherein The adopting multi-level gain adjustment includes: Defining the gain adjustment interval as an integer multiple of the frame duration; Setting the total gain change amount and the number of adjustment steps; Each time the gain adjustment amount is the total gain change amount divided by the number of adjustment steps.

5. The method according to claim 1, wherein, When a large-volume scene is detected, adopting multi-level gain adjustment to gradually reduce the hardware analog gain in stages. Reducing the hardware analog gain is the volume reduction stage, and the method includes: Immediately performing the first gain reduction when the large volume is first detected. The subsequent adjustment interval of each time in the volume reduction stage is set according to the duration when the large volume is first detected. Among them, during the gain adjustment in the volume reduction stage, the large-volume detection is paused.

6. The method according to claim 1, wherein: The adopting the multi-level gain adjustment to gradually restore the hardware analog gain in stages. Restoring the hardware analog gain is the volume increase stage, and the method includes: The adjustment interval of each time in the volume increase stage is set according to the adjustment interval of the volume reduction stage, and the gain is restored while maintaining the same number of steps as the volume reduction stage; After the gain adjustment in the volume increase stage is completed, reactivate the large-volume detection.

7. The method according to claim 1, wherein: The entering the holding stage after the reduction of the hardware analog gain is completed includes: Starting from the last volume reduction adjustment, and resetting the sliding window queue and updating the preset energy threshold of the sampling points during the holding period, wherein the holding duration of the holding stage is an integer multiple of the frame duration.

8. The method according to claim 1, wherein: Re-enabling the sliding window mechanism to detect whether the large volume disappears after the end of the holding stage includes: If the high volume is detected, it is necessary to maintain the holding stage and restart the timing; If the high volume is not detected, enter the volume increasing stage.

9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1 to 8.

10. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Microphone automatic gain control method and device and storage medium

    CN110650410A