Sound mixing control method and system

The adaptive audio mixing control method, which combines real-time peak detection and dynamic threshold linkage, solves the problems of delay and false triggering in existing automatic audio mixing technologies. It enables rapid response and accurate judgment of sudden speech signals, thereby improving the user's auditory experience.

CN120998220APending Publication Date: 2025-11-21HANSONG NANJING TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511350569.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing automatic audio mixing technology suffers from delays and high false trigger rates when responding to sudden speech signals, making it difficult to distinguish between human voices and environmental noise, resulting in a poor listening experience.

Method used

An adaptive audio mixing control method that combines real-time peak detection and dynamic threshold linkage identifies energy spikes in the mixed audio signal and adjusts the background volume based on the trigger threshold. The trigger threshold is then optimized using a machine learning model to adapt to complex environments.

Benefits of technology

It effectively reduces volume adjustment lag time, improves the accuracy of sound judgment and the smoothness of volume adjustment, and enhances the user's auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998220A_ABST
    Figure CN120998220A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a sound mixing control method, and the method comprises the steps: obtaining a background audio signal, and a mixed audio signal containing potential voice and environment noise; carrying out peak detection on the mixed audio signal so as to identify a candidate signal interval with an energy bump point on a target frequency band of the mixed audio signal; and reducing the volume of the background audio signal in response to the energy characteristic of the candidate signal interval exceeding the trigger threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This manual relates to the field of audio control, and in particular to a mixing control method and system. Background Technology

[0002] In applications such as in-car navigation, live streaming, and voice assistants, automatic audio mixing technology automatically reduces the volume of background audio sources (such as music) when the main audio source (such as human voice) appears, thus ensuring the clarity of the main audio source. However, most current automatic mixing technologies are based on the Root Mean Square (RMS) algorithm, which determines whether to trigger mixing by calculating the average energy of the audio. This can lead to high response delays and an inability to react promptly to sudden speech signals. Furthermore, this algorithm struggles to effectively distinguish between human voices and environmental noise (such as applause or shouting), resulting in a high false trigger rate, while also posing a risk of missing triggers for short speech segments. In addition, most other volume switching methods are abrupt and unnatural, easily producing an jarring listening experience, and their fixed trigger thresholds make them difficult to adapt to complex and changing environments.

[0003] Therefore, there is an urgent need for a mixing control method and system that can solve the above problems, respond quickly to sudden speech signals, accurately judge and naturally switch between speech and background music, reduce latency and false triggering problems, and intelligently adjust trigger flexibility to improve adaptive environment capabilities and enhance the auditory experience. Summary of the Invention

[0004] This specification provides one or more embodiments of a mixing control method, the method comprising: acquiring a background audio signal and a mixed audio signal containing potential speech and environmental noise; performing peak detection on the mixed audio signal to identify a candidate signal interval with an energy spike point in a target frequency band of the mixed audio signal; and reducing the volume of the background audio signal in response to the energy characteristics of the candidate signal interval exceeding a trigger threshold.

[0005] This specification provides one or more embodiments of a mixing control system, the system comprising: an acquisition module configured to acquire a background audio signal and a mixed audio signal including potential speech and ambient noise; an identification module configured to perform peak detection on the mixed audio signal to identify candidate signal intervals with energy spikes in a target frequency band of the mixed audio signal; and a volume reduction module configured to reduce the volume of the background audio signal in response to the energy characteristics of the candidate signal intervals exceeding a trigger threshold. Attached Figure Description

[0006] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1 These are schematic diagrams illustrating application scenarios of the mixing control system according to some embodiments of this specification; Figure 2 This is an exemplary block diagram of a mixing control system according to some embodiments of this specification; Figure 3 This is an exemplary flowchart of a mixing control method according to some embodiments of this specification; Figure 4 This is an exemplary schematic diagram illustrating the determination of a trigger threshold according to some embodiments of this specification; Figure 5 This is an exemplary schematic diagram illustrating the determination of candidate signal intervals according to some embodiments of this specification. Detailed Implementation

[0007] The accompanying drawings used in the description of the embodiments will be briefly introduced below. The drawings do not represent all embodiments.

[0008] The terms “system,” “device,” “unit,” and / or “module” as used herein are one method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0009] Unless the context clearly indicates an exception, words such as "a," "an," "a kind," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0010] Ducking technology, especially when applied to scenarios where background music is reduced to trigger human voices (e.g., lowering the music volume when navigation prompts), can significantly improve the auditory experience. Currently widely used audio mixing technology based on RMS energy averaging may suffer from problems such as lag in background music volume adjustment, inability to effectively distinguish between human voices and ambient noise, missed triggers, and abrupt volume transitions, which seriously affect the user experience.

[0011] This invention provides a mixing control method and system that effectively reduces volume adjustment lag time, improves sound judgment accuracy and the smoothness and naturalness of volume adjustment by performing adaptive audio mixing control with real-time peak detection and dynamic threshold linkage. It can be applied to complex environmental scenarios and meets user experience requirements.

[0012] Figure 1 This is a schematic diagram illustrating an application scenario of a mixing control system according to some embodiments of this specification. In some embodiments, such as Figure 1 As shown, the application scenario 100 of the audio mixing control system may include a processor 110, an audio acquisition device 120, an audio playback device 130, a network 140, and a storage device 150.

[0013] In some embodiments, the application scenario 100 of the mixing control system may include indoor environments (such as home theaters, conference rooms, cinemas, stadiums, shopping malls, live broadcast rooms, etc.) or other environments (such as inside a vehicle).

[0014] In some embodiments, the processor 110 can process data and / or information obtained from components of the mixing control system application scenario 100 or other external devices. The processor can execute program instructions based on this data, information, and / or processing results to perform one or more functions described in this application. For example, acquiring background audio signals and mixed audio signals, identifying candidate signal intervals, reducing the volume of the background audio signal, etc.

[0015] In some embodiments, processor 110 may be a local processor relative to the mixing control system or an external processor. In some embodiments, processor 110 may be a computer, a user console, a single processor, or a processor group, etc. The processor group may be centralized or distributed. In some embodiments, processor 110 may be implemented on a cloud platform. For example, the cloud platform may include one or any combination of private cloud, public cloud, hybrid cloud, etc.

[0016] Audio acquisition device 120 refers to a device capable of acquiring audio signals. In some embodiments, audio acquisition device 120 may integrate a microphone array, a WIFI module, a Bluetooth module, and a network module, etc. For example, audio acquisition device 120 may be a microphone 121, a recording device 122, etc., or any combination thereof.

[0017] In some embodiments, the audio acquisition device 120 can be used to acquire background audio signals, mixed audio signals, etc.

[0018] Audio playback device 130 refers to a device that plays audio signals. For example, speakers, audio equipment, audio players, etc.

[0019] Network 140 includes any suitable network capable of facilitating information and / or data exchange within the application scenario 100 of the mixing control system. In some embodiments, one or more components of the application scenario 100 of the mixing control system (e.g., processor 110, audio acquisition device 120, audio playback device 130, and storage device 150, etc.) can exchange information and / or data via network 140.

[0020] Network 140 can be any one or more of wired or wireless networks. For example, network 140 may include Bluetooth, WIFI, cable network, fiber optic network, telecommunications network, cable connection, etc., or any combination thereof.

[0021] In some embodiments, storage device 150 may be used to store data and / or instructions. For example, storage device 150 may store background audio signals and mixed audio signals, etc.

[0022] Storage device 150 may include one or more storage components, each of which may be a separate device or part of another device. In some embodiments, storage device 150 may include random access memory (RAM), read-only memory (ROM), etc. In some embodiments, storage device 150 may be implemented on a cloud platform. By way of example only, a cloud platform may include a private cloud, a public cloud, a hybrid cloud, etc., or any combination thereof.

[0023] In some embodiments, the storage device 150 can communicate with one or more components in the application scenario 100 of the mixing control system via the network 140.

[0024] For further explanation of the parameters mentioned above, such as background audio signal, mixed audio signal, and candidate signal range, please refer to [link / reference]. Figure 3 The relevant description in the document.

[0025] It should be noted that the above description of application scenario 100 is for convenience only and should not be construed as limiting this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of this system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles. In some embodiments, the modules may share a single storage module, or each module may have its own separate storage module. Such modifications are all within the scope of this specification.

[0026] Figure 2 This is an exemplary block diagram of a mixing control system according to some embodiments of this specification.

[0027] In some embodiments, such as Figure 2As shown, the mixing control system 200 may include an acquisition module 210, an identification module 220, and a volume reduction module 230. In some embodiments, some or all of the acquisition module 210, the identification module 220, and the volume reduction module 230 may be configured in a processor.

[0028] The acquisition module 210 refers to the module used to acquire audio signals.

[0029] In some embodiments, the acquisition module 210 may be configured to acquire a background audio signal and a mixed audio signal containing potential speech and environmental noise.

[0030] The recognition module 220 refers to a module used for recognizing and / or processing audio signals.

[0031] In some embodiments, the identification module 220 is configured to perform real-time peak detection on the mixed audio signal to identify candidate signal intervals with energy spikes in the target frequency band of the mixed audio signal.

[0032] In some embodiments, the identification module 220 is further configured to determine the extreme point of the energy change rate in the target frequency band; and to determine the candidate signal range based on the extreme point and the release threshold.

[0033] Volume reduction module 230 refers to a module used to reduce volume.

[0034] In some embodiments, the volume reduction module 230 is configured to reduce the volume of the background audio signal in response to the energy characteristics of the candidate signal interval exceeding a trigger threshold.

[0035] In some embodiments, the volume reduction module 230 is further configured to estimate the baseline level of ambient noise based on the background audio signal and the mixed audio signal; and to determine a trigger threshold based on the baseline level.

[0036] In some embodiments, the volume reduction module 230 is further configured to determine a trigger threshold based on historical trigger data using a parametric model, wherein the parametric model is a machine learning model, and the historical trigger data includes a noise change rate determined based on a baseline level.

[0037] In some embodiments, the volume reduction module 230 is further configured to determine a reduction sequence; based on the reduction sequence, reduce the volume of the background audio signal.

[0038] For further explanation of the above parameters (such as candidate signal range, trigger threshold, and amplitude reduction sequence), please refer to [link to relevant documentation]. Figures 3-5 Related descriptions.

[0039] It should be noted that the above description of the mixing control system and its modules is for ease of description only and should not be construed as limiting this specification to the embodiments described. It is understood that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles.

[0040] Figure 3 This is an exemplary flowchart of a mixing control method according to some embodiments of this specification.

[0041] In some embodiments, the mixing control method may be executed by a processor. For example... Figure 3 As shown, process 300 may include steps 310-330.

[0042] Step 310: Acquire the background audio signal and the mixed audio signal containing potential speech and environmental noise.

[0043] Background audio signals refer to the audio signals played by audio playback devices. Examples include background music and game sound effects played by audio playback devices.

[0044] Mixed audio signals refer to the raw audio signals acquired by audio acquisition devices.

[0045] In some embodiments, the mixed audio signal may include potential speech and ambient noise.

[0046] Latent speech refers to the speech signal that needs to be determined to be the target speech. Target speech refers to the effective speech signal in the mixed audio that needs to be enhanced or prioritized. Examples include navigator voices, conference speakers' voices, and animal sounds.

[0047] Environmental noise refers to interfering sound signals from the surrounding environment. Examples include the hum of an air conditioner, the sound of keyboard typing, and the sound of wind.

[0048] In some embodiments, the mixed audio signal may further include a background audio re-sampling signal.

[0049] Background audio re-acquisition signal refers to the sound signal that is re-captured by the audio acquisition device after the background audio signal has been played.

[0050] For example, in a vehicle interior environment, the mixed audio signal can be a combination of the navigation voice, the air conditioning hum, and the background music recaptured by the audio acquisition device.

[0051] In some embodiments, the background audio signal can be retrieved by the processor from a storage device or obtained from a third-party platform (such as music software); the mixed audio signal can be acquired by an audio acquisition device and transmitted to the processor. In some embodiments, the background audio signal and the mixed audio signal can also be acquired by other means, such as manual input, etc., which are not limited here.

[0052] Step 320: Peak detection is performed on the mixed audio signal to identify candidate signal intervals with energy spikes in the target frequency band of the mixed audio signal.

[0053] Peak detection refers to the process of identifying the transient characteristics of the initial sound by analyzing the instantaneous changes in the energy of an audio signal in real time.

[0054] The target frequency band refers to the frequency band where the energy of the target speech is concentrated. For example, when the target speech is a navigation voice, the target frequency band can be 300Hz-3kHz.

[0055] An energy spike point is a sampling point where the energy of a sound signal undergoes a sudden change. For example, a sampling point where the magnitude of the change in sound signal energy exceeds a rate of change threshold. The rate of change threshold can be preset manually.

[0056] In some embodiments, the processor can perform spectral analysis on the mixed audio signal to determine the target frequency band of the mixed audio signal and identify energy spikes in the target frequency band. Spectral analysis methods may include, but are not limited to, Fast Fourier Transform (FFT) algorithms, power spectral density analysis algorithms, etc.

[0057] Candidate signal intervals refer to audio signal intervals extracted from the target frequency band that contain energy spikes and are to be verified as containing the target speech.

[0058] In some embodiments, the processor can determine the candidate signal interval using various methods. For example, the processor can determine the candidate signal interval by taking the energy spike point identified in the target frequency band of the mixed audio signal as the starting time and the point when the sound energy is continuously lower than a preset energy threshold for a preset duration as the ending time.

[0059] In some embodiments, the processor can determine the extreme points of the rate of energy change in the target frequency band; based on the extreme points and the release threshold, a candidate signal interval is determined. For further explanation of this section, see [link to relevant documentation]. Figure 5 And related descriptions.

[0060] Step 330: In response to the energy characteristics of the candidate signal region exceeding the trigger threshold, the volume of the background audio signal is reduced.

[0061] Energy characteristics refer to parameters related to the energy of the audio signal within a candidate signal range. Examples include peak amplitude, range RMS value, and energy integral.

[0062] Peak amplitude refers to the maximum energy of the audio signal within the candidate signal interval. Interval RMS value refers to the average energy of the audio signal within the candidate signal interval. Energy integral refers to the sum of the squares of the energy values ​​of the audio signal at all sampling time points within the candidate signal interval.

[0063] In some embodiments, the processor can determine energy characteristics by performing spectral analysis on candidate signal intervals.

[0064] A trigger threshold is a preset threshold related to energy characteristics, used to determine whether to trigger a pitch reduction operation. For example, a trigger threshold may include at least one of a peak amplitude threshold, an interval RMS threshold, and an energy integral threshold.

[0065] In some embodiments, the trigger threshold can be preset manually. In some embodiments, the processor can estimate the baseline level of ambient noise based on the background audio signal and the mixed audio signal; and determine the trigger threshold based on the baseline level. For more information on this section, please refer to [link to relevant content]. Figure 4 Related descriptions.

[0066] Volume reduction refers to lowering the volume of the background audio signal.

[0067] In some embodiments, in response to the energy characteristics of a candidate signal region exceeding a trigger threshold, the processor can reduce the volume of the background audio signal using a variety of methods.

[0068] For example, the processor can reduce the volume of the background audio signal according to a preset total volume reduction margin. This preset total volume reduction margin can be set to the magnitude by which the energy characteristic exceeds a trigger threshold. For instance, if the interval RMS value exceeds the interval RMS threshold by 20%, then the preset total volume reduction margin is 20%.

[0069] In some embodiments, the processor may also determine a reduction sequence; based on the reduction sequence, the volume of the background audio signal is reduced.

[0070] A volume reduction sequence is a sequence of data composed of multiple volume reduction amplitudes across multiple volume reduction periods during a volume reduction process, arranged chronologically. The total duration of the multiple volume reduction periods is the same as the duration of the entire volume reduction process, and all volume reduction periods have the same duration. The volume reduction amplitude can be expressed as a percentage or a decimal.

[0071] In some embodiments, the volume reduction magnitudes in the reduction sequence are arranged in descending order or in a step-descending order, and the sum of the multiple volume reduction magnitudes equals the target total reduction magnitude. The target total reduction magnitude refers to the total volume reduction of the background audio signal to meet the user's listening needs. For example, if the target total reduction magnitude is 0.8, the reduction sequence can be represented as (0.4, 0.2, 0.15, 0.05) or (0.3, 0.3, 0.1, 0.1), where the former is an example of multiple volume reduction magnitudes arranged in descending order and the latter is an example of multiple volume reduction magnitudes arranged in a step-descending order.

[0072] Taking the volume reduction sequence (0.4, 0.2, 0.15, 0.05) as an example, it represents that the entire volume reduction process includes four volume reduction periods, with the volume reduction increments from beginning to end being 0.4, 0.2, 0.15, and 0.05 respectively. Here, 0.4 indicates that the volume is reduced by 40% in the first volume reduction period (i.e., the volume at the end of the first volume reduction period is 40% lower than the volume at the beginning of the first volume reduction period), and the rest follow the same logic.

[0073] It's important to note that during the volume reduction process, the volume decrease is not continuous but rather achieved through a series of discrete, minute steps. For example, if the initial volume of the first volume reduction period is 60dB, and 0.4 represents a 40% volume reduction during that period, the processor can perform multiple volume reductions at various sampling points within the first volume reduction period, lowering the volume from 60dB to 36dB. 36dB is the final volume of the first volume reduction period. The number of sampling points is extremely large to achieve a completely smooth volume change in the perceived sound.

[0074] In some embodiments, the processor may randomly generate a reduction sequence that satisfies preset sequence conditions. The preset sequence conditions are that the multiple volume reduction magnitudes are arranged in descending order or in a step-down order, and the sum of the multiple volume reduction magnitudes equals the target total reduction magnitude.

[0075] In some embodiments, the processor can discretely reduce the volume according to a reduction sequence to achieve a smooth attenuation of the background audio signal's volume. For example, the reduction sequence can be represented as (0.4, 0.2, 0.15, 0.05). The processor can then select multiple reduction points within the first reduction period, with equal or similar intervals between any adjacent reduction points; reduction is performed at each reduction point, with the volume reduction amplitude at multiple reduction points being the same or similar, and the sum being 0.4. "Similar" can mean that the interval difference is not greater than a first difference threshold, and the difference in volume reduction amplitude is not greater than a second difference threshold. The first and second difference thresholds can be preset manually. The same applies to other reduction periods.

[0076] In some embodiments, by determining a reasonable reduction sequence to perform volume reduction operation, a non-linear smooth volume decay is achieved to improve the user's listening experience.

[0077] In some embodiments of this specification, candidate signal intervals with energy spikes are accurately identified in the target frequency band through real-time peak detection. This can accurately filter non-human noise. By combining the baseline level of dynamic environmental noise to set a reasonable trigger threshold, the timing of volume reduction of the background audio signal is adaptively determined, effectively reducing trigger delay, false trigger rate, and missed trigger rate. Furthermore, when the energy characteristics of the candidate signal interval exceed the trigger threshold, the volume of the background audio signal is smoothly attenuated according to a non-linear reduction sequence, eliminating auditory discomfort caused by sudden audio drop, achieving intelligent suppression of background audio volume, and improving the user's auditory experience.

[0078] Figure 4 This is an exemplary schematic diagram illustrating the determination of a trigger threshold according to some embodiments of this specification.

[0079] In some embodiments, such as Figure 4 The processor can estimate the baseline level 430 of ambient noise based on the background audio signal 410 and the mixed audio signal 420; and determine the trigger threshold 460 based on the baseline level 430.

[0080] For more information on background audio signals, mixed audio signals, ambient noise, and trigger thresholds, please refer to [link / reference needed]. Figure 3 And its related descriptions.

[0081] The baseline level refers to the average energy representation of ambient noise. In some embodiments, the baseline level can represent the real-time average energy of ambient noise after excluding interference from background audio re-sampling signals.

[0082] In some embodiments, the processor may obtain a predicted echo signal based on the background audio signal using an echo model; determine a near-end clean signal based on the mixed audio signal and the predicted echo signal; and estimate the baseline level of ambient noise based on speech-free segments of the near-end clean signal.

[0083] An echo model is a model used to simulate echoes. For example, an echo model can be the physical transfer function of the acoustic propagation process, simulating sound output from an audio playback device, its propagation through space, and its acquisition by an audio acquisition device. For instance, an echo model may include adaptive filtering algorithms, etc.

[0084] The predicted echo signal refers to the theoretical value of the background audio re-sampling signal calculated through echo model simulation. In some embodiments, the processor may treat the predicted echo signal as the background audio re-sampling signal.

[0085] Near-end clean signal refers to the sound signal retained after the mixed audio signal has undergone echo cancellation.

[0086] In some embodiments, the processor can determine the near-end clean signal as the residual after subtracting the predicted echo signal from the mixed audio signal. The near-end clean signal includes potential speech and ambient noise.

[0087] Speechless segments refer to segments in the near-end clean signal that do not contain human voices.

[0088] In some embodiments, the processor can perform frame processing on the near-end clean signal based on a preset window length and calculate the RMS of each frame; frames that meet preset conditions are identified as speech-free frames; and N consecutive speech-free frames are identified as speech-free segments. The preset conditions and the value of N can be set by default by the processor or preset manually based on experience. For example, the preset conditions can be RMS less than the silence threshold, N=10, etc.

[0089] In some embodiments, the processor can determine the average energy of the speechless segments as the baseline level of the ambient noise.

[0090] In some embodiments, the processor may update the baseline level every first preset time interval, wherein the first preset time interval may be preset by the system or a technician, for example, 1 second.

[0091] In some embodiments, the processor may determine the trigger threshold based on a baseline level and a safety margin, the safety margin being determined based on the noise stability value over a preset time period.

[0092] The preset time period refers to a fixed observation window used to calculate the noise stability value. It can be set by the system default or preset by technicians according to their needs.

[0093] A noise stability value is a numerical value used to measure the degree of fluctuation in ambient noise. In some embodiments, the noise stability value may be inversely proportional to the baseline level. For example, the processor may determine the noise stability value as the reciprocal of the variance of the baseline level.

[0094] Safety margin refers to a pre-set decibel offset to avoid false triggering by noise fluctuations. For example, +5dB.

[0095] In some embodiments, the processor may determine a safety margin based on the noise stability value over a preset time period.

[0096] For example, the processor can determine a safety margin based on the noise stability values ​​over a preset time period by querying a preset table. This preset table can include the correspondence between noise stability values ​​and safety margins, with the safety margin being negatively correlated with the noise stability values. For instance, in a highly stable environment with a stable baseline (e.g., a quiet room with only continuous air conditioning noise), a larger noise stability value allows for a smaller safety margin; conversely, in a low-stability environment with significant baseline fluctuations (e.g., a noisy exhibition venue), a smaller noise stability value allows for a larger safety margin. The preset table can be manually pre-built.

[0097] In some embodiments, the processor may determine the trigger threshold as the sum of the baseline level and the safety margin.

[0098] In some embodiments of this specification, a safety margin is determined based on the noise stability value, thereby enabling the trigger threshold to adaptively adjust with environmental changes. In environments with stable noise, the safety margin is reduced to ensure effective triggering of faint human voices; in scenarios with drastic noise fluctuations, the margin is automatically increased to suppress false triggering. Users do not need to manually adjust the trigger threshold, significantly improving sensitivity and robustness while reducing configuration costs.

[0099] In some embodiments, such as Figure 4 As shown, the processor can also determine the trigger threshold 460 based on historical trigger data 440 through parameter model 450. Parameter model 450 is a machine learning model. Historical trigger data 440 includes noise change rate 441 determined based on baseline level 430.

[0100] Historical trigger data refers to historical data related to trigger events collected within a preset historical time period. The preset historical time period refers to a pre-defined period preceding the current moment. For example, 60 seconds prior to the current moment.

[0101] In some embodiments, the processor can directly retrieve historical trigger data from the storage device.

[0102] In some embodiments, historical trigger data may include noise change rate.

[0103] The noise variation rate refers to the instantaneous fluctuation intensity that characterizes the baseline level of environmental noise.

[0104] In some embodiments, the noise change rate can be quantified by the first derivative of the baseline level with respect to time. For example, if the baseline level rises from -40 dB to -38 dB in 0.1 s, the noise change rate is +20 dB / s.

[0105] A parametric model is a model used to determine the trigger threshold. In some embodiments, the parametric model can be a machine learning model, such as any one or a combination of Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), etc.

[0106] In some embodiments, the input to the parameter model may include historical trigger data, and the output may include trigger thresholds.

[0107] In some embodiments, the parametric model can be trained using a large number of first training samples and first labels corresponding to the first training samples. In some embodiments, the processor can input multiple first training samples with first labels into the initial parametric model, construct a loss function using the first labels and the results of the initial parametric model, and iteratively update the parameters of the initial parametric model based on the loss function using gradient descent or other methods. When preset training conditions are met, the model training is complete, and a trained parametric model is obtained. These preset training conditions may include loss function convergence, the number of iterations reaching a threshold, etc.

[0108] The first training sample can be historical trigger data that only includes the rate of change of sample noise. The first label is the optimal historical trigger threshold corresponding to the first training sample.

[0109] In some embodiments, both the first training sample and the first label can be obtained based on historical experimental data. For example, for an experimental object (such as a historical audio file), the processor can use its noise change rate as the first training sample and perform multiple volume reduction triggering experiments on it using different historical triggering thresholds. Each volume reduction triggering experiment corresponding to a historical triggering threshold will produce a corresponding experimental effect; the historical triggering threshold corresponding to the best experimental effect is used as the first label corresponding to the first training sample.

[0110] The best experimental results can be evaluated manually or judged according to preset rules. For example, the highest sensitivity (the smallest trigger delay of the target speech, etc.) is considered.

[0111] In some embodiments of this specification, by introducing a machine learning model and combining noise change rate, false triggering rate, and missed triggering rate, the triggering threshold is finely adjusted in real time. The triggering threshold is automatically raised when there is noise and appropriately lowered when there is quiet, so as to achieve continuous self-optimization of the threshold by the system.

[0112] In some embodiments, historical trigger data may further include false trigger rate 442 and missed trigger rate 443.

[0113] The false trigger rate refers to the ratio of false triggers within a preset time period to the total number of triggers. A false trigger occurs when the system incorrectly identifies a non-target speech signal as the target speech and triggers a pitch reduction operation.

[0114] In some embodiments, the parameter model increases the safety margin in response to an increase in the false trigger rate.

[0115] Missed trigger rate refers to the ratio of missed triggers within a preset time period to the total number of triggers. Missed triggers occur when the target speech is not effectively recognized by the system and triggers a reduction in volume.

[0116] In some embodiments, the parameter model reduces the safety margin in response to an increase in the leak trigger rate.

[0117] In some embodiments, the processor can directly retrieve the false trigger rate and the missed trigger rate from the historical data stored in the storage device.

[0118] In some embodiments, the first training sample may further include the sample false trigger rate and the sample missed trigger rate.

[0119] In some embodiments, historical triggering data, including sample false triggering rate, sample missed triggering rate, and sample noise change rate, is used as the first training sample. The first label (optimal historical triggering threshold) corresponding to the first training sample can be obtained from historical experimental data. For example, the processor can perform multiple volume reduction triggering experiments on an experimental object (such as a historical audio file) using different historical triggering thresholds. Each volume reduction triggering experiment corresponding to a historical triggering threshold will produce a corresponding false triggering rate and missed triggering rate. The historical triggering data composed of these two rates and the noise change rate of the experimental object can be used as a first training sample. For this first training sample, historical triggering thresholds with false triggering rates less than the first threshold and missed triggering rates less than the second threshold can be used as their corresponding candidate first labels. The candidate first label with the smallest overall triggering error is selected as the first label.

[0120] The first and second thresholds can be set by system default or preset by technical personnel. Trigger comprehensive error refers to the overall deviation of the system's triggering decisions. In some embodiments, trigger comprehensive error can be the sum or product of the false trigger rate and the missed trigger rate.

[0121] In some embodiments of this specification, by introducing false trigger rate and missed trigger rate as model inputs, the trigger threshold evolves in real time with actual usage feedback, significantly reducing false suppression and missed suppression, and eliminating the need for manual adjustment in the long term.

[0122] In some embodiments of this specification, by estimating the environmental noise baseline in real time and dynamically adjusting the safety margin, the system can automatically calculate the appropriate trigger threshold in quiet or noisy scenarios without requiring manual fine-tuning by the user, thus avoiding both false triggering and missed triggering.

[0123] Figure 5 This is an exemplary schematic diagram illustrating the determination of candidate signal intervals according to some embodiments of this specification.

[0124] In some embodiments, such as Figure 5 As shown, the processor performs real-time peak detection on the mixed audio signal to identify candidate signal intervals with energy spikes in the target frequency band of the mixed audio signal, which may also include steps 510-520.

[0125] Step 510: Determine the extreme point 512 of the energy change rate on the target frequency band 511.

[0126] For more information on the target frequency band, please refer to [link / reference]. Figure 3 Related descriptions.

[0127] The rate of change of energy characterizes the instantaneous change intensity of sound energy and can be represented by the energy difference of the sound signal in the target frequency band within adjacent time windows. For example, if the sound signal energy in the target frequency band is 20 units at t0 = 0 ms and 25 units at t1 = 10 ms, then the rate of change of energy is +5 units / 10 ms.

[0128] An extreme point is a local maximum point in a sequence of multiple energy change rates.

[0129] In some embodiments, the processor can slice the audio signal of the target frequency band to obtain multiple slices. The slicing method can be Short-Time Fourier Transform (STFT), etc. The processor calculates the energy change rate between adjacent slices. When the energy change rate is a local maximum, the end time point of the previous slice (i.e., the start point of the next slice) corresponding to that energy change rate is determined as the extreme point. The preset neighborhood threshold can be set by the processor by default or preset manually based on experience. A local maximum can refer to the maximum value among R consecutive energy change rates. The value of R can be preset manually. The energy change rate between adjacent slices can be represented by the ratio of the difference between the average energy of the next slice and the average energy of the previous slice, to the average duration of the two slices.

[0130] Step 520: Based on the extreme point 512 and the release threshold 513, determine the candidate signal interval 514.

[0131] The release threshold is the energy threshold used to determine the end of a candidate signal interval.

[0132] In some embodiments, the release threshold can be determined based on an extreme point. For example, the processor can determine the release threshold as the product of the sound signal energy corresponding to the extreme point and the attenuation coefficient. The attenuation coefficient can be preset manually.

[0133] In some embodiments, the release threshold may be determined based on a baseline level of ambient noise and a release offset value. More information on ambient noise can be found at [link to relevant documentation]. Figure 3 For more information on baseline levels, please refer to the relevant descriptions. Figure 4 Related descriptions.

[0134] The release offset value refers to the decibel offset set to prevent misjudging the end of the target speech due to a short pause. For example, +10dB.

[0135] In some embodiments, the release offset value can be determined based on energy fluctuation values. For example, the processor can calculate the variance of energy at multiple time points within a current preset time period and determine it as the energy fluctuation value; the release offset value is positively correlated with the energy fluctuation value. The duration of the current preset time period and the selection of time points can be preset manually.

[0136] In some embodiments, the release offset is related to the current duration.

[0137] The current duration refers to the time span from the starting point of the candidate signal interval to the current moment.

[0138] In some embodiments, the release offset is positively correlated with the current duration. For example, the release offset can be the product of an increment factor and the current duration. The increment factor is the decibel increment per unit time, which can be set by system default or manually preset, for example, 2 dB / s. For instance, if the current duration is 5 seconds and the increment factor is 2 dB / s, then the release offset is +10 dB.

[0139] In some embodiments, the processor may determine the release threshold as the sum of the baseline level and the release offset value.

[0140] In some embodiments of this specification, by increasing the release offset value with the current duration, the offset is automatically increased when long sentences are continuous, and decreased when short speech is short, thus avoiding false trailing and enhancing the state stability in complex dialogues.

[0141] In some embodiments of this specification, by setting the release threshold as the sum of the noise baseline and the offset, an energy buffer is formed, thereby automatically raising or lowering the speech termination threshold in the same environment, avoiding truncation or false recovery caused by a fixed threshold, and significantly enhancing the stability of system decision-making and speech integrity.

[0142] In some embodiments, the processor can use an extreme point as the starting point of a candidate signal interval and continuously monitor the sound signal energy within the target frequency band; in response to a decrease in energy that remains below a preset release threshold for a second preset time period, the processor determines that the current target speech has ended and uses the end point of the second preset time period as the end point of the candidate signal interval; a candidate signal interval is formed based on the starting and ending points of the candidate signal interval. The second preset time period and the preset release threshold can be set by default by the processor or preset manually based on experience.

[0143] For example, the second preset time period is 20ms, the preset release threshold is −35dB, and the extreme point is... =20ms, then =20ms was determined as the starting point of the candidate signal interval, and the processor monitored the energy within the target frequency band. =30ms to =Continuously decreased for 60ms, and from =Starting at 60ms, until =80ms, if it remains below -35dB for 20 consecutive ms, then =80ms is determined as the end point of the candidate signal interval, and the candidate signal interval is... =20ms to =80ms.

[0144] In some embodiments of this specification, by first locking the extreme point of the energy change rate on the target frequency band, and then using the release threshold as the end threshold, the extreme point is used as the starting point of the candidate signal interval. The candidate signal interval is accurately located on the target frequency band, and the interval is immediately closed when the energy drops back and remains below the release threshold. This achieves millisecond-level voice triggering accuracy in noisy or quiet environments, and significantly improves the mixing response speed and stability.

[0145] The embodiments in this specification are merely illustrative and not intended to limit the scope of this specification. Various modifications and alterations that can be made by those skilled in the art under the guidance of this specification remain within its scope.

[0146] Furthermore, certain features, structures, or characteristics in one or more embodiments of this specification may be appropriately combined.

[0147] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0148] In the event of any inconsistency or conflict between the descriptions, definitions, and / or terms used in the supplementary materials to this specification and the contents of this specification, the descriptions, definitions, and / or terms used in this specification shall prevail.

Claims

1. A mixing control method, characterized in that, The method includes: Acquire background audio signals, as well as mixed audio signals containing potential speech and environmental noise; Peak detection is performed on the mixed audio signal to identify candidate signal intervals with energy spikes within the target frequency band of the mixed audio signal; and In response to the energy characteristics of the candidate signal range exceeding the trigger threshold, the volume of the background audio signal is reduced.

2. The method as described in claim 1, characterized in that, The method further includes: Based on the background audio signal and the mixed audio signal, estimate the baseline level of the ambient noise; The trigger threshold is determined based on the baseline level.

3. The method as described in claim 2, characterized in that, Determining the trigger threshold based on the baseline level includes: The trigger threshold is determined based on historical trigger data through a parametric model, which is a machine learning model. The historical trigger data includes the noise change rate determined based on the baseline level.

4. The method as described in claim 1, characterized in that, The step of performing peak detection on the mixed audio signal to identify candidate signal intervals with energy spikes in the target frequency band of the mixed audio signal includes: Determine the extreme point of the rate of energy change in the target frequency band; The candidate signal range is determined based on the extreme points and the release threshold.

5. The method as described in claim 1, characterized in that, Reducing the volume of the background audio signal includes: Determine the rate of decline sequence; Based on the amplitude reduction sequence, the volume of the background audio signal is reduced.

6. A mixing control system, characterized in that, The system includes: The acquisition module is configured to acquire background audio signals, as well as mixed audio signals containing potential speech and environmental noise; The identification module is configured to perform peak detection on the mixed audio signal to identify candidate signal intervals with energy spikes in the target frequency band of the mixed audio signal; and The volume reduction module is configured to reduce the volume of the background audio signal in response to the energy characteristics of the candidate signal interval exceeding a trigger threshold.

7. The system as described in claim 6, characterized in that, The volume reduction module is further configured to: Based on the background audio signal and the mixed audio signal, estimate the baseline level of the ambient noise; The trigger threshold is determined based on the baseline level.

8. The system as described in claim 7, characterized in that, The volume reduction module is further configured to: The trigger threshold is determined based on historical trigger data through a parametric model, which is a machine learning model. The historical trigger data includes the noise change rate determined based on the baseline level.

9. The system as described in claim 6, characterized in that, The identification module is further configured to: Determine the extreme point of the rate of energy change in the target frequency band; The candidate signal range is determined based on the extreme points and the release threshold.

10. The system as described in claim 6, characterized in that, The volume reduction module is further configured to: Determine the rate of decline sequence; Based on the amplitude reduction sequence, the volume of the background audio signal is reduced.