Multi-voice converging type anti-interference bluetooth earphone translation system

By identifying and adjusting the sound source weights in real time, the dynamic range imbalance problem of the multi-voice convergence anti-interference Bluetooth headset translation system during abnormal sonic booms was solved, achieving stability and robustness of the audio processing link and ensuring the continuity and accuracy of the translation task.

CN120676282BActive Publication Date: 2026-06-12VISION INTELLIGENCE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VISION INTELLIGENCE CO LTD
Filing Date
2025-06-03
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing multi-voice convergence anti-interference Bluetooth headset translation systems may experience an imbalance in the dynamic range of the audio processing link when faced with sudden high-decibel shouts, physical collisions, or strong background noise. This can lead to weakened voice signals or system interruptions, affecting the continuity and accuracy of the translation task.

Method used

By identifying audio dynamic range imbalance in real time, the system employs audio dynamic state acquisition, feature preprocessing, abnormal feature quantification analysis, and sound source weight adaptive adjustment modules to adjust sound source priority and weight, suppress interference channels, and maintain the stability of the main speech channel.

Benefits of technology

It effectively ensures the continuity of translation tasks and the accuracy of recognition, significantly improves the practicality and interaction quality of the system in multi-speaker, multilingual interaction scenarios, and avoids translation interruption and speech recognition failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676282B_ABST
    Figure CN120676282B_ABST
Patent Text Reader

Abstract

The application discloses a multi-voice co-converging type anti-interference Bluetooth earphone translation system, and relates to the technical field of voice processing, comprising an audio dynamic state acquisition module, an audio feature preprocessing and data construction module, a dynamic range abnormal feature extraction and quantitative analysis module, an audio system state intelligent evaluation module and a sound source weight self-adaptive regulation and control module; the audio dynamic state acquisition module, during the execution of a real-time language translation task of the Bluetooth earphone, through a high-precision microphone array, a digital signal processing chip and an audio processing module integrated in the earphone, real-time monitoring and capturing audio dynamic range state feature data generated by each sound source channel in the running process. The application realizes local inhibition of an interference channel and stable maintenance of a main voice channel by real-time identification of an audio dynamic range imbalance state and self-adaptive adjustment of sound source weight, and effectively guarantees the continuity and recognition accuracy of a translation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice processing technology, specifically to a multi-voice convergence anti-interference Bluetooth headset translation system. Background Technology

[0002] "Multi-voice convergence anti-interference Bluetooth headset translation" refers to a Bluetooth headset system capable of simultaneously receiving and processing voice information from multiple speakers. During processing, it employs an intelligent interference suppression mechanism to ensure that each voice segment is clearly separated, accurately identified, and efficiently translated. Based on a multi-channel voice input fusion mechanism (convergence), the system automatically distinguishes the voice characteristics of different voice sources. It combines noise reduction algorithms and sound source localization technology for voice signal separation and interference filtering. An integrated multilingual real-time translation engine then rapidly translates the extracted voice and outputs audio / text. This system is widely applicable to multilingual, multi-speaker communication scenarios (such as international conferences, tourism, and multi-person conversations). The core of the system is to achieve an intelligent interactive experience that ensures "uninterrupted multi-speaker communication, translation of diverse languages, and controllable interference environments."

[0003] Existing technologies have the following shortcomings: In multi-speech convergence speech processing systems, automatic gain control (AGC) mechanisms are typically introduced to dynamically adjust audio gain in order to adapt to differences in speech intensity between different sound sources. However, in practical applications, when a sudden event (such as a high-decibel shout, physical collision, or strong background noise) causes an "abnormal sonic boom" signal with an instantaneous sound pressure level far exceeding that of other sound sources, the AGC module in the system may abnormally trigger a rapid gain compression response due to sensing extreme energy mutations, thereby causing an imbalance in the overall dynamic range of the audio processing link. During this process, in order to protect the system hardware or maintain stable output amplitude, AGC often synchronously lowers the gain of all input channels, resulting in a significant weakening or even complete submersion of speech signals within the normal sound pressure range. In severe cases, continuous energy peaks exceeding limits may trigger the system's channel protection mechanism or automatic input link shutdown mechanism, causing an interruption of the entire multi-speech acquisition and translation function. Once this problem occurs, it can easily lead to loss of voice information, translation errors, or communication interruption. Especially in multilingual, multi-speaker concurrent interaction environments (such as international conferences, medical consultations, emergency command, etc.), it will cause irreversible communication barriers and significantly affect the system's usability and security.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-voice convergence anti-interference Bluetooth headset translation system. By identifying the imbalance state of audio dynamic range in real time and adaptively adjusting the sound source weight, it achieves local suppression of interference channels and stable maintenance of the main voice channel, effectively ensuring the continuity of translation tasks and recognition accuracy, thereby solving the problems in the background art mentioned above.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a multi-voice convergence anti-interference Bluetooth headset translation system, comprising an audio dynamic state acquisition module, an audio feature preprocessing and data construction module, a dynamic range anomaly feature extraction and quantification analysis module, an audio system state intelligent evaluation module, and a sound source weight adaptive adjustment module:

[0007] The audio dynamic state acquisition module, during the real-time language translation task performed by the Bluetooth headset, monitors and captures the audio dynamic range state characteristic data generated by each sound source channel in real time through the high-precision microphone array, digital signal processing chip and audio processing module integrated inside the headset.

[0008] The audio feature preprocessing and data construction module preprocesses the raw audio dynamic range feature data acquired in real time, thereby establishing an efficient and reliable dataset to support subsequent feature engineering and intelligent analysis processes.

[0009] The dynamic range anomaly feature extraction and quantification analysis module applies feature engineering techniques to the preprocessed dataset to accurately extract key indicators characterizing the overall dynamic range imbalance of the audio processing link, and conducts in-depth comprehensive analysis of the extracted key indicators to quantify the severity of the overall imbalance of the audio processing link.

[0010] The audio system status intelligent assessment module takes the comprehensively analyzed indicators as feature vectors and inputs them into the pre-trained machine learning model. The model then performs a rapid and intelligent assessment of the current operating status of the audio processing system to determine whether the current audio processing system has entered a dynamic range imbalance state.

[0011] The sound source weight adaptive control module immediately adjusts the sound source priority when it detects a dynamic range imbalance in the audio processing system. It temporarily sets the sound source channel identified as having strong interference characteristics to "low priority" and weakens its weight in the mixed output. At the same time, it automatically increases the proportion of the remaining sound sources in the overall output to suppress the interference of sudden anomalies on the normal conversation channel.

[0012] Preferably, the specific steps for acquiring audio dynamic range state feature data during real-time language translation tasks performed by the Bluetooth headset are as follows:

[0013] The high-precision microphone array inside the headphones synchronously collects multi-channel raw speech signals from different directions or speakers to ensure coverage of all sound source information.

[0014] The acquired analog audio signal is input to the digital signal processing chip, where A / D conversion, signal enhancement, and preliminary noise reduction are performed to generate a computable digital audio stream.

[0015] The audio processing module analyzes the digital audio stream and extracts the state characteristics related to the dynamic range of the automatic gain control response parameters of each channel.

[0016] The feature data is cached in a time-series structure as the input basis for subsequent dynamic range imbalance recognition, speech interference judgment, and intelligent translation stability control.

[0017] Preferably, feature engineering techniques are applied to the preprocessed dataset to accurately extract key indicators characterizing the overall dynamic range imbalance of the audio processing link. The extracted indicators include the output amplitude change after the dynamic compressor response and the amplitude change rate between consecutive frames of the audio waveform. The output amplitude change after the dynamic compressor response and the amplitude change rate between consecutive frames of the audio waveform are comprehensively analyzed under the detection window to generate the channel dynamic compression imbalance index and the amplitude gradient mutation index, respectively. The severity of the overall imbalance of the audio processing link is determined by the channel dynamic compression imbalance index and the amplitude gradient mutation index.

[0018] Preferably, the specific steps for generating the channel dynamic compression imbalance index by comprehensively analyzing the output amplitude change after the dynamic compressor response within the detection window are as follows:

[0019] Within the detection window, for all valid sound source channels in the Bluetooth headset, the maximum compression gradient value of each channel in the output amplitude envelope curve after processing by the dynamic compressor is obtained. Based on the maximum compression gradient values ​​of all channels, a compression difference mapping matrix M between channels is constructed, defined as follows:

[0020]

[0021] In the formula, M i,j These are elements in the compression difference mapping matrix M, representing the normalized difference in compression response intensity between sound source channel i and sound source channel j, where i∈1,2,3,...,n, and n is the total number of sound source channels. G i G is the maximum compression gradient value of sound source channel i. j It is the maximum compression gradient value of the sound source channel j, max(G i G j ) is the normalization reference factor, which represents the maximum of the maximum compression gradient values ​​of source channel i and source channel j;

[0022] After obtaining the compression difference mapping matrix M, statistical analysis is performed on all channel combinations to calculate the channel dynamic compression imbalance index. The calculation expression is as follows:

[0023]

[0024] In the formula, CDCI is the dynamic compression imbalance index of the vocal tract. It is the total number of non-repeating channel pairs, σ(M) i,j ) is a nonlinear enhancement function, representing that for each M i,j The weight of medium-to-high degree differences is enhanced by applying nonlinear transformation. The calculation method is as follows: σ(x)=tanh(k·x), where k is the amplification coefficient used to adjust the response sensitivity to abrupt changes in differences, and tanh(x) is the hyperbolic tangent function.

[0025] Preferably, the specific steps for generating an amplitude gradient abrupt change index by comprehensively analyzing the amplitude change rate between consecutive frames of the audio waveform within a detection window are as follows:

[0026] Within the detection window, the amplitude features of each frame of audio signal are extracted. Based on the amplitude changes between consecutive frames, a sudden change response factor is generated, using the following formula:

[0027] ΔΦ a =tanh(λ·|A a+1 -A a | β )·sgn(A a+1 -A a )

[0028] In the formula, A a It is the amplitude characteristic of the audio signal in the a-th frame of the detection window, A a+1 It is the amplitude characteristic of the audio signal in the (a+1)th frame of the detection window, that is, the amplitude characteristic of the audio signal in the next frame. λ is the amplitude response sensitivity coefficient, β is the power exponent of the abrupt change response, and sgn(A a+1 -A a ) is the mutation direction preservation factor, ΔΦ a It is a mutation response factor;

[0029] After obtaining the mutation response factor ΔΦ of all adjacent frames a Then, the acquired response values ​​are mapped to the nonlinear enhancement domain. Through amplification and integration, a quantitative index characterizing the sum of abrupt change in intensity within the entire detection window is formed, namely the amplitude gradient abrupt change index. The calculation expression is as follows:

[0030]

[0031] In the formula, AVGI is the amplitude gradient mutation exponent, γ is the mutation response amplification factor, e is the natural base, m is the number of frames contained in the detection window, and Z is the normalization constant.

[0032] Preferably, the channel dynamic compression imbalance index and amplitude gradient mutation index, which have been comprehensively analyzed, are input as feature vectors into a pre-trained machine learning model. The model generates a link dynamic range imbalance coefficient, and the current operating status of the audio processing system is quickly and intelligently evaluated based on the link dynamic range imbalance coefficient to determine whether the current audio processing system has entered a dynamic range imbalance state.

[0033] Preferably, the link dynamic range imbalance coefficient generated by the pre-trained machine learning model during the rapid and intelligent evaluation of the current operating status of the audio processing system is compared and analyzed with a pre-set reference threshold for the link dynamic range imbalance coefficient to determine whether the current audio processing system has entered a dynamic range imbalance state. The determination logic is as follows:

[0034] If the link dynamic range imbalance coefficient is greater than the preset reference threshold for the link dynamic range imbalance coefficient, the current audio processing system is determined to have entered a dynamic range imbalance state; if the link dynamic range imbalance coefficient is less than or equal to the preset reference threshold for the link dynamic range imbalance coefficient, the current audio processing system is determined not to have entered a dynamic range imbalance state.

[0035] Preferably, when a dynamic range imbalance is detected in the audio processing system, the sound source priority is immediately adjusted, and the sound source channel identified as having strong interference characteristics is temporarily set to "low priority" and its weight in the mixed output is weakened. At the same time, the proportion of the remaining sound sources in the overall output is automatically increased. The specific steps are as follows:

[0036] After identifying a dynamic range imbalance in the audio processing system, the interference intensity of all sound source channels in the system is first assessed. By calculating the interference index of each channel, its impact on the imbalance is quantified. The expression for calculating the interference index is as follows:

[0037] ψ i =α·ΔA i +β·δH i

[0038] In the formula, ψ i It is the interference index, representing the comprehensive interference intensity score of the dynamic range imbalance caused by the i-th sound source channel, ΔA. i It is the amplitude abrupt change intensity, representing the maximum rate of change of the speech signal amplitude between adjacent frames of sound source channel i, δH i It is the gain offset, and α is the amplitude abrupt change weighting coefficient;

[0039] After obtaining the interference index ψ of all channels i Then, the weights of each sound source channel in the mixed output are dynamically adjusted to weaken the output contribution of strong interference sources and increase the proportion of the remaining channels. The weight adjustment formula is as follows:

[0040]

[0041] In the formula, W i It represents the source channel mixing output weight, indicating the weight value of the i-th source channel in the current audio mixing output. d is the average interference index of all sound source channels, e is the natural base, and Λ is the link dynamic range imbalance coefficient. t It is the reference threshold for the link dynamic range imbalance coefficient;

[0042] After adjusting the weights of each channel, to ensure the consistency of amplitude and the accuracy of signal proportions in the mixed output, the weights W of the source channel mixed output are adjusted. i Normalization is performed, and the mixed output signal is reconstructed based on the normalization weights. The calculation formula is as follows:

[0043]

[0044] In the formula, Y min It is the mixing output signal, W j This represents the weight value of the j-th audio source channel in the current audio mix output, where n is the total number of audio source channels, and S... i It is the original speech signal of the i-th sound source channel.

[0045] The technical effects and advantages provided by the present invention in the above technical solution are as follows:

[0046] This invention effectively achieves real-time identification and adaptive control of the dynamic range imbalance of the audio processing link by introducing key modules such as audio dynamic state perception, feature preprocessing, abnormal feature quantification analysis, intelligent state evaluation, and adaptive weight adjustment. This significantly improves the system's stability and robustness in complex multi-sound source environments. The system constructs a high-quality feature dataset through audio dynamic state acquisition and preprocessing, and intelligently evaluates the audio state using feature engineering and machine learning models. When sonic booms or abnormal sound source disturbances occur, it can quickly identify the interference source and dynamically adjust its weight, preventing its interference from spreading to the global AGC mechanism and ensuring the continuity and translatability of other normal speech channels. Compared to the shortcomings of traditional systems that are susceptible to high-energy interference and overall failure, this system implements a differentiated control strategy of "local suppression of interference channels and stable maintenance of the main speech link," effectively avoiding problems such as translation task interruption and speech recognition failure. This significantly improves the practicality and interaction quality of Bluetooth headsets in multi-speaker, multi-language interaction scenarios. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0048] Figure 1 This is a schematic diagram of a multi-voice convergence anti-interference Bluetooth headset translation system according to the present invention. Detailed Implementation

[0049] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0050] This invention provides, for example Figure 1 The multi-voice convergence anti-interference Bluetooth headset translation system shown includes an audio dynamic status acquisition module, an audio feature preprocessing and data construction module, a dynamic range anomaly feature extraction and quantitative analysis module, an audio system status intelligent evaluation module, and a sound source weight adaptive adjustment module.

[0051] The audio dynamic state acquisition module, during the real-time language translation task performed by the Bluetooth headset, monitors and captures the audio dynamic range state characteristic data generated by each sound source channel in real time through the high-precision microphone array, digital signal processing chip (DSP) and audio processing module integrated inside the headset.

[0052] During real-time language translation tasks performed by the Bluetooth headset, the system relies on a high-precision microphone array integrated within the headset to pick up raw speech signals from different directions and speakers. Combined with a digital signal processing chip (DSP) and a built-in audio processing module, these speech signals are analyzed in real time to extract operational status characteristic data related to the audio dynamic range. Specifically, the acquired data includes: short-time energy, instantaneous peak amplitude, average sound pressure level (SPL), signal-to-noise ratio (SNR), automatic gain control (AGC) parameter changes, and dynamic compression ratio for each sound source channel. This data reflects the system's processing behavior and response characteristics to changes in the intensity of the input speech signal at different time points. Its core function is to provide crucial evidence for subsequent judgment of whether abnormal interference (such as sonic booms) or audio link dynamic range imbalance exists, thus laying a data foundation for intelligent gain adjustment and stable execution of the translation task.

[0053] The core function of this step is to provide accurate raw data support for subsequent sound source interference analysis, ensuring that subsequent diagnosis and control are based on real and comprehensive audio operating status data, thus guaranteeing the authenticity and comprehensiveness of the data from the source.

[0054] The specific steps for acquiring audio dynamic range state feature data during real-time language translation tasks performed by Bluetooth headsets are as follows:

[0055] The first step is to simultaneously collect multi-channel raw speech signals from different directions or speakers through a high-precision microphone array installed inside the headphones, ensuring coverage of all sound source information;

[0056] The second step involves inputting the acquired analog audio signal into a digital signal processing chip (DSP) to perform A / D conversion, signal enhancement, preliminary noise reduction, and other processing to generate a computable digital audio stream.

[0057] The third step involves the audio processing module analyzing the digital audio stream and extracting dynamic range-related state characteristics for each channel, such as short-time energy, peak amplitude, average sound pressure level, and automatic gain control (AGC) response parameters.

[0058] The fourth step involves caching the aforementioned feature data in a time-series structure, serving as the input basis for subsequent dynamic range imbalance recognition, speech interference judgment, and intelligent translation stability adjustment. This process enables real-time quantitative perception of the speech environment in which the headphones are located and dynamic representation of the system's operating status.

[0059] The audio feature preprocessing and data construction module preprocesses the raw audio dynamic range feature data acquired in real time, thereby establishing an efficient and reliable dataset to support subsequent feature engineering and intelligent analysis processes.

[0060] Preprocessing of the raw audio dynamic range feature data acquired in real time mainly includes the following specific steps: First, data denoising is performed by using filters (such as median filtering and Kalman filtering) or adaptive algorithms to remove background noise and instantaneous spike interference from the feature data, thereby improving data stability. Second, missing data interpolation is performed to complete the missing data by using linear interpolation, spline interpolation, or time series prediction models (such as ARIMA) to repair feature value loss caused by instantaneous packet loss or hardware jitter during the acquisition process, ensuring data continuity. Then, data normalization or standardization is performed to uniformly map feature data of different dimensions or amplitude ranges to a fixed range (such as [0, 1] or a mean of 0 and a standard deviation of 1), avoiding feature bias from affecting the model learning effect. Further, time window segmentation is performed, dividing the continuously acquired data into time segments according to a sliding window or fixed frame length, which facilitates the capture of short-term dynamic trends and the construction of time series feature vectors. Finally, outlier detection and removal, as well as correlation analysis between features, are performed to identify extreme outliers and remove redundant features or construct derived indices based on correlation evaluation, thereby improving the quality of feature expression and the model's discriminative ability. The core function of this preprocessing process is to transform the raw feature data from a mixed, discrete, and noisy state into a clear, continuous, and standardized analytical input, providing a high-quality data foundation for subsequent feature extraction, model training, and state recognition.

[0061] By systematically preprocessing real-time acquired audio dynamic range feature data (such as denoising, completion, normalization, and slicing), the originally messy and flawed raw data is transformed into a dataset with complete structure, stable temporal sequence, and strong feature consistency. This dataset possesses high accuracy, high continuity, and high expressiveness, and can serve as standard input for subsequent algorithmic processing. Its main function is to provide a solid data foundation for feature engineering, ensuring that the extracted key indicators have physical rationality and statistical stability. It also provides high-quality input for the training, evaluation, and inference of machine learning models, state recognition algorithms, or translation strategy optimization models, enabling the entire audio processing and intelligent recognition system to have stronger robustness, accuracy, and adaptability. In other words, this dataset is a crucial "information foundation" in the "intelligent analysis" chain.

[0062] The key role of this step is to improve the quality and analyzability of the data by eliminating clutter and interference, thereby establishing an efficient and reliable dataset to support subsequent feature engineering and intelligent analysis processes.

[0063] The dynamic range anomaly feature extraction and quantification analysis module applies feature engineering techniques to the preprocessed dataset to accurately extract key indicators characterizing the overall dynamic range imbalance of the audio processing link, and conducts in-depth comprehensive analysis of the extracted key indicators to quantify the severity of the overall imbalance of the audio processing link.

[0064] Feature engineering techniques are applied to the preprocessed dataset to accurately extract key indicators characterizing the overall dynamic range imbalance of the audio processing link. The extracted indicators include the output amplitude change after the dynamic compressor response and the amplitude change rate between consecutive frames of the audio waveform. The output amplitude change after the dynamic compressor response and the amplitude change rate between consecutive frames of the audio waveform are comprehensively analyzed under the detection window to generate the channel dynamic compression imbalance index and the amplitude gradient mutation index. The severity of the overall imbalance of the audio processing link is determined by the channel dynamic compression imbalance index and the amplitude gradient mutation index.

[0065] Abnormal changes in the output amplitude of a dynamic compressor after its response directly indicate an imbalance in the overall dynamic range of the audio processing link. This is because the dynamic compressor, as the core unit in an audio system used to control the amplitude changes of the input signal, largely reflects the system's ability to adapt to and adjust to the energy differences of different sound sources. Under normal conditions, the dynamic compressor makes smooth and gradual gain adjustments based on the input intensity of the sound source to ensure that all channels operate within a unified and reasonable dynamic range. However, when problems such as sudden sonic booms, abnormal enhancement of a single sound source, or abnormal system feedback occur, the compressor may produce a drastic nonlinear response to some channels, manifesting as a sudden drop in output amplitude, over-compression, or adjustment lag. These abnormal changes not only affect the speech quality of the compressed channel but also affect the overall dynamic behavior of other sound sources through the converging link, disrupting the dynamic consistency between channels and causing a system-level dynamic range imbalance. Therefore, abnormal changes in the output amplitude of the dynamic compressor are not just a problem of a single channel but one of the important external manifestations of overall link imbalance, possessing high system-level representativeness and diagnostic value.

[0066] The specific steps for generating the channel dynamic compression imbalance index by comprehensively analyzing the output amplitude change after the dynamic compressor response within the detection window are as follows:

[0067] Within the detection window, for all valid sound source channels in the Bluetooth headset, the maximum compression gradient value of each channel in the output amplitude envelope curve after processing by the dynamic compressor is obtained. This value represents the compression response intensity of the channel within the current detection window, that is, the maximum amplitude change rate produced by the compressor on that channel. Based on the maximum compression gradient values ​​of all channels, a compression difference mapping matrix M between channels is constructed, defined as follows:

[0068]

[0069] In the formula, M i,jThese are elements in the compression difference mapping matrix M, representing the normalized difference in compression response intensity between sound source channel i and sound source channel j. This value quantifies the degree of compression inconsistency between any two channels. A larger value indicates a greater difference in compression response between channels, suggesting inconsistent control of multiple sound sources and a potential risk of imbalance. i∈1, 2, 3, ..., n, where n is the total number of sound source channels. G i G represents the maximum compression gradient value of sound source channel i, indicating the maximum rate of change of amplitude (i.e., the maximum compression response slope) in the output amplitude curve of the audio signal after processing by the dynamic compressor within the detection time window. j is the maximum compression gradient value of source channel j, used for differential comparison with channel i to identify the degree of inconsistency in the compression response between the two, max(G i G j ) is the normalization reference factor, which represents the maximum of the maximum compression gradient values ​​of sound source channel i and sound source channel j, achieving normalization processing so that the difference G i -G j Compare with its reference amplitude to avoid amplifying the essential slight differences due to a higher absolute amplitude of the channel.

[0070] This mapping matrix comprehensively considers the compression inconsistencies between channels and eliminates the influence of absolute amplitude values ​​through normalization, thus focusing on the structural manifestation of "relative differences." This matrix provides the basic data support for the subsequent generation of the channel compression imbalance index.

[0071] After obtaining the compression difference mapping matrix M, statistical analysis is performed on all channel combinations to calculate the channel dynamic compression imbalance index. The calculation expression is as follows:

[0072]

[0073] In the formula, CDCI is the channel dynamic compression imbalance index, which measures the degree of difference in response between multiple sound source channels under dynamic compressor processing. The value ranges from [0, 1]. The larger the value, the more significant the difference in compression behavior between channels, and the more likely the system is to be in a state of dynamic range imbalance. It is the total number of non-repeating channel pairs (normalization factor), representing the number of combinations for comparing any two channels out of n channels, σ(M i,j ) is a nonlinear enhancement function, representing that for each M i,j The weight of medium-to-high degree differences is enhanced by applying nonlinear transformation. The calculation method is as follows: σ(x)=tanh(k·x), where k is the amplification coefficient used to adjust the response sensitivity to abrupt changes in differences, and tanh(x) is the hyperbolic tangent function.

[0074] Nonlinear augmentation functions are mathematical functions used in feature processing or model computation to perform nonlinear transformations on input data. Their purpose is not simple scaling, but rather differentiated amplification or suppression based on different input amplitude ranges. Their core function is to enhance the influence of medium-to-high amplitude features on the model or index, while suppressing the interference effects of low-amplitude changes, thereby improving the system's sensitivity to identifying "abnormal states" or "critical imbalances." In dynamic range imbalance identification, small-amplitude compression differences may be normal fluctuations within the system and require no intervention; however, medium-to-high amplitude differences often indicate a loss of control in the system's compression mechanism or sonic booms. Therefore, nonlinear functions are needed to amplify these differences, making them more dominant in the weight calculation. Commonly used nonlinear augmentation functions include the hyperbolic tangent function tanh(x) and the exponential function 1-e^(-x). -x Sigmoid function Among them, tanh(x) has good discriminative power and continuity, and is often used to smooth and enhance the response of medium and high values, while avoiding excessive amplification of low values. It is widely used in scenarios such as audio imbalance detection, image edge enhancement, and neural network activation.

[0075] By aggregating and normalizing the difference values ​​of all non-repeating channel pairs in the matrix, the Channel Dynamic Compression Imbalance Index (CDCI) can effectively characterize the consistency of the system's multi-channel compression response within the current detection window. An increase in the CDCI indicates a greater difference in compression behavior among multiple channels, suggesting a significant dynamic range imbalance in the audio processing link. Conversely, a lower CDCI value indicates that the compression control of each channel is basically coordinated, and the audio processing link is in a stable and balanced state.

[0076] The output amplitude changes after the dynamic compressor's response are comprehensively analyzed within a detection window to generate a channel dynamic compression imbalance index. If the dynamic compressor's output amplitude response to each sound source channel differs significantly—that is, some channels are drastically compressed while others are not significantly affected—it will lead to inconsistent compression behavior between channels, resulting in a significant increase in the channel dynamic compression imbalance index. At this point, the system's internal dynamic range control strategy is unable to coordinate and uniformly adjust multiple sound sources, reflecting a structural imbalance in the audio link. Therefore, the larger the channel dynamic compression imbalance index, the more it indicates unbalanced compressor control, disordered overall system dynamic range, and a threat to recognition and translation stability. Conversely, if the C-channel dynamic compression imbalance index is small, it means the compressor's response to all sound source channels is relatively consistent, the system is in a balanced adjustment state, the dynamic range is within a reasonable control range, and no imbalance has occurred.

[0077] When the rate of amplitude change between consecutive frames of an audio waveform increases sharply within a short period, it usually indicates an abnormal state of overall dynamic range imbalance in the audio processing link. This drastic fluctuation means that the system's automatic gain control (AGC) or dynamic compressor has been abnormally triggered by a strong interfering sound source (such as a sonic boom or sudden high sound pressure level), causing the gain adjustment of the audio signal to become unstable, leading to nonlinear amplitude jumps in the speech signal between frames. Under normal conditions, the rate of amplitude change between consecutive frames should remain within a smooth range, reflecting the system's consistent and stable response to different sound sources. Sudden surges in amplitude usually mean that the system attempted to suppress high-energy interference but failed to synchronously regulate other sound sources, disrupting the dynamic balance of multiple channels in the link. This abnormal fluctuation not only disrupts the continuity and recognizability of speech but also affects the accuracy of mixing output and semantic extraction, and is a typical sign that the audio link has entered a "non-steady-state imbalance zone." Therefore, a sharp increase in the rate of amplitude change between consecutive frames is an important dynamic behavioral characteristic for identifying dynamic range imbalance.

[0078] The specific steps for generating the amplitude gradient abrupt change index by comprehensively analyzing the amplitude change rate between consecutive frames of an audio waveform within a detection window are as follows:

[0079] Within the detection window, the amplitude features (such as short-time energy amplitude or logarithmic amplitude) of each frame of audio signal are extracted. Based on the amplitude changes between consecutive frames, a mutation response factor is generated to sensitively capture the degree and direction of mutations between adjacent frames. The generation formula is as follows:

[0080] ΔΦ a =tanh(λ·|A a+1 -A a | β )·sgn(A a+1 -A a )

[0081] In the formula, A a This detects the amplitude characteristics of the audio signal in frame a of the detection window. Optional amplitude characteristics include: Short-Time Energy (STE): the sum of the squared values ​​of the signal in each frame; Logarithmic Energy: taking the logarithm of the energy in each frame to compress the dynamic range; Frequency Domain Amplitude Peak or Envelope: used to emphasize non-stationary segments. a+1 This refers to the amplitude characteristics of the audio signal in the (a+1)th frame of the detection window, i.e., the amplitude characteristics of the audio signal in the next frame. λ is the amplitude response sensitivity coefficient, used to control the amplification degree of amplitude difference input before entering the tanh nonlinear function, with a value range of 1.5-3. β is the abrupt response power exponent, used to nonlinearly enhance the asymmetry of amplitude changes, with a value range of 1.2-2. sgn(A a+1 -A a) is the mutation direction preservation factor (sign function), which preserves the "directional information" of the mutation and is used to distinguish whether the system is subjected to "surge disturbance" or "decay disturbance". Used in conjunction with tanh and the power exponent, it helps to construct a complete "mutation scene feature map". ΔΦ a It is a mutation response factor, with a value range of (-1 < ΔΦ). i <1), if ΔΦ i →+1 indicates a strong abrupt increase in amplitude between the frame pairs; if ΔΦ i →-1 indicates a strong "decline" abrupt change; if ΔΦ i ≈0: Indicates that the inter-frame changes are smooth and the system is stable;

[0082] tanh is the hyperbolic tangent function with an output range of (-1, 1). It is an sigmoid, continuous, differentiable, nonlinear function that is approximately linear with small input values, gradually saturating and approaching ±1 with larger input values. The core function of tanh in audio dynamics analysis is to nonlinearly compress sudden amplitude changes, ensuring a linear response to small changes and rapid saturation for large changes. This avoids numerical "explosions" caused by abnormal spikes (such as sonic booms) dominating the overall judgment, while still maintaining sensitivity to medium-amplitude abrupt changes. Compared to functions like ReLU and sigmoid, tanh retains positive and negative symmetry and boundedness, making it suitable for handling positive and negative amplitude abrupt changes between consecutive frames. It suppresses extreme values ​​without losing directionality, making it an ideal compression function for characterizing dynamic range abrupt changes.

[0083] The sign function (denoted as sgn(x)) is a fundamental piecewise function in mathematics used to determine the sign of a real number. Its function is to preserve the direction of numerical change while ignoring the magnitude of the change. It is commonly used in signal processing, mathematical modeling, or control systems to identify trends of increase, stagnation, or decrease in values. In this application scenario, sgn(A a+1 -A a The sign function represents the "directional difference" between the amplitudes of two consecutive audio signal frames: if the amplitude of the next frame is greater than that of the current frame, it is +1, indicating a sudden increase; if it is smaller, it is -1, indicating a sudden decrease; if they are equal, it is 0, indicating no change. The main purpose of choosing the sign function is to retain the upward or downward trend information of the inter-frame change when calculating the amplitude abrupt change exponent. This allows for the identification of not only "whether there is a sudden change" but also "the direction of the change" in subsequent judgments. This is crucial for judging compression and burst interference in dynamic range imbalance.

[0084] By constructing a nonlinearly enhanced mutation response factor, the amplitude mutation behavior and its directional characteristics between consecutive frames of an audio waveform are accurately captured. This step provides highly sensitive, direction-aware basic feature inputs for subsequent determination of the severity of dynamic range imbalance.

[0085] After obtaining the mutation response factor ΔΦ of all adjacent frames a Then, the acquired response values ​​are mapped to the nonlinear enhancement domain. Through amplification and integration, a quantitative index characterizing the sum of abrupt change in intensity within the entire detection window is formed, namely the amplitude gradient abrupt change index. The calculation expression is as follows:

[0086]

[0087] In the formula, AVGI is the amplitude gradient abrupt change index, which represents the overall strength of the amplitude abrupt change trend between audio frames within the detection window. It is used to determine whether there is a dynamic range imbalance and takes a value between 0 and 1. γ is the abrupt change response amplification coefficient, which controls the abrupt change response factor |ΔΦ. a In the exponential mapping function, the "sensitivity amplification" ranges from 1 to 5, where e is the natural base, m is the number of frames contained in the detection window, and Z is a normalization constant used to normalize the accumulated result, ensuring that the final output value falls within a controllable range (such as [0,1], [0,100]), preventing AVGI from being affected by the length of the detection window, and ensuring that the output has comparability and a stable numerical range, suitable for different time scales and model input specifications.

[0088] In the AVGI index, the exponential mapping function refers to... It is a nonlinear enhancement function constructed based on the natural exponential function, and its core function is to enhance the inter-frame abrupt response factor |ΔΦ. a Mapping to the (0,1) interval makes the system highly sensitive to large mutations and relatively suppressive of small fluctuations. When the mutation value is small, the exponential term approaches 1, and the activation value approaches 0; when the mutation value is large, the exponential term rapidly approaches 0, and the activation value approaches 1, thus amplifying the expression of abnormal mutation behavior. This function is an inverse transformation of the classic exponential decay function, commonly found in signal processing and neural network activation functions. It can significantly improve the system's ability to discriminate dynamic range imbalance states and is a key enhancement mechanism in AVGI construction.

[0089] The larger the amplitude gradient abrupt change index, generated by comprehensively analyzing the amplitude change rate between consecutive frames of an audio waveform within a detection window, the more severe the amplitude fluctuations between adjacent frames and the increased waveform discontinuity. This typically reflects that the system experienced abnormal sound intensity interference (such as sonic booms or severe signal disturbances) within a certain time period, triggering a nonlinear response in the gain control mechanism and compressor adjustment chaos, thereby disrupting the amplitude balance between channels and leading to a chain reaction of dynamic range imbalance. Therefore, the larger the amplitude gradient abrupt change index, the stronger the system's imbalance trend. Conversely, if the amplitude gradient abrupt change index is at a low stability level, it indicates that the amplitude changes between audio frames are stable and the system's dynamic compression consistency is good. In this case, the audio processing link is operating normally and has not yet experienced imbalance.

[0090] The audio system status intelligent assessment module takes the comprehensively analyzed indicators as feature vectors and inputs them into a pre-trained machine learning model (such as random forest, XGBoost, CNN or LSTM network). The model performs a fast and intelligent assessment of the current operating status of the audio processing system and determines whether the current audio processing system has entered a dynamic range imbalance state.

[0091] The comprehensive analysis of the channel dynamic compression imbalance index and amplitude gradient mutation index is used as feature vectors and input into a pre-trained machine learning model (such as random forest, XGBoost, CNN or LSTM network). The model generates a link dynamic range imbalance coefficient. Based on the link dynamic range imbalance coefficient, the current operating status of the audio processing system is quickly and intelligently evaluated to determine whether the current audio processing system has entered a dynamic range imbalance state.

[0092] "Pre-trained machine learning models" refer to discriminative models built and trained in advance by developers before the deployment of audio processing systems. These models are based on a large amount of historical audio processing data, simulated interference samples, or real user interaction data, and are used to identify the dynamic state of the audio link. This model learns from massive amounts of feature data (such as the duct dynamic compression imbalance index and amplitude gradient mutation index) to extract potential feature patterns and classification boundaries, thereby enabling it to identify whether the audio link is in a state of normal, mild, or severe dynamic imbalance. During training, model structures such as Random Forest, XGBoost, Convolutional Neural Networks (CNN), or Long Short-Term Memory (LSTM) networks are typically used. Supervised learning is performed using labeled training datasets, and through continuous iterative optimization, the model masters the nonlinear mapping relationship between the imbalance state and the feature data. This training process is usually completed offline on high-performance computing platforms, ultimately outputting a "deployable model" that has undergone parameter freezing and generalization capability verification for subsequent online inference tasks.

[0093] In actual operation, this model does not require retraining but is invoked in real-time as an "embedded intelligent decision-making module" within the Bluetooth headset audio processing system. When the Channel Dynamic Compression Imbalance Index (CDCI) and Amplitude Gradient Abrupt Change Index (AVGI) are calculated in real-time by the system, they are packaged into feature vectors and input into the pre-trained model. The model immediately outputs a link dynamic range imbalance coefficient reflecting the current system state. This coefficient is a quantitative evaluation value; a threshold can be set to determine whether the system is currently in an imbalanced state. For example, an imbalance coefficient greater than 0.7 may indicate "high risk," while less than 0.3 indicates a "stable state," serving as the basis for subsequent intervention decisions (such as source weight adjustment, gain strategy switching, etc.). Because the model has learned various complex scenarios during training, including different user voice styles, background noise types, and interference intensity, it possesses strong robustness and scene adaptability. In summary, the introduction of this pre-trained model enables the Bluetooth headset system to possess data-driven intelligent state perception and response capabilities, significantly improving the stability and self-regulation level of the audio link in dynamic and complex environments.

[0094] The link dynamic range imbalance coefficient generated by the pre-trained machine learning model during the rapid and intelligent evaluation of the current operating status of the audio processing system is compared with a pre-set reference threshold for the link dynamic range imbalance coefficient to determine whether the current audio processing system has entered a dynamic range imbalance state. The determination logic is as follows:

[0095] If the link dynamic range imbalance coefficient is greater than the preset reference threshold for the link dynamic range imbalance coefficient, the current audio processing system is determined to have entered a dynamic range imbalance state; if the link dynamic range imbalance coefficient is less than or equal to the preset reference threshold for the link dynamic range imbalance coefficient, the current audio processing system is determined not to have entered a dynamic range imbalance state.

[0096] The sound source weight adaptive control module immediately adjusts the sound source priority when it detects a dynamic range imbalance in the audio processing system. It temporarily sets the sound source channel identified as having strong interference characteristics to "low priority" and weakens its weight in the mixed output. At the same time, it automatically increases the proportion of the remaining sound sources in the overall output to suppress the interference of sudden abnormalities on the normal conversation channel.

[0097] Upon detecting a dynamic range imbalance in the audio processing system, a dynamic adjustment mechanism for sound source priorities is immediately executed. Its core function is to quickly block the destructive impact of abnormal sound sources on the overall audio processing chain and stably maintain the output quality of normal speech signals and system availability. In multi-source converged Bluetooth headset voice translation systems, each sound source channel operates collaboratively within the same processing framework, and the system typically assigns equal weight to each channel by default. However, in sudden abnormal situations, such as sonic booms, loud noises, or sharp non-speech interference from a sound source, its high-amplitude signal may be misinterpreted by the system as an overall energy increase, inducing abnormal triggering of automatic gain control (AGC) or an over-response of the dynamic compression mechanism. This suppresses the output amplitude of all channels, causing speech recognition failure, translation interruption, or system protective shutdown.

[0098] To address this, this step uses an intelligent identification mechanism to mark abnormal sound source channels as "low priority" and proactively reduce their weight in the mixing output. This operation is equivalent to "unmasking suppression" of the channel—instead of directly cutting off its input data stream, it reduces its participation and influence on the system's decision-making process, thereby preventing it from continuing to drive AGC responses, polluting the mixing output, or interfering with semantic judgment. Simultaneously, the system automatically increases the output proportion of other sound source channels in a normal state, strengthening the signal expression of the main semantic channel at the mixing level, ensuring that the system's speech recognition, machine translation, and speech synthesis modules can still receive sufficiently clear and decodeable audio information. This dynamic sound source control strategy of "suppressing interference + enhancing main semantics" not only prevents the system from falling into a state of overall suppression collapse but also significantly improves the fault tolerance and robustness of speech understanding in complex scenarios.

[0099] In summary, the fundamental purpose of this step is to introduce adaptive intervention logic based on sound source behavior characteristics. Without interrupting all channels, it reduces the weight of high-risk channels and preserves the semantics of low-risk channels, achieving a "partial isolation, global stability" operating strategy for the system in highly complex, multi-interference, and multi-speaker scenarios. This mechanism is not only applicable to handling sudden interference but also provides operational space for subsequent modules such as speech signal reconstruction, channel recovery, and self-learning optimization. It is a crucial disaster recovery and self-healing control link in the entire audio processing system.

[0100] Once a dynamic range imbalance is detected in the audio processing system, the sound source priorities are immediately adjusted. Sound source channels identified as having strong interference characteristics are temporarily set to "low priority" and their weight in the mixed output is weakened. Simultaneously, the proportion of other sound sources in the overall output is automatically increased. The specific steps are as follows:

[0101] After identifying a dynamic range imbalance in the audio processing system (i.e., the link dynamic range imbalance coefficient is greater than a preset threshold), the interference intensity of all sound source channels in the audio processing system is first assessed. By calculating the interference index of each channel, its impact on the imbalance state is quantified. The expression for calculating the interference index is as follows:

[0102] ψ i =α·ΔA i +β·δH i

[0103] In the formula, ψ i This is the interference index, representing the comprehensive interference intensity score caused by the i-th sound source channel to disrupt the dynamic range. It serves as a direct criterion for determining whether a sound source is a "strong interference source." The larger the value, the greater the impact of the channel on the current imbalance state. ΔA i It is the amplitude abrupt change intensity, representing the maximum rate of change of the speech signal amplitude between adjacent frames of sound source channel i, δH i This is the gain offset, representing the absolute magnitude of the gain change before and after the gain control module (such as AGC or DRC) of sound source channel i. If a sound source triggers a strong response to AGC, its gain will exhibit a significant jump, δH. i The value will also increase accordingly. It is an adjustment offset characteristic index for identifying abnormal system gain. α is the amplitude mutation weighting coefficient, which controls the amplitude mutation characteristic ΔA. i The contribution of the interference index is set by the system design parameters and can be configured based on model training or experience. For example, α = 0.6 indicates that the mutation feature is dominant, which is used to flexibly adjust the system's sensitivity to "acoustic slope interference".

[0104] The interference index reflects whether there are sudden strong interference signals in each channel, providing a quantitative basis for subsequent priority control.

[0105] After obtaining the interference index ψ of all channels i Then, the weights of each sound source channel in the mixed output are dynamically adjusted to weaken the output contribution of strong interference sources and increase the proportion of the remaining channels. The weight adjustment formula is as follows:

[0106]

[0107] In the formula, W i This represents the output weight of the i-th audio source channel in the current audio mix output. It directly determines the proportion of the speech signal from that channel in the final system output, and its value ranges from 0 to 1. The larger the value, the more prominent the channel; the smaller the value, the more its output will be suppressed. It is used to dynamically adjust the output contribution of the channel, reduce the output influence of interfering sound sources, and improve the overall recognition stability of the system under dynamic range imbalance. is the average interference index of all sound source channels, used to construct a baseline value so that the interference intensity of each channel can be evaluated in a "center-offset" manner. It is used to normalize the interference level of the current channel and serves as the input to the Sigmoid function, causing the modulation behavior to exhibit a non-linear mechanism of "suppressing strong interference and preserving weak signals." d is a non-linear suppression factor that controls the weight adjustment function (Sigmoid) on... The steepness of the response curve is used to construct a dynamic nonlinear weight reduction curve, enabling the system to tolerate "slight disturbances" and rapidly suppress "extreme disturbances," achieving gradient adaptation of the response. It is a positive number; the larger the value, the steeper the function and the more aggressive the adjustment result; the smaller the value, the smoother the function transition. e is the natural base, and Λ is the link dynamic range imbalance coefficient. t It is the reference threshold for the link dynamic range imbalance coefficient;

[0108] The sigmoid function in the weighting adjustment function is a commonly used sigmoid nonlinear mapping function, mathematically expressed as:

[0109]

[0110] Its output value always lies between (0,1) and has the following characteristics: when the input x is very small, the output value approaches 0; when x is very large, the output value approaches 1; and it is most sensitive to changes in the range where the input is close to 0, exhibiting a steep curve transition. Therefore, the Sigmoid function is often used to smoothly compress any input value into a controllable range, forming a nonlinear response pattern of "mild in the middle and extreme at both ends".

[0111] In audio weighting, the Sigmoid function maps the deviation between the interference intensity of each channel and the average interference level. This results in slightly suppressed sources with slightly higher interference indices, while sources with significantly higher interference indices are suppressed more severely, thus constructing a "non-linear, progressive" suppression strategy. Compared to linear functions, the Sigmoid avoids over-responding to small fluctuations while providing a fast and powerful response to sudden strong interference, making it a very suitable function model for constructing dynamic suppression weighting mechanisms.

[0112] This formula constructs an inhibitory weighting mechanism based on the intensity of disturbance and introduces an imbalance control factor, which makes the system more strongly suppressive when the imbalance is high and more moderate in the state of slight imbalance.

[0113] After adjusting the weights of each channel, to ensure the consistency of amplitude and the accuracy of signal proportions in the mixed output, the weights W of the source channel mixed output are adjusted. i Normalization is performed, and the mixed output signal is reconstructed based on the normalization weights. The calculation formula is as follows:

[0114]

[0115] In the formula, Y min This is the mixed output signal, representing the final composite speech signal output by the system. It is the result of weighted synthesis from multiple sound source channels, and is a continuously changing speech waveform data stream over time. W j This represents the weight value of the j-th audio source channel in the current audio mix output, where n is the total number of audio source channels, and S... i is the original speech signal of the i-th sound source channel, and is the unprocessed audio data collected from each sound source channel.

[0116] This formula ensures that the energy distribution of the total output audio remains logically consistent, while minimizing the interference of abnormal channels on the final audio result, significantly improving the stability and accuracy of the speech recognition and translation system under unbalanced conditions.

[0117] This invention effectively achieves real-time identification and adaptive control of the dynamic range imbalance of the audio processing link by introducing key modules such as audio dynamic state perception, feature preprocessing, abnormal feature quantification analysis, intelligent state evaluation, and adaptive weight adjustment. This significantly improves the system's stability and robustness in complex multi-sound source environments. The system constructs a high-quality feature dataset through audio dynamic state acquisition and preprocessing, and intelligently evaluates the audio state using feature engineering and machine learning models. When sonic booms or abnormal sound source disturbances occur, it can quickly identify the interference source and dynamically adjust its weight, preventing its interference from spreading to the global AGC mechanism and ensuring the continuity and translatability of other normal speech channels. Compared to the shortcomings of traditional systems that are susceptible to high-energy interference and overall failure, this system implements a differentiated control strategy of "local suppression of interference channels and stable maintenance of the main speech link," effectively avoiding problems such as translation task interruption and speech recognition failure. This significantly improves the practicality and interaction quality of Bluetooth headsets in multi-speaker, multi-language interaction scenarios.

[0118] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0119] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0120] It should be noted that, in this document, the use of relational terms such as "first" and "second" is merely for distinguishing one entity or operation from another, and does not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0121] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0122] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0123] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0124] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0125] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0126] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0127] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A multi-voice convergence anti-interference Bluetooth headset translation system, characterized in that, It includes an audio dynamic state acquisition module, an audio feature preprocessing and data construction module, a dynamic range anomaly feature extraction and quantification analysis module, an audio system state intelligent evaluation module, and a sound source weight adaptive adjustment module: The audio dynamic state acquisition module, during the real-time language translation task performed by the Bluetooth headset, monitors and captures the audio dynamic range state characteristic data generated by each sound source channel in real time through the high-precision microphone array, digital signal processing chip and audio processing module integrated inside the headset. The audio feature preprocessing and data construction module preprocesses the raw audio dynamic range feature data acquired in real time, thereby establishing an efficient and reliable dataset to support subsequent feature engineering and intelligent analysis processes. The dynamic range anomaly feature extraction and quantification analysis module applies feature engineering techniques to the preprocessed dataset to accurately extract key indicators characterizing the overall dynamic range imbalance of the audio processing link, and conducts in-depth comprehensive analysis of the extracted key indicators to quantify the severity of the overall imbalance of the audio processing link. The audio system status intelligent assessment module takes the comprehensively analyzed indicators as feature vectors and inputs them into the pre-trained machine learning model. The model then performs a rapid and intelligent assessment of the current operating status of the audio processing system to determine whether the current audio processing system has entered a dynamic range imbalance state. The sound source weight adaptive control module immediately adjusts the sound source priority when it detects a dynamic range imbalance in the audio processing system. It temporarily sets the sound source channel identified as having strong interference characteristics to "low priority" and weakens its weight in the mixed output. At the same time, it automatically increases the proportion of the remaining sound sources in the overall output to suppress the interference of sudden abnormalities on the normal conversation channel. The specific steps for immediately adjusting the sound source priority after identifying a dynamic range imbalance in the audio processing system, temporarily setting the sound source channel identified as having strong interference characteristics to "low priority" and weakening its weight in the mixed output, while automatically increasing the proportion of the remaining sound sources in the overall output, are as follows: After identifying a dynamic range imbalance in the audio processing system, the interference intensity of all sound source channels in the system is first assessed. By calculating the interference index of each channel, its impact on the imbalance is quantified. The expression for calculating the interference index is as follows: In the formula, It is the interference index, representing the first... The overall interference intensity score for dynamic range imbalance caused by each sound source channel. It is the intensity of the amplitude change, representing the sound source channel. The maximum rate of change of the amplitude of the speech signal between adjacent frames, It is the gain offset. It is the amplitude mutation weighting coefficient; Obtain the interference index for all channels Then, the weights of each sound source channel in the mixed output are dynamically adjusted to weaken the output contribution of strong interference sources and increase the proportion of the remaining channels. The weight adjustment formula is as follows: In the formula, It is the source channel mixing output weight, representing the first... The weight values ​​of each sound source channel in the current audio mix output. It is the average value of the interference index of all sound source channels. It is a nonlinear inhibitory factor. It is the natural base. It is the link dynamic range imbalance coefficient. It is the reference threshold for the link dynamic range imbalance coefficient; After adjusting the weights of each channel, to ensure the consistency of amplitude and the accuracy of signal proportions in the mixed output, the weights of the source channel mixing outputs are adjusted. Normalization is performed, and the mixed output signal is reconstructed based on the normalization weights. The calculation formula is as follows: In the formula, It is the mixing output signal. It means the first The weight values ​​of each sound source channel in the current audio mix output. This is the total number of sound source channels. It is the first The original speech signal from each sound source channel.

2. The multi-voice convergence anti-interference Bluetooth headset translation system according to claim 1, characterized in that, The specific steps for acquiring audio dynamic range state feature data during real-time language translation tasks performed by Bluetooth headsets are as follows: The high-precision microphone array inside the headphones synchronously collects multi-channel raw speech signals from different directions or speakers to ensure coverage of all sound source information. The acquired analog audio signal is input to the digital signal processing chip, where A / D conversion, signal enhancement, and preliminary noise reduction are performed to generate a computable digital audio stream. The audio processing module analyzes the digital audio stream, extracts the automatic gain control response parameters for each channel, and obtains the state characteristics related to the dynamic range. The feature data is cached in a time-series structure as the input basis for subsequent dynamic range imbalance recognition, speech interference judgment, and intelligent translation stability control.

3. The multi-voice convergence anti-interference Bluetooth headset translation system according to claim 1, characterized in that, Feature engineering techniques are applied to the preprocessed dataset to accurately extract key indicators characterizing the overall dynamic range imbalance of the audio processing link. The extracted indicators include the output amplitude change after the dynamic compressor response and the amplitude change rate between consecutive frames of the audio waveform. The output amplitude change after the dynamic compressor response and the amplitude change rate between consecutive frames of the audio waveform are comprehensively analyzed under the detection window to generate the channel dynamic compression imbalance index and the amplitude gradient mutation index, respectively. The severity of the overall imbalance of the audio processing link is quantified by the channel dynamic compression imbalance index and the amplitude gradient mutation index.

4. The multi-voice convergence anti-interference Bluetooth headset translation system according to claim 3, characterized in that, The specific steps for generating the channel dynamic compression imbalance index by comprehensively analyzing the output amplitude change after the dynamic compressor response within the detection window are as follows: Within the detection window, for all valid sound source channels in the Bluetooth headset, the maximum compression gradient value of each channel in the output amplitude envelope curve after processing by the dynamic compressor is obtained. Based on the maximum compression gradient values ​​of all channels, a compression difference mapping matrix between channels is constructed. The definition is as follows: In the formula, It is a compressed difference mapping matrix The elements in the text represent the sound source channels. With the sound source channel The degree of normalized difference in compressive response intensity , This is the total number of sound source channels. It is the sound source channel The maximum compression gradient value, It is the sound source channel The maximum compression gradient value, It is a normalized reference factor, representing the normalization reference factor for the sound source channel. With the sound source channel The maximum compression gradient value is taken; Obtaining the compressed difference mapping matrix Then, statistical analysis was performed on all channel combinations to calculate the channel dynamic compression imbalance index. The calculation expression is as follows: In the formula, It is the dynamic compression imbalance index of the vocal tract. It is the total number of channels without duplicates. It is a nonlinear enhancement function, representing that for each The weights for medium- to high-level differences are enhanced by applying a nonlinear transformation, and the calculation method is as follows: , It is the amplification factor, used to adjust the sensitivity to differential abrupt changes. It is the hyperbolic tangent function.

5. A multi-voice convergence anti-interference Bluetooth headset translation system according to claim 3, characterized in that, The specific steps for generating the amplitude gradient abrupt change index by comprehensively analyzing the amplitude change rate between consecutive frames of an audio waveform within a detection window are as follows: Within the detection window, the amplitude features of each frame of audio signal are extracted. Based on the amplitude changes between consecutive frames, a sudden change response factor is generated, using the following formula: In the formula, It is the detection window number Amplitude characteristics of frame audio signals, It is the detection window number The amplitude characteristics of the audio signal in one frame, that is, the amplitude characteristics of the audio signal in the next frame. It is the amplitude response sensitivity coefficient. It is the power exponent of the mutation response. It is a mutation direction preservation factor. It is a mutation response factor; Obtain the mutation response factor of all adjacent frames. Then, the acquired response values ​​are mapped to the nonlinear enhancement domain. Through amplification and integration, a quantitative index characterizing the sum of abrupt change in intensity within the entire detection window is formed, namely the amplitude gradient abrupt change index. The calculation expression is as follows: In the formula, It is the magnitude gradient abrupt change index. It is the amplification factor of the sudden change response. It is the natural base. It is the number of frames contained in the detection window. It is a normalized constant.

6. A multi-voice convergence anti-interference Bluetooth headset translation system according to claim 3, characterized in that, The dynamic compression imbalance index and amplitude gradient mutation index of the duct, which have been comprehensively analyzed, are input as feature vectors into a pre-trained machine learning model. The model generates a link dynamic range imbalance coefficient, and the current operating status of the audio processing system is quickly and intelligently evaluated based on the link dynamic range imbalance coefficient to determine whether the current audio processing system has entered a dynamic range imbalance state.

7. A multi-voice convergence anti-interference Bluetooth headset translation system according to claim 6, characterized in that, The link dynamic range imbalance coefficient generated by the pre-trained machine learning model during the rapid and intelligent evaluation of the current operating status of the audio processing system is compared with a pre-set reference threshold for the link dynamic range imbalance coefficient to determine whether the current audio processing system has entered a dynamic range imbalance state. The determination logic is as follows: If the link dynamic range imbalance coefficient is greater than the preset reference threshold for the link dynamic range imbalance coefficient, the current audio processing system is determined to have entered a dynamic range imbalance state; if the link dynamic range imbalance coefficient is less than or equal to the preset reference threshold for the link dynamic range imbalance coefficient, the current audio processing system is determined not to have entered a dynamic range imbalance state.