Multi-voice common exchange type anti-interference Bluetooth earphone translation system
By real-time monitoring of the audio dynamic range status and adaptively adjusting the sound source weight, the dynamic range imbalance problem of the multi-voice co-convergence anti-interference Bluetooth headset translation system during abnormal sonic booms is solved, the system stability and continuity of the translation task are achieved, and the practicality of multi-language multi-speaker interaction is improved.
Patent Information
- Application Number
- CN202510729140.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-03
AI Technical Summary
When faced with abnormal sonic booms such as sudden high-decibel shouting, physical collisions, or strong background noise, the existing multi-voice co-convergence anti-interference Bluetooth headset translation system may cause an imbalance in the dynamic range of the audio processing link, resulting in voice signal weakening or system interruption, affecting the continuity and accuracy of the translation task.
The audio dynamic state acquisition module monitors the audio dynamic range status in real time. Combined with feature preprocessing, dynamic range anomaly feature extraction and quantitative analysis, the machine learning model is used for intelligent evaluation, adaptively adjusting the sound source weight, suppressing strong interference channels and maintaining the stability of the main voice channel.
It achieves real-time recognition and adaptive control of dynamic range imbalances in the audio processing chain, significantly improving the stability and robustness of the system in complex multi-sound source environments, avoiding translation task interruptions and speech recognition failures, and improving practicality and interaction quality in multi-speaker and multi-language interaction scenarios.
Smart Images

Figure CN120676282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice processing, and in particular to a multi-voice co-convergence anti-interference Bluetooth headset translation system. Background Art
[0002] "Multi-voice co-convergence anti-interference Bluetooth headset translation" refers to a Bluetooth headset system that can simultaneously receive and process voice information from multiple speakers, and adopts an intelligent interference suppression mechanism during the processing to ensure that each voice segment can be clearly separated, accurately identified and efficiently translated. The system is based on a multi-channel voice input fusion mechanism (co-convergence), which can automatically distinguish the voice characteristics of different voice sources, combine noise reduction algorithms with sound source localization technology to separate voice signals and filter interference, and use an integrated multi-language real-time translation engine to quickly translate the extracted voice and output audio / text. It is widely applicable to multi-language and multi-speaker communication scenarios (such as international conferences, tourism, multi-person conversations, etc.). The core of the system is to achieve an intelligent interactive experience of "multiple people speaking without confusion, multiple languages being able to be translated, and interference environment being controllable."
[0003] The existing technology has the following shortcomings: In a multi-voice co-convergent speech processing system, in order to adapt to the differences in speech intensity between different sound sources, an automatic gain control (AGC) mechanism is usually introduced to achieve dynamic audio gain adjustment. However, in actual application scenarios, when an unexpected event (such as high-decibel shouting, physical collision, strong background noise, etc.) causes an "abnormal sonic boom" signal with an instantaneous sound pressure far higher than that of other sound sources, the AGC module in the system may abnormally trigger a rapid gain compression response due to sensing the extreme energy mutation, thereby causing an imbalance in the overall dynamic range of the audio processing chain. In this process, in order to protect the system hardware or maintain a stable output amplitude, the AGC often synchronously lowers the gain of all input channels. As a result, the speech signal within the normal sound pressure range is significantly weakened or even completely drowned out. In severe cases, the continuous energy peak exceeds the limit, triggering the system's channel protection mechanism or the input link automatic shutdown mechanism, resulting in the interruption of the entire multi-voice collection and translation function. Once this problem occurs, it can easily lead to voice information loss, translation logic errors or communication interruption. Especially in multi-language, multi-speaker concurrent interaction environments (such as international conferences, medical consultations, emergency command, etc.), it will cause irreversible communication barriers and significantly affect the practicality and security of the system.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-voice co-convergence anti-interference Bluetooth headset translation system, which realizes local suppression of interference channels and stable maintenance of main voice channels by real-time identification of audio dynamic range imbalance and adaptive adjustment of sound source weights, effectively ensuring the continuity of translation tasks and recognition accuracy, so as to solve the problems in the above-mentioned background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solutions: a multi-voice co-convergence anti-interference Bluetooth headset translation system, comprising an audio dynamic state acquisition module, an audio feature preprocessing and data construction module, a dynamic range anomaly feature extraction and quantitative analysis module, an audio system state intelligent assessment module, and a sound source weight adaptive control module:
[0007] The audio dynamic state acquisition module monitors and captures the audio dynamic range state characteristic data generated by each sound source channel in real time during the Bluetooth headset's real-time language translation task through the high-precision microphone array, digital signal processing chip, and audio processing module integrated into the headset;
[0008] The audio feature preprocessing and data construction module preprocesses the raw audio dynamic range feature data collected in real time to establish an efficient and reliable data set to support subsequent feature engineering and intelligent analysis processes;
[0009] The dynamic range anomaly feature extraction and quantitative analysis module applies feature engineering techniques to the preprocessed data set to accurately extract key indicators that represent the overall dynamic range imbalance of the audio processing chain. It then conducts in-depth comprehensive analysis of the extracted key indicators to quantify the severity of the overall imbalance in the audio processing chain.
[0010] The audio system status intelligent assessment module inputs the comprehensively analyzed indicators as feature vectors into a pre-trained machine learning model. The model then conducts a rapid and intelligent assessment of the current operating status of the audio processing system to determine whether the current audio processing system has entered a state of dynamic range imbalance.
[0011] The sound source weight adaptive control module immediately adjusts the sound source priority when it detects a dynamic range imbalance in the audio processing system. It temporarily sets the sound source channel identified as having strong interference characteristics to "low priority" and weakens its weight in the mixed output. At the same time, it automatically increases the proportion of remaining sound sources in the overall output, suppressing the interference of sudden anomalies on normal conversation channels.
[0012] Preferably, during the execution of the real-time language translation task by the Bluetooth headset, the specific steps of obtaining the audio dynamic range state characteristic data are as follows:
[0013] The high-precision microphone array inside the headset synchronously collects multi-channel original voice signals from different directions or speakers to ensure coverage of all sound source information;
[0014] The collected analog voice signal is input into the digital signal processing chip, which performs A / D conversion, signal enhancement, and preliminary noise reduction to generate a computable digital audio stream;
[0015] The audio processing module analyzes the digital audio stream and extracts the state characteristics related to the dynamic range of the automatic gain control response parameters of each channel;
[0016] The feature data is cached in a structured time series format, serving as the input basis for subsequent dynamic range imbalance identification, speech interference judgment, and intelligent translation stability control.
[0017] Preferably, feature engineering technology is applied to the preprocessed data set to accurately extract key indicators that characterize the overall dynamic range imbalance of the audio processing link. The extracted indicators include the output amplitude change after the dynamic compressor responds and the amplitude change rate between consecutive frames of the audio waveform. The output amplitude change after the dynamic compressor responds and the amplitude change rate between consecutive frames of the audio waveform are comprehensively analyzed under the detection window to generate the channel dynamic compression imbalance index and the amplitude gradient mutation index respectively. The severity of the overall imbalance of the audio processing link is determined through the channel dynamic compression imbalance index and the amplitude gradient mutation index.
[0018] Preferably, the specific steps of comprehensively analyzing the output amplitude change after the dynamic compressor responds in the detection window to generate the channel dynamic compression imbalance index are as follows:
[0019] Within the detection window, for all valid sound source channels in the Bluetooth headset, the maximum compression gradient value in the output amplitude envelope curve of each channel after being processed by the dynamic compressor is obtained. Based on the maximum compression gradient value of all channels, the compression difference mapping matrix M between channels is constructed, which is defined as follows:
[0020]
[0021] , where M i,j It is an element in the compression difference mapping matrix M, which represents the normalized difference between the compression response strength of the sound source channel i and the sound source channel j, i∈1,2,3,…,n, n is the total number of sound source channels, G i is the maximum compression gradient value of the sound source channel i, G j is the maximum compression gradient value of the sound source channel j, max(G i , G j ) is the normalized reference factor, which represents the maximum compression gradient value of the sound source channel i and the sound source channel j;
[0022] After obtaining the compression difference mapping matrix M, statistical analysis is performed on all channel combinations to calculate the channel dynamic compression imbalance index. The calculation expression is as follows:
[0023]
[0024] , where CDCI is the channel dynamic compression imbalance index, is the total number of non-repeated channel pairs, σ(M i,j ) is a nonlinear enhancement function, which means that for each M i,j A nonlinear transformation is applied to enhance the weights of medium and high degree differences. The calculation method is as follows: σ(x) = tanh(k·x), where k is the amplification factor used to adjust the response sensitivity to differential mutations, and tanh(x) is the hyperbolic tangent function.
[0025] Preferably, the specific steps of comprehensively analyzing the amplitude change rate between consecutive frames of the audio waveform under the detection window to generate the amplitude gradient mutation index are as follows:
[0026] Within the detection window, the amplitude characteristics of each frame of audio signal are extracted, and the mutation response factor is generated based on the amplitude changes between consecutive frames. The generation formula is as follows:
[0027] ΔΦ a =tanh(λ·|A a+1 -A a | β )·sgn(A a+1 -A a )
[0028] , where A a is the amplitude feature of the audio signal of the ath frame in the detection window, A a+1 is the amplitude feature of the audio signal of the a+1th frame in the detection window, that is, the amplitude feature of the next frame of audio signal, λ is the amplitude response sensitivity coefficient, β is the mutation response power exponent, sgn(A a+1 -A a ) is the mutation direction maintaining factor, ΔΦ a is the mutation response factor;
[0029] After obtaining the mutation response factor ΔΦ of all adjacent frames a After that, the acquired response value is mapped to the nonlinear enhancement domain. Through amplification and integration, a quantitative index representing the sum of the mutation intensity in the entire detection window is formed, namely the amplitude gradient mutation index. The calculation expression is as follows:
[0030]
[0031] , where AVGI is the amplitude gradient mutation index, γ is the mutation response amplification coefficient, e is the natural base, m is the number of frames contained in the detection window, and Z is the normalization constant.
[0032] Preferably, the channel dynamic compression imbalance index and amplitude gradient mutation index that have undergone comprehensive analysis are input as feature vectors into a pre-trained machine learning model, and the link dynamic range imbalance coefficient is generated by the model. Based on the link dynamic range imbalance coefficient, a rapid and intelligent evaluation of the current operating status of the audio processing system is performed to determine whether the current audio processing system has entered a dynamic range imbalance state.
[0033] Preferably, a link dynamic range imbalance coefficient generated when a pre-trained machine learning model is used to quickly and intelligently evaluate the current operating state of the audio processing system is compared with a pre-set link dynamic range imbalance coefficient reference threshold to determine whether the current audio processing system has entered a dynamic range imbalance state. The judgment logic is as follows:
[0034] If the link dynamic range imbalance coefficient is greater than a preset link dynamic range imbalance coefficient reference threshold, it is determined that the current audio processing system has entered a dynamic range imbalance state; if the link dynamic range imbalance coefficient is less than or equal to a preset link dynamic range imbalance coefficient reference threshold, it is determined that the current audio processing system has not entered a dynamic range imbalance state.
[0035] Preferably, when a dynamic range imbalance is identified in the audio processing system, the sound source priority is immediately adjusted, the sound source channel identified as having strong interference characteristics is temporarily set to "low priority" and its weight in the mixed output is weakened, and the proportion of the remaining sound sources in the overall output is automatically increased. The specific steps are as follows:
[0036] After identifying the dynamic range imbalance in the audio processing system, the interference intensity of all sound source channels in the audio processing system is first evaluated. The interference index of each channel is calculated to quantify its impact on the imbalance state. The interference index calculation expression is as follows:
[0037] ψ i =α·ΔA i +β·δH i
[0038] , where ψ i is the interference index, which represents the comprehensive interference intensity score of the i-th sound source channel causing the dynamic range imbalance, ΔA i is the amplitude mutation intensity, which represents the maximum rate of change of the speech signal amplitude between adjacent frames of the sound source channel i, δH i is the gain offset, α is the amplitude mutation weight coefficient;
[0039] After obtaining the interference index ψ of all channels i After that, the weight of each sound source channel in the mixed output is dynamically adjusted to weaken the output contribution of the strong interference sound source and increase the proportion of the remaining channels. The weight adjustment formula is as follows:
[0040]
[0041] , where W i Is the sound source channel mixing output weight, which represents the weight value of the i-th sound source channel in the current audio mixing output. is the average value of the interference index of all sound source channels, d is the nonlinear suppression factor, e is the natural base, Λ is the link dynamic range imbalance coefficient, Λ t is the link dynamic range imbalance coefficient reference threshold;
[0042] After completing the weight adjustment of each channel, in order to ensure the amplitude consistency and signal ratio accuracy of the mixed output, the sound source channel mixed output weight W i Normalization is performed and the mixed output signal is reconstructed according to the normalized weights. The calculation formula is as follows:
[0043]
[0044] , where Y min is the mixed output signal, W j It represents the weight value of the jth sound source channel in the current audio mixing output, n is the total number of sound source channels, S i is the original speech signal of the i-th sound source channel.
[0045] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0046] By introducing key modules such as audio dynamic state perception, feature preprocessing, abnormal feature quantification analysis, intelligent state assessment, and adaptive weight control, the present invention effectively achieves real-time identification and adaptive control of the dynamic range imbalance state of the audio processing link, thereby significantly improving the system's stability and robustness in complex multi-sound source environments. The system constructs a high-quality feature dataset through audio dynamic state acquisition and preprocessing, and intelligently assesses the audio state by combining feature engineering and machine learning models. When a sonic boom or abnormal sound source disturbance occurs, it can quickly identify the interference source and dynamically adjust its weight to prevent its interference from spreading to the global AGC mechanism, thereby ensuring the continuity and translatability of other normal voice channels. Compared with the disadvantage of traditional systems that are susceptible to high-energy interference and overall failure, this system implements a differentiated control strategy of "locally suppressing the interference channel and maintaining the stability of the main voice link", effectively avoiding problems such as translation task interruption and speech recognition failure, and significantly improving the practicality and interaction quality of Bluetooth headsets in multi-speaker, multi-language interaction scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0048] Figure 1 This is a module schematic diagram of a multi-voice co-convergence anti-interference Bluetooth headset translation system of the present invention. DETAILED DESCRIPTION
[0049] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that the description of this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0050] The present invention provides Figure 1 The multi-voice co-convergence anti-interference Bluetooth headset translation system shown in the figure includes an audio dynamic state acquisition module, an audio feature preprocessing and data construction module, a dynamic range anomaly feature extraction and quantitative analysis module, an audio system state intelligent assessment module, and a sound source weight adaptive control module:
[0051] The audio dynamic state acquisition module monitors and captures the audio dynamic range state characteristic data generated by each sound source channel in real time during the Bluetooth headset's real-time language translation task through the high-precision microphone array, digital signal processing chip (DSP), and audio processing module integrated into the headset;
[0052] When Bluetooth headsets perform real-time language translation tasks, the system relies on the high-precision microphone array integrated into the headset to pick up original voice signals from different directions and speakers. Combined with the digital signal processing chip (DSP) and the built-in audio processing module, these voice signals are analyzed in real time to extract operating status characteristic data related to the audio dynamic range. Specific data obtained includes: short-term energy, instantaneous peak amplitude, average sound pressure level (SPL), signal-to-noise ratio (SNR), automatic gain control (AGC) parameter changes, dynamic compression ratio, etc. of each sound source channel. This data reflects the system's processing behavior and response characteristics to changes in input voice signal intensity at different time points. Its core role is to provide key basis for subsequent judgment of whether there is abnormal interference (such as sonic booms) or imbalance in the dynamic range of the audio link, thereby laying a data foundation for the stable execution of intelligent gain adjustment and translation tasks.
[0053] The core function of this step is to provide accurate raw data support for subsequent sound source interference analysis, ensuring that subsequent diagnosis and control are based on real and comprehensive audio operation status data, and ensuring the authenticity and comprehensiveness of the data from the source.
[0054] When a Bluetooth headset performs a real-time language translation task, the specific steps for obtaining audio dynamic range state feature data are as follows:
[0055] The first step is to use a high-precision microphone array inside the headset to synchronously collect multi-channel original voice signals from different directions or speakers to ensure coverage of all sound source information;
[0056] In the second step, the collected analog voice signal is input into the digital signal processing chip (DSP), which performs A / D conversion, signal enhancement, preliminary noise reduction and other processing to generate a computable digital audio stream;
[0057] In the third step, the audio processing module analyzes the digital audio stream and extracts dynamic range-related state features such as short-term energy, peak amplitude, average sound pressure level, and automatic gain control (AGC) response parameters of each channel;
[0058] The fourth step is to cache this feature data in a structured time series format, serving as the input for subsequent dynamic range imbalance detection, speech interference assessment, and intelligent translation stability control. This process enables real-time quantitative perception of the headset's speech environment and dynamic characterization of the system's operating status.
[0059] The audio feature preprocessing and data construction module preprocesses the raw audio dynamic range feature data collected in real time to establish an efficient and reliable data set to support subsequent feature engineering and intelligent analysis processes;
[0060] Preprocessing is performed on the raw audio dynamic range feature data collected in real time, which mainly includes the following specific steps: first, data denoising is performed. Background noise and instantaneous spike interference in the feature data are removed through filters (such as median filtering, Kalman filtering) or adaptive algorithms to improve data stability; second, missing data interpolation is performed. Linear interpolation, spline interpolation, or time series prediction models (such as ARIMA) are used to repair missing feature values caused by instantaneous packet loss or hardware jitter during the acquisition process to ensure data continuity; then, data normalization or normalization operations are performed to uniformly map feature data of different dimensions or amplitude intervals to a fixed range (such as [0, 1] or a mean of 0 and a standard deviation of 1) to avoid feature bias affecting model learning effects; further time window segmentation is performed to divide the continuously collected data into time segments according to sliding windows or fixed frame lengths to facilitate capturing short-term dynamic trends and constructing time series feature vectors; finally, outlier detection and removal and correlation analysis between features are performed to identify extreme outliers and remove redundant features or construct derived indicators based on correlation evaluation to improve feature expression quality and model discrimination ability. The core function of this preprocessing process is to transform the original feature data from its mixed, discrete, and noisy original form into a clearly structured, continuous, and standardized analysis input, providing a high-quality data foundation for subsequent feature extraction, model training, and state recognition.
[0061] By implementing systematic preprocessing (such as denoising, completion, normalization, slicing, etc.) on the audio dynamic range feature data collected in real time, the originally messy and defective raw data is converted into a data set with complete structure, stable time series, and strong feature consistency. This data set has high accuracy, high continuity and high expressiveness, and can be used as a standard input for subsequent algorithm processing. Its main function is to provide a solid data foundation for feature engineering, ensuring that the extracted key indicators are physically reasonable and statistically stable. It also provides high-quality input for the training, evaluation and reasoning of machine learning models, state recognition algorithms or translation strategy optimization models, making the entire audio processing and intelligent recognition system more robust, accurate and adaptable. In other words, this data set is a crucial "information foundation" in the "intelligent analysis" chain.
[0062] The key role of this step is to improve the quality and analyzability of the data by eliminating clutter and interference factors in the data, thereby establishing an efficient and reliable data set to support subsequent feature engineering and intelligent analysis processes.
[0063] The dynamic range anomaly feature extraction and quantitative analysis module applies feature engineering techniques to the preprocessed data set to accurately extract key indicators that represent the overall dynamic range imbalance of the audio processing chain. It then conducts in-depth comprehensive analysis of the extracted key indicators to quantify the severity of the overall imbalance in the audio processing chain.
[0064] Feature engineering technology is applied to the preprocessed dataset to accurately extract key indicators that characterize the overall dynamic range imbalance of the audio processing chain. The extracted indicators include the output amplitude change after the dynamic compressor responds and the amplitude change rate between consecutive frames of the audio waveform. The output amplitude change after the dynamic compressor responds and the amplitude change rate between consecutive frames of the audio waveform are comprehensively analyzed under the detection window to generate the channel dynamic compression imbalance index and amplitude gradient mutation index respectively. The channel dynamic compression imbalance index and amplitude gradient mutation index are used to measure the severity of the overall imbalance of the audio processing chain.
[0065] Abnormal output amplitude fluctuations after a dynamic compressor's response can directly indicate an imbalance in the overall dynamic range of the audio processing chain. This is because the dynamic compressor, as the core unit in an audio system that controls input signal amplitude fluctuations, largely reflects the system's ability to adapt and adjust to energy differences between different sound sources. Under normal conditions, the dynamic compressor smoothly and gradually adjusts gain based on the sound source input intensity, ensuring that all channels operate within a uniform and reasonable dynamic range. However, when problems such as sudden sonic booms, abnormal enhancement of a single sound source, or abnormal system feedback occur, the compressor may exhibit a dramatic nonlinear response in some channels, manifesting as abnormalities such as sudden drops in output amplitude, excessive compression, or delayed adjustment. These abnormal fluctuations not only affect the voice quality of the compressed channel but also affect the overall dynamic behavior of other sound sources through the co-sinking link, disrupting dynamic consistency between channels and causing system-level dynamic range imbalance. Therefore, abnormal fluctuations in the dynamic compressor's output amplitude are not just a single-channel issue; rather, they are a key manifestation of overall chain imbalance, providing high system-level diagnostic value.
[0066] The specific steps for comprehensively analyzing the output amplitude changes after the dynamic compressor response in the detection window to generate the channel dynamic compression imbalance index are as follows:
[0067] Within the detection window, for all valid sound source channels in the Bluetooth headset, the maximum compression gradient value in the output amplitude envelope curve of each channel after being processed by the dynamic compressor is obtained. This value represents the compression response strength of the channel within the current detection window, that is, the maximum amplitude change rate caused by the compressor acting on the channel. Based on the maximum compression gradient values of all channels, the compression difference mapping matrix M between channels is constructed, which is defined as follows:
[0068]
[0069] , where M i,jIt is an element in the compression difference mapping matrix M, which represents the normalized difference in the compression response intensity between the sound source channel i and the sound source channel j. It is used to quantify the degree of compression inconsistency between any two channels. The larger the value, the greater the difference in compression response between the channels, indicating that the system is inconsistent in regulating multiple sound sources and there is a potential risk of imbalance. i∈1,2,3,…,n, where n is the total number of sound source channels. G i is the maximum compression gradient value of the sound source channel i, which indicates the maximum amplitude change rate (i.e., the maximum compression response slope) in the output amplitude curve of the audio signal after being processed by the dynamic compressor within the detection time window of the sound source channel i. j is the maximum compression gradient value of the sound source channel j, which is used to make a differential comparison with channel i to identify the degree of inconsistency in the compression response between the two. max(G i , G j ) is the normalized reference factor, which means taking the maximum compression gradient value of the sound source channel i and the sound source channel j as the maximum, and realizing the normalization processing so that the difference G i -G j Compare with its reference amplitude to avoid amplifying essentially weak differences due to high absolute channel amplitude;
[0070] This mapping matrix comprehensively considers compression inconsistencies between channels and eliminates the influence of absolute amplitude values through normalization, thereby focusing on the structural manifestation of "relative differences." This matrix provides basic data support for the subsequent generation of the channel compression imbalance index.
[0071] After obtaining the compression difference mapping matrix M, statistical analysis is performed on all channel combinations to calculate the channel dynamic compression imbalance index. The calculation expression is as follows:
[0072]
[0073] Where CDCI is the channel dynamic compression imbalance index, which is used to measure the degree of response difference between multiple sound source channels under dynamic compressor processing. The value range is [0, 1]. The larger the value, the more significant the difference in compression behavior between channels, and the more likely the system is in a dynamic range imbalance state. is the total number of non-repeated channel pairs (normalization factor), which represents the number of combinations of comparing any two channels in n channels, σ(M i,j ) is a nonlinear enhancement function, which means that for each M i,j A nonlinear transformation is applied to enhance the weights of medium and high degree differences. The calculation method is as follows: σ(x) = tanh(k·x), where k is the amplification factor used to adjust the response sensitivity to differential mutations, and tanh(x) is the hyperbolic tangent function.
[0074] A nonlinear enhancement function refers to a mathematical function used to perform nonlinear transformations on input data in feature processing or model calculations. Its purpose is not simple scaling, but differential amplification or suppression according to different input amplitude ranges. Its core function is to enhance the influence of medium and high amplitude features on models or indicators, while suppressing the interference effects of low amplitude changes, thereby improving the system's recognition sensitivity to "abnormal states" or "critical imbalances." In the identification of dynamic range imbalances, small amplitude compression differences may be normal fluctuations within the system and do not require intervention; while medium and high amplitude differences often indicate that the system's compression mechanism is out of control or a sonic boom impact, so they need to be amplified through nonlinear functions to make them more dominant in the calculation of weights. Commonly available nonlinear enhancement functions include the hyperbolic tangent function tanh(x), the exponential function 1-e -x , Sigmoid function Among them, tanh(x) has good discrimination and continuity, and is often used to smoothly enhance medium and high value responses while avoiding excessive amplification of low values. It is widely used in scenarios such as audio imbalance detection, image edge enhancement, and neural network activation.
[0075] By aggregating and normalizing the difference values of all non-overlapping channel pairs in the matrix, the CDCI effectively characterizes the consistency of the system's multi-channel compression response within the current detection window. A high CDCI indicates increasing differences in compression behavior across multiple channels, indicating a significant dynamic range imbalance in the audio processing chain. A low CDCI indicates that compression control across the system's channels is generally coordinated, and the audio processing chain is in a stable and balanced state.
[0076] The dynamic compression imbalance index (CDI) is generated by comprehensively analyzing the output amplitude changes after the dynamic compressor's response within the detection window. If the dynamic compressor's output amplitude response to each sound source channel differs significantly—meaning some channels are severely compressed while others are not significantly affected—this will lead to inconsistent compression behavior between channels, significantly increasing the CDI. In this case, the system's internal dynamic range control strategy is unable to coordinate and uniformly adjust multiple sound sources, indicating a structural imbalance in the audio link. Therefore, a larger CDI indicates uneven compressor control, disrupting the overall system dynamic range and threatening recognition and translation stability. Conversely, a smaller CDI indicates that the compressor's response to all sound source channels is consistent, the system is in a balanced state, and the dynamic range is within a reasonable control range, indicating no imbalance.
[0077] When the amplitude change rate between consecutive audio waveform frames increases dramatically within a short period of time, it typically indicates an abnormal state of overall dynamic range imbalance within the audio processing chain. This dramatic fluctuation suggests that the system's automatic gain control (AGC) or dynamic compressor has been abnormally triggered by a strong interfering sound source (such as a sonic boom or sudden high sound pressure), causing the audio signal's gain adjustment to become unstable, leading to nonlinear amplitude jumps between frames. Normally, the amplitude change rate between consecutive frames should remain within a stable range, reflecting the system's consistent and stable response to different sound sources. However, sudden amplitude increases typically indicate that the system, attempting to suppress high-energy interference, has failed to simultaneously control other sound sources, disrupting the dynamic balance of multiple channels within the chain. This abnormal fluctuation not only disrupts speech continuity and intelligibility but also affects the accuracy of the mixed output and semantic extraction, signaling that the audio chain has entered a "non-steady state imbalance." Therefore, a dramatic increase in the amplitude change rate between consecutive frames is an important dynamic behavioral characteristic for identifying dynamic range imbalance.
[0078] The specific steps for comprehensively analyzing the amplitude change rate between consecutive frames of the audio waveform under the detection window to generate the amplitude gradient mutation index are as follows:
[0079] Within the detection window, the amplitude characteristics of each frame of audio signal (such as short-time energy amplitude or logarithmic amplitude) are extracted. Based on the amplitude changes between consecutive frames, a mutation response factor is generated to sensitively capture the degree and direction of mutations between adjacent frames. The generation formula is as follows:
[0080] ΔΦ a =tanh(λ·|A a+1 -A a | β )·sgn(A a+1 -A a )
[0081] , where A a It is the amplitude feature of the audio signal of the ath frame in the detection window. The optional amplitude features include: Short-time Energy (STE): calculates the sum of the square values of the signal of each frame; Log Energy: takes the logarithm of the energy of each frame to compress the dynamic range; Frequency domain amplitude peak or envelope value: used to emphasize non-stationary segments, A a+1 is the amplitude feature of the audio signal of the a+1th frame in the detection window, that is, the amplitude feature of the next frame of audio signal. λ is the amplitude response sensitivity coefficient, which is used to control the amplification degree of the amplitude difference input before entering the tanh nonlinear function. The value range is 1.5-3. β is the mutation response power exponent, which is used to nonlinearly enhance the asymmetry of the amplitude change. The value range is 1.2-2. sgn(A a+1 -A a) is the mutation direction preservation factor (sign function), which retains the "directional information" of the mutation and is used to distinguish whether the system is subject to "surge interference" or "attenuation interference". When used with tanh and power exponents, it helps to build a complete "mutation scenario feature map". a is the mutation response factor, with a value range of (-1<ΔΦ i <1), if ΔΦ i →+1 indicates that a strong amplitude "increase" mutation occurs between the frame pairs; if ΔΦ i →-1 indicates a strong "downward" mutation; if ΔΦ i ≈0: indicates smooth inter-frame changes and a stable system;
[0082] Tanh is a hyperbolic tangent function with an output range of (-1, 1). It is an S-shaped, continuous, differentiable nonlinear function that is approximately linear for small input values and gradually saturates to ±1 for large input values. The core function of choosing tanh in audio dynamic analysis is to perform nonlinear compression on sudden amplitude changes, making small changes respond linearly and large changes quickly saturate. This prevents the numerical "explosion" caused by abnormal spikes (such as sonic booms) from dominating the entire judgment result, while still maintaining sensitive recognition of medium-amplitude mutations. Compared with functions such as ReLU and sigmoid, tanh retains positive and negative symmetry and boundedness, making it suitable for processing positive and negative amplitude mutations between consecutive frames. It suppresses extreme values without losing directionality, making it an ideal compression function for characterizing dynamic range mutation behavior.
[0083] Sign function (sgn(x)) is a basic piecewise function in mathematics that is used to determine the positive or negative sign of a real number. It is used to ignore the magnitude of the change while preserving the direction of the value change. It is often used in signal processing, mathematical modeling, or control systems to identify the growth, stability, or decline of a value. In this application scenario, sgn(A a+1 -A a ) represents the "directional difference" between the amplitudes of two consecutive audio frames: if the amplitude of the next frame is larger than the current frame, it is +1, indicating a sudden increase; if it is smaller, it is -1, indicating a sudden decrease; if they are equal, it is 0, indicating no change. The main purpose of selecting a sign function is to preserve the upward or downward trend of inter-frame changes when calculating the amplitude mutation index. This allows subsequent judgment to identify not only whether a mutation occurred, but also the direction of the mutation. This is crucial for distinguishing between compression-type and burst-type interference in dynamic range imbalance.
[0084] By constructing a nonlinearly enhanced mutation response factor, we accurately capture the amplitude mutation behavior and directional characteristics between consecutive frames of the audio waveform. This step provides a highly sensitive and directionally sensitive basic feature input for subsequent determination of the severity of dynamic range imbalance.
[0085] After obtaining the mutation response factor ΔΦ of all adjacent frames a After that, the acquired response value is mapped to the nonlinear enhancement domain. Through amplification and integration, a quantitative index representing the sum of the mutation intensity in the entire detection window is formed, namely the amplitude gradient mutation index. The calculation expression is as follows:
[0086]
[0087] , where AVGI is the amplitude gradient mutation index, which represents the overall strength of the amplitude mutation trend between audio frames in the detection window. It is used to determine whether there is a dynamic range imbalance and has a value range of 0-1. γ is the mutation response amplification factor, which controls the mutation response factor |ΔΦ a In the exponential mapping function, the "sensitivity gain" ranges from 1 to 5. e is the natural base, m is the number of frames in the detection window, and Z is the normalization constant used to normalize the accumulated results to ensure that the final output value falls within a controllable range (such as [0, 1], [0, 100]). This prevents AVGI from being affected by the detection window length and ensures that the output is comparable and has a stable numerical range, suitable for different time scales and model input specifications.
[0088] In the AVGI index, the exponential mapping function is It is a nonlinear enhancement function constructed based on the natural exponential function. Its core function is to convert the inter-frame mutation response factor |ΔΦ a Mapping the signal to the (0,1) interval makes the system highly sensitive to large mutations and relatively suppressive to small fluctuations. When the mutation value is small, the exponential term approaches 1, and the activation value approaches 0. When the mutation value is large, the exponential term rapidly approaches 0, and the activation value approaches 1, thereby amplifying abnormal mutation behavior. This function is an inverse variation of the classic exponential decay function, commonly used in signal processing and neural network activation functions. It can significantly improve the system's ability to detect dynamic range imbalances and is a key enhancement mechanism in the AVGI architecture.
[0089] A larger amplitude gradient mutation index, generated by comprehensively analyzing the amplitude change rate between consecutive frames of the audio waveform within the detection window, indicates more severe amplitude fluctuations and increased waveform discontinuity between adjacent frames. This typically reflects that the system experienced abnormal sound intensity interference (such as a sonic boom or severe signal disturbance) during a certain period of time, triggering a nonlinear response in the gain control mechanism and chaotic compressor adjustments, thereby disrupting the amplitude balance between channels and leading to a chain reaction of dynamic range imbalance. Therefore, a larger amplitude gradient mutation index value indicates a strong tendency towards system imbalance. Conversely, a low and stable amplitude gradient mutation index indicates stable amplitude changes between audio frames and good system dynamic compression consistency. In this case, the audio processing chain is operating normally and no imbalance has occurred.
[0090] The audio system status intelligent assessment module inputs the comprehensively analyzed indicators as feature vectors into a pre-trained machine learning model (such as a random forest, XGBoost, CNN, or LSTM network). The model then performs a rapid and intelligent assessment of the current operating status of the audio processing system to determine whether the system has entered a state of dynamic range imbalance.
[0091] The comprehensively analyzed channel dynamic compression imbalance index and amplitude gradient mutation index are input as feature vectors into a pre-trained machine learning model (such as random forest, XGBoost, CNN or LSTM network). The link dynamic range imbalance coefficient is generated by the model. Based on the link dynamic range imbalance coefficient, the current operating status of the audio processing system is quickly and intelligently evaluated to determine whether the current audio processing system has entered a dynamic range imbalance state.
[0092] A "pre-trained machine learning model" refers to a discriminative model built and trained before the audio processing system is deployed, using a large amount of historical audio processing data, simulated interference samples, or real user interaction data. This discriminative model is used to identify the dynamic state of the audio link. This model learns from a large amount of feature data (such as the channel dynamic compression imbalance index and the amplitude gradient mutation index) to extract potential characteristic patterns and classification boundaries, enabling it to identify whether the audio link is in a state of normal dynamic range, mild imbalance, or severe imbalance. The training process typically employs model structures such as random forests, gradient boosting trees (XGBoost), convolutional neural networks (CNNs), or long short-term memory networks (LSTMs). Supervised learning is performed using labeled training datasets, and through continuous iterative optimization, the model learns the nonlinear mapping relationship between imbalance and feature data. This training process is typically completed offline on a high-performance computing platform, ultimately outputting a "deployable model" with parameter freezing and verified generalization capabilities for subsequent online inference tasks.
[0093] During actual operation, the model requires no retraining and is instead invoked in real time as an embedded intelligent decision-making module within the Bluetooth headset audio processing system. After the Channel Dynamic Compression Imbalance Index (CDCI) and Amplitude Gradient Hurt Index (AVGI) are calculated in real time, they are packaged into a feature vector and fed into the pre-trained model. The model then immediately outputs a link dynamic range imbalance coefficient, reflecting the current system state. This coefficient is a quantitative assessment, and a threshold can be set to determine whether the current imbalance is present. For example, an imbalance coefficient greater than 0.7 may indicate "high risk," while a value less than 0.3 indicates "stable." This coefficient serves as the basis for subsequent intervention decisions (such as adjusting source weights and switching gain strategies). Because the model has already learned various complex scenarios during training, including different user voice styles, background noise types, and interference intensity, it exhibits strong robustness and adaptability. In short, the introduction of this pre-trained model provides the Bluetooth headset system with data-driven intelligent state awareness and response capabilities, significantly improving the stability and self-regulation of the audio link in dynamic and complex environments.
[0094] The link dynamic range imbalance coefficient generated by a pre-trained machine learning model during a rapid and intelligent assessment of the current operating status of the audio processing system is compared with a pre-set link dynamic range imbalance coefficient reference threshold to determine whether the current audio processing system has entered a dynamic range imbalance state. The judgment logic is as follows:
[0095] If the link dynamic range imbalance coefficient is greater than a preset link dynamic range imbalance coefficient reference threshold, it is determined that the current audio processing system has entered a dynamic range imbalance state; if the link dynamic range imbalance coefficient is less than or equal to a preset link dynamic range imbalance coefficient reference threshold, it is determined that the current audio processing system has not entered a dynamic range imbalance state.
[0096] The adaptive sound source weighting control module immediately adjusts the sound source priority upon detecting a dynamic range imbalance in the audio processing system. It temporarily sets the sound source channel identified as having strong interference characteristics to "low priority" and weakens its weight in the mixed output. At the same time, it automatically increases the proportion of remaining sound sources in the overall output, suppressing the interference of sudden anomalies on normal conversation channels.
[0097] When a dynamic range imbalance is detected in the audio processing system, the dynamic adjustment mechanism for sound source priority is immediately implemented. Its core function is to quickly block the destructive impact of abnormal sound sources on the overall audio processing chain and stably maintain the output quality of normal voice signals and system availability. In a multi-source co-converging Bluetooth headset voice translation system, each sound source channel operates collaboratively within the same processing framework, and the system generally gives equal weight to each channel by default. However, in sudden abnormal situations, such as sonic booms, strong noise, or sharp non-speech interference from a sound source, its high-amplitude signal will be mistakenly interpreted by the system as an overall energy increase, inducing abnormal triggering of the automatic gain control (AGC) or excessive response of the dynamic compression mechanism, thereby suppressing the output amplitude of all channels, causing speech recognition failure, translation interruption, or protective system shutdown.
[0098] To this end, this step uses an intelligent recognition mechanism to mark the abnormal sound source channel as "low priority" and actively weakens its weight in the mixing output. This operation is equivalent to "non-shielded suppression" of the channel - it does not directly cut off its input data stream, but reduces its participation and influence on the system decision-making link, thereby preventing it from continuing to drive the AGC response, polluting the mixing output, or interfering with semantic judgment. At the same time, the system also automatically increases the output proportion of other sound source channels in normal state, strengthens the signal expression of the main semantic channel at the mixing level, and ensures that the system's speech recognition, machine translation, speech synthesis and other modules can still receive sufficiently clear and decodable audio information. This "interference suppression + main semantic enhancement" dynamic sound source control strategy not only avoids the system from falling into a state of overall suppression collapse, but also significantly improves the fault tolerance and robustness of speech understanding in complex scenarios.
[0099] In summary, the fundamental purpose of this step is to introduce adaptive intervention logic based on the behavioral characteristics of the sound source. Without interrupting all channels, this approach reduces the weight of high-risk channels and preserves the semantics of low-risk channels, thus achieving the system's "partial isolation, global stability" operational strategy in highly complex, multi-interference, and multi-speaker scenarios. This mechanism is not only suitable for handling sudden interference but also provides operational space for subsequent modules such as voice signal reconstruction, channel recovery, and self-learning optimization. It is a key disaster recovery and self-healing control link in the entire audio processing system.
[0100] When a dynamic range imbalance is detected in the audio processing system, the priority of the sound source is immediately adjusted. The sound source channel identified as having strong interference characteristics is temporarily set to "low priority" and its weight in the mixed output is weakened. At the same time, the proportion of the remaining sound sources in the overall output is automatically increased. The specific steps are as follows:
[0101] After identifying a dynamic range imbalance in the audio processing system (i.e., the link dynamic range imbalance coefficient is greater than a preset threshold), the interference intensity of all sound source channels in the audio processing system is first evaluated. The interference index of each channel is calculated to quantify its impact on the imbalance state. The interference index calculation expression is as follows:
[0102] ψ i =α·ΔA i +β·δH i
[0103] , where ψ i Is the interference index, which represents the comprehensive interference intensity score of the i-th sound source channel causing the dynamic range imbalance. It is used as a direct criterion to judge whether the sound source is a "strong interference source". The larger the value, the greater the impact of the channel on the current imbalance state. ΔA i is the amplitude mutation intensity, which represents the maximum rate of change of the speech signal amplitude between adjacent frames of the sound source channel i, δH i Is the gain offset, which indicates the absolute amplitude of the gain change before and after the gain control module (such as AGC or DRC) of the sound source channel i. If a sound source triggers a strong response from the AGC, its gain will jump significantly. δH i The value will also increase accordingly, which is the characteristic index of the adjustment offset for identifying the abnormal gain of the system. α is the amplitude mutation weight coefficient, which controls the amplitude mutation characteristic ΔA i The contribution of the interference index is set by the system design parameters and can be configured based on model training or experience. For example, α = 0.6 indicates that the mutation feature is dominant, which is used to flexibly adjust the system's sensitivity to "acoustic slope interference";
[0104] The interference index reflects whether there are sudden strong interference signals in each channel, providing a quantitative basis for subsequent priority control.
[0105] After obtaining the interference index ψ of all channels i After that, the weight of each sound source channel in the mixed output is dynamically adjusted to weaken the output contribution of the strong interference sound source and increase the proportion of the remaining channels. The weight adjustment formula is as follows:
[0106]
[0107] , where W i It is the sound source channel mixing output weight, which represents the weight value of the i-th sound source channel in the current audio mixing output. It directly determines the proportion of the channel's voice signal in the final output of the system. The value range is between 0 and 1. The larger the value, the more prominent the channel is, and the smaller the value, the more its output will be suppressed. It is used to dynamically adjust the channel output contribution, reduce the output impact of interfering sound sources, and improve the overall recognition stability of the system under the state of dynamic range imbalance. It is the average value of the interference index of all sound source channels, which is used to construct the benchmark value so that the interference intensity of each channel can be evaluated in the form of "center offset". It is used to normalize the interference degree of the current channel and serves as the input of the Sigmoid function, so that the control behavior presents a nonlinear "suppress strong interference and retain weak signals" mechanism. d is the nonlinear inhibition factor, which controls the weight adjustment function (Sigmoid) to The steepness of the response curve is used to construct a dynamic nonlinear weight reduction curve, so that the system can tolerate "slight interference" and quickly suppress "extreme interference" to achieve gradient adaptation of the response. It is a positive number. The larger the value, the higher the function steepness and the more radical the adjustment result. The smaller the value, the smoother the function transition. e is the natural base number, Λ is the link dynamic range imbalance coefficient, and Λ t is the link dynamic range imbalance coefficient reference threshold;
[0108] The Sigmoid function in the weight adjustment function is a commonly used S-shaped nonlinear mapping function, and its mathematical expression is:
[0109]
[0110] Its output value always lies between (0, 1) and exhibits the following characteristics: when the input x is very small, the output approaches 0; when x is very large, the output approaches 1; and it is most sensitive to changes in the range close to 0, exhibiting a steep curve transition. Therefore, the Sigmoid function is often used to smoothly compress any input value into a controllable range, forming a nonlinear response pattern with a moderate center and extreme ends.
[0111] In audio weight adjustment, the Sigmoid function maps the deviation between the interference intensity of each channel and the average interference level, so that sound sources with slightly higher interference indices are slightly suppressed, while sound sources with significantly higher interference indices are significantly suppressed, thus establishing a "nonlinear, progressive" suppression strategy. Compared to linear functions, the Sigmoid function avoids over-responding to small fluctuations while providing a quick and powerful response to sudden strong interference. This makes it a well-suited function model for building dynamic suppression weight adjustment mechanisms.
[0112] This formula constructs an inhibitory weighting mechanism based on interference intensity and introduces an imbalance degree control factor, so that the system suppresses more strongly when the imbalance degree is high and adjusts more gently when the imbalance degree is slight.
[0113] After completing the weight adjustment of each channel, in order to ensure the amplitude consistency and signal ratio accuracy of the mixed output, the sound source channel mixed output weight W i Normalization is performed and the mixed output signal is reconstructed according to the normalized weights. The calculation formula is as follows:
[0114]
[0115] , where Y min It is the mixed output signal, which represents the final composite voice signal output by the system. It is the result of weighted synthesis of multiple sound source channels. It is a voice waveform data stream that changes continuously over time. j It represents the weight value of the jth sound source channel in the current audio mixing output, n is the total number of sound source channels, S i is the original speech signal of the i-th sound source channel, and is the unprocessed audio data collected from each sound source channel.
[0116] This formula ensures that the energy distribution of the total output audio remains logically consistent, while minimizing the interference of abnormal channels on the final audio results, significantly improving the stability and accuracy of speech recognition and translation systems in unbalanced states.
[0117] By introducing key modules such as audio dynamic state perception, feature preprocessing, abnormal feature quantification analysis, intelligent state assessment, and adaptive weight control, the present invention effectively achieves real-time identification and adaptive control of the dynamic range imbalance state of the audio processing link, thereby significantly improving the system's stability and robustness in complex multi-sound source environments. The system constructs a high-quality feature dataset through audio dynamic state acquisition and preprocessing, and intelligently assesses the audio state by combining feature engineering and machine learning models. When a sonic boom or abnormal sound source disturbance occurs, it can quickly identify the interference source and dynamically adjust its weight to prevent its interference from spreading to the global AGC mechanism, thereby ensuring the continuity and translatability of other normal voice channels. Compared with the disadvantage of traditional systems that are susceptible to high-energy interference and overall failure, this system implements a differentiated control strategy of "locally suppressing the interference channel and maintaining the stability of the main voice link", effectively avoiding problems such as translation task interruption and speech recognition failure, and significantly improving the practicality and interaction quality of Bluetooth headsets in multi-speaker, multi-language interaction scenarios.
[0118] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0119] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.
[0120] It should be noted that, in this document, if there are relational terms such as first and second, etc., they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0121] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0122] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0123] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0124] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0125] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0126] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0127] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.
Claims
1. A multi-voice co-convergence anti-interference Bluetooth headset translation system, characterized in that: It includes an audio dynamic state acquisition module, an audio feature preprocessing and data construction module, a dynamic range abnormality feature extraction and quantitative analysis module, an audio system state intelligent assessment module, and a sound source weight adaptive control module: The audio dynamic state acquisition module monitors and captures the audio dynamic range state characteristic data generated by each sound source channel in real time during the Bluetooth headset's real-time language translation task through the high-precision microphone array, digital signal processing chip, and audio processing module integrated into the headset; The audio feature preprocessing and data construction module preprocesses the raw audio dynamic range feature data collected in real time to establish an efficient and reliable data set to support subsequent feature engineering and intelligent analysis processes; The dynamic range anomaly feature extraction and quantitative analysis module applies feature engineering techniques to the preprocessed data set to accurately extract key indicators that represent the overall dynamic range imbalance of the audio processing chain. It then conducts in-depth comprehensive analysis of the extracted key indicators to quantify the severity of the overall imbalance in the audio processing chain. The audio system status intelligent assessment module inputs the comprehensively analyzed indicators as feature vectors into a pre-trained machine learning model. The model then conducts a rapid and intelligent assessment of the current operating status of the audio processing system to determine whether the current audio processing system has entered a state of dynamic range imbalance. The sound source weight adaptive control module immediately adjusts the sound source priority upon detecting a dynamic range imbalance in the audio processing system. It temporarily sets the sound source channel identified as having strong interference characteristics to "low priority" and weakens its weight in the mixed output. At the same time, it automatically increases the proportion of remaining sound sources in the overall output, suppressing the interference of sudden anomalies on normal conversation channels.
2. The multi-voice co-convergence anti-interference Bluetooth headset translation system according to claim 1, characterized in that: When a Bluetooth headset performs a real-time language translation task, the specific steps for obtaining audio dynamic range state feature data are as follows: The high-precision microphone array inside the headset synchronously collects multi-channel original voice signals from different directions or speakers to ensure coverage of all sound source information; The collected analog voice signal is input into the digital signal processing chip, which performs A / D conversion, signal enhancement, and preliminary noise reduction to generate a computable digital audio stream; The audio processing module analyzes the digital audio stream and extracts the state characteristics related to the dynamic range of the automatic gain control response parameters of each channel; The feature data is cached in a structured time series format, serving as the input basis for subsequent dynamic range imbalance identification, speech interference judgment, and intelligent translation stability control.
3. The multi-voice co-convergence anti-interference Bluetooth headset translation system according to claim 1, characterized in that: Feature engineering technology is applied to the preprocessed dataset to accurately extract key indicators that characterize the overall dynamic range imbalance of the audio processing chain. The extracted indicators include the output amplitude change after the dynamic compressor responds and the amplitude change rate between consecutive frames of the audio waveform. The output amplitude change after the dynamic compressor responds and the amplitude change rate between consecutive frames of the audio waveform are comprehensively analyzed under the detection window to generate the channel dynamic compression imbalance index and amplitude gradient mutation index respectively. The channel dynamic compression imbalance index and amplitude gradient mutation index are used to measure the severity of the overall imbalance of the audio processing chain.
4. The multi-voice co-convergence anti-interference Bluetooth headset translation system according to claim 3, characterized in that: The specific steps for comprehensively analyzing the output amplitude changes after the dynamic compressor response in the detection window to generate the channel dynamic compression imbalance index are as follows: Within the detection window, for all valid sound source channels in the Bluetooth headset, the maximum compression gradient value in the output amplitude envelope curve of each channel after being processed by the dynamic compressor is obtained. Based on the maximum compression gradient value of all channels, the compression difference mapping matrix M between channels is constructed, which is defined as follows: , Where M i,j It is an element in the compression difference mapping matrix M, which represents the normalized difference between the compression response strength of the sound source channel i and the sound source channel j, i∈1,2,3,…,n, n is the total number of sound source channels, G i is the maximum compression gradient value of the sound source channel i, G j is the maximum compression gradient value of the sound source channel j, max(G i , G j ) is the normalized reference factor, which represents the maximum compression gradient value of the sound source channel i and the sound source channel j; After obtaining the compression difference mapping matrix M, statistical analysis is performed on all channel combinations to calculate the channel dynamic compression imbalance index. The calculation expression is as follows: , Where CDCI is the channel dynamic compression imbalance index, is the total number of non-repeated channel pairs, σ(M i,j ) is a nonlinear enhancement function, which means that for each M i,j A nonlinear transformation is applied to enhance the weights of medium and high degree differences. The calculation method is as follows: σ(x) = tank(k·x), where k is the amplification factor used to adjust the response sensitivity to differential mutations, and tanh(x) is the hyperbolic tangent function.
5. The multi-voice co-convergence anti-interference Bluetooth headset translation system according to claim 3, characterized in that: The specific steps for comprehensively analyzing the amplitude change rate between consecutive frames of the audio waveform under the detection window to generate the amplitude gradient mutation index are as follows: Within the detection window, the amplitude characteristics of each frame of audio signal are extracted, and the mutation response factor is generated based on the amplitude change between consecutive frames. The generation formula is as follows: DF a =tanj(λ·|A a+1 -A a | β )·sgn(A a+1 -A a ), Where A a is the amplitude feature of the audio signal of the ath frame in the detection window, A a+1 is the amplitude feature of the audio signal of the a+1th frame in the detection window, that is, the amplitude feature of the next frame of audio signal, λ is the amplitude response sensitivity coefficient, β is the mutation response power exponent, sgn(A a+1 -A a ) is the mutation direction maintaining factor, ΔΦ a is the mutation response factor; After obtaining the mutation response factor ΔΦ of all adjacent frames a After that, the acquired response value is mapped to the nonlinear enhancement domain. Through amplification and integration, a quantitative index representing the sum of the mutation intensity in the entire detection window is formed, namely the amplitude gradient mutation index. The calculation expression is as follows: , Where AVGI is the amplitude gradient mutation index, γ is the mutation response amplification coefficient, e is the natural base, m is the number of frames contained in the detection window, and Z is the normalization constant.
6. The multi-voice co-convergence anti-interference Bluetooth headset translation system according to claim 3, characterized in that: The comprehensively analyzed channel dynamic compression imbalance index and amplitude gradient mutation index are input as feature vectors into a pre-trained machine learning model. The link dynamic range imbalance coefficient is generated by the model. Based on the link dynamic range imbalance coefficient, a rapid and intelligent evaluation of the current operating status of the audio processing system is performed to determine whether the current audio processing system has entered a dynamic range imbalance state.
7. The multi-voice co-convergence anti-interference Bluetooth headset translation system according to claim 6, characterized in that: The link dynamic range imbalance coefficient generated by a pre-trained machine learning model during a rapid and intelligent assessment of the current operating status of the audio processing system is compared with a pre-set link dynamic range imbalance coefficient reference threshold to determine whether the current audio processing system has entered a dynamic range imbalance state. The judgment logic is as follows: If the link dynamic range imbalance coefficient is greater than a preset link dynamic range imbalance coefficient reference threshold, it is determined that the current audio processing system has entered a dynamic range imbalance state; if the link dynamic range imbalance coefficient is less than or equal to a preset link dynamic range imbalance coefficient reference threshold, it is determined that the current audio processing system has not entered a dynamic range imbalance state.
8. The multi-voice co-convergence anti-interference Bluetooth headset translation system according to claim 7, characterized in that: When a dynamic range imbalance is detected in the audio processing system, the priority of the sound source is immediately adjusted. The sound source channel identified as having strong interference characteristics is temporarily set to "low priority" and its weight in the mixed output is weakened. At the same time, the proportion of the remaining sound sources in the overall output is automatically increased. The specific steps are as follows: After identifying the dynamic range imbalance in the audio processing system, the interference intensity of all sound source channels in the audio processing system is first evaluated. The interference index of each channel is calculated to quantify its impact on the imbalance state. The interference index calculation expression is as follows: ψ i =a·ΔA i +β·δH i , Where, ψ i is the interference index, which represents the comprehensive interference intensity score of the i-th sound source channel causing the dynamic range imbalance, ΔA i is the amplitude mutation intensity, which represents the maximum rate of change of the speech signal amplitude between adjacent frames of the sound source channel i, δH i is the gain offset, α is the amplitude mutation weight coefficient; After obtaining the interference index ψ of all channels i After that, the weight of each sound source channel in the mixed output is dynamically adjusted to weaken the output contribution of the strong interference sound source and increase the proportion of the remaining channels. The weight adjustment formula is as follows: , Where W i Is the sound source channel mixing output weight, which represents the weight value of the i-th sound source channel in the current audio mixing output. is the average value of the interference index of all sound source channels, d is the nonlinear suppression factor, e is the natural base, Λ is the link dynamic range imbalance coefficient, Λ t is the link dynamic range imbalance coefficient reference threshold; After completing the weight adjustment of each channel, in order to ensure the amplitude consistency and signal ratio accuracy of the mixed output, the sound source channel mixed output weight W i Normalization is performed and the mixed output signal is reconstructed according to the normalized weights. The calculation formula is as follows: , Where Y min is the mixed output signal, W j It represents the weight value of the jth sound source channel in the current audio mixing output, n is the total number of sound source channels, S i is the original speech signal of the i-th sound source channel.
Citation Information
Patent Citations
Collaboratively processing audio between headset and source to mask distracting noise
CN106464998A
Multi-voice common exchange type anti-interference Bluetooth earphone translation system
CN117292702A
TWS Bluetooth earphone control method
CN117294985A
Two-channel trunking terminal voice intercommunication method
CN118250738A
Earphone intelligent noise reduction method and device based on AIGC
CN118338184A
Cited By
Sound cloning method for translation earphone
CN120673742A
A method for sound cloning in translation headphones
CN120673742B