Speech translation system based on Bluetooth earphone
By introducing speech speed detection, rhythm evaluation and speech control modules into the Bluetooth headset-based speech translation system, the speech recognition error problem caused by user speech speed fluctuations is solved, and a more accurate and coherent speech translation effect is achieved.
Patent Information
- Application Number
- CN202510639144.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
Smart Images

Figure CN120164455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice translation for Bluetooth headsets, and particularly to a voice translation system based on Bluetooth headsets. Background Art
[0002] A voice translation system based on Bluetooth headsets is an intelligent communication system that integrates voice recognition, machine translation, and speech synthesis technologies. Its core lies in the real-time collection, processing, and output of voice through Bluetooth headsets. Users can transmit the spoken language to a mobile terminal or edge computing device in real time by wearing Bluetooth headsets. The system automatically recognizes the voice content, performs cross-language translation, and feeds back the target language result in the form of voice, thus achieving instant communication between two languages or even multiple languages. This system has the characteristics of high portability, hands-free operation, and natural operation, and is particularly suitable for scenarios such as international travel, international business, foreign language teaching, and auxiliary communication for special groups (such as hearing-impaired people). Its significance lies in breaking down language barriers and improving the communication efficiency between people, and it is one of the important directions for the integrated application of artificial intelligence and wearable devices.
[0003] The existing voice translation technology based on Bluetooth headsets usually uses the Bluetooth headset as the front-end voice collection and output device, and cooperates with an intelligent terminal (such as a smartphone or a dedicated translation device) to achieve a complete voice translation process. Its voice translation process mainly includes the following key links: First, the user speaks the original language through the Bluetooth headset, and the headset transmits the voice signal to the terminal device in the form of audio data; Second, the built-in voice recognition module (ASR) in the terminal extracts features and performs modeling analysis on the received audio data, and converts it into the corresponding text content; Subsequently, the translation engine (MT) semantically understands the source language text and converts it into the text information of the target language, usually combined with neural network machine translation technology to improve accuracy and naturalness; Then, the speech synthesis module (TTS) converts the target language text into a natural voice signal; Finally, the translated voice is transmitted back to the user through the Bluetooth headset to achieve a complete cross-language interaction from voice to voice. The entire system may also be equipped with pre-processing modules such as voice enhancement, noise reduction, and echo cancellation to ensure high translation accuracy and usability in complex environments.
[0004] The existing technology has the following deficiencies: During the process of a user using a Bluetooth headset for voice translation, if there are obvious fluctuations in the speaking speed during their expression, such as the speaking speed suddenly slowing down significantly when organizing language, or resuming expression quickly after a pause, it will cause a sudden change in the rhythm of the voice signal in the time dimension. Such situations are common when the user is speaking while thinking, the expression content is complex, or the sentence structure is changed temporarily. It is difficult to maintain a stable time interval and beat structure of the voice. Since the Bluetooth headset is only a front-end audio acquisition device and does not have the function of perceiving rhythm changes itself, and the voice processing module at the back end of the system usually segments and processes the audio stream based on fixed frame cutting windows, silence thresholds, and endpoint detection algorithms, frame cutting errors are likely to occur when the speaking speed changes suddenly. For example, a semantically complete speech segment may be misclassified into two segments, or two semantically independent speech segments may be wrongly merged, resulting in misalignment of the input to the speech recognition model and problems such as syllable miscutting and phoneme overlapping. The existing voice translation technology based on Bluetooth headsets cannot dynamically adjust the voice processing strategy according to the intensity of the sudden change in the sentence rhythm when there are obvious fluctuations in the user's speaking speed during the expression, and still performs unified frame processing according to the conventional speaking speed mode, which will cause incomplete extraction of voice features or mixing of information between upper and lower sentences, and then lead to speech recognition errors, such as disordered word order, missing keywords, and broken semantic structures, seriously affecting the accuracy of the final translation content and the continuity of user communication.
[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The object of the present invention is to provide a voice translation system based on a Bluetooth headset to solve the problems in the above background art.
[0007] To achieve the above object, the present invention provides the following technical solution: A voice translation system based on a Bluetooth headset, comprising a voice acquisition module, a speaking speed detection module, a rhythm evaluation module, a voice regulation module, and a text verification module; The voice acquisition module collects the user's voice signal through the Bluetooth headset, preprocesses the collected voice signal, and divides the voice signal into a continuous frame sequence according to a preset frame length and frame shift after preprocessing; The speaking speed detection module establishes a time window based on the continuous frame sequence, analyzes the change state of the speaking speed within the time window, determines the time period when there are obvious fluctuations in the user's speaking speed during the expression, and labels this time period as the speaking speed fluctuation time period, where the obvious speaking speed fluctuation means that the fluctuation amplitude of the speaking speed exceeds a preset threshold; A rhythm evaluation module extracts rhythm feature information from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, analyzes it, and evaluates the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period; A voice control module selects and executes corresponding voice processing strategies according to the evaluation results and performs dynamic control during the execution process; A text verification module, after completing speech recognition, verifies and adjusts the language structure and word order of the recognized text to ensure that the translated text conforms to the grammar rules of the target language and maintains semantic coherence.
[0008] Preferably, in the speech rate detection module, the size of the time window is set based on the continuous frame sequence, and the time window is divided into several time periods; by analyzing the speech rate change state of each time period, specifically including analyzing the time interval between adjacent frames and the change in speech energy, the speech rate fluctuation amplitude of each time period is calculated; the time period with the speech rate fluctuation amplitude exceeding the preset threshold is determined as the time period when the user's speech rate fluctuates significantly during the expression, and it is marked as the speech rate fluctuation period; the speech rate fluctuation amplitude is determined by calculating the ratio of the time interval between adjacent frames to the difference in speech energy.
[0009] Preferably, in the rhythm evaluation module, rhythm feature information is extracted from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, and after acquisition, it is preprocessed; acoustic fluctuation feature information and spectral configuration distribution information are extracted from the preprocessed rhythm feature information, and after extraction, they are analyzed to generate a speech energy fluctuation coefficient and a speech spectrum dispersion index respectively; a rhythm mutation intensity evaluation model is constructed for the generated speech energy fluctuation coefficient and speech spectrum dispersion index, and a rhythm mutation index is generated through weighted summation; a preset rhythm mutation index threshold interval is determined, and after determination, it is compared with the generated rhythm mutation index, and the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period is evaluated according to the comparison result.
[0010] Preferably, the acquisition logic of the speech energy fluctuation coefficient is as follows: Acoustic fluctuation feature information is extracted from the preprocessed rhythm feature information, specifically including the energy value and zero-crossing rate of each frame in the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, and they are respectively marked as and , represents the energy value of the th frame in the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, represents the zero-crossing rate of the th frame in the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, , is a positive integer; Calculate the voice energy fluctuation coefficient. The specific calculation formula is as follows: In the formula, is the voice energy fluctuation coefficient.
[0011] Preferably, the acquisition logic of the voice spectrum discrete index is as follows: Extract the spectrum configuration distribution information from the preprocessed rhythm feature information, specifically including: obtaining the spectrum feature information of each frame in the user voice signal collected by the Bluetooth headset during the speech rate fluctuation period, and constructing a Mel frequency cepstral coefficient vector based on the spectrum feature information of each frame, expressed as: In the formula, is the Mel frequency cepstral coefficient vector of the th frame, including components, represents the Mel frequency cepstral coefficient corresponding to the th cepstral component in the th frame, , , and are all positive integers; Based on the Mel frequency cepstral coefficient vector of each frame, calculate the following two types of characteristic values: First, the average offset value of the cepstral feature, according to the formula: In the formula, is the average value of the Mel frequency cepstral coefficients corresponding to the th cepstral component in all frames, represents the average offset value of the cepstral feature of the th frame; Second, the dispersion value of the cepstral distribution, according to the formula: In the formula, is the mean value of the Mel frequency cepstral coefficients corresponding to all cepstral components in the th frame, is the dispersion value of the cepstral distribution of the th frame; Calculate the voice spectrum discrete index. The specific calculation formula is as follows: In the formula, is the discrete index of the speech spectrum.
[0012] Preferably, for the generated speech energy fluctuation coefficient and the discrete index of the speech spectrum construct a rhythm mutation intensity evaluation model, and generate a rhythm mutation index through weighted summation. The specific calculation formula is as follows: In the formula, is the rhythm mutation index, and are the non-zero weight coefficients of the speech energy fluctuation coefficient and the discrete index of the speech spectrum respectively, and .
[0013] Preferably, determine a preset rhythm mutation index threshold interval , and after determination, compare it with the generated rhythm mutation index to evaluate the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period according to the comparison result. The specific comparison and analysis are as follows: If , the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period is low intensity; If , the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period is medium intensity; If , the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period is high intensity.
[0014] Preferably, in the speech regulation module, according to the evaluation result, select and execute the corresponding speech processing strategy, specifically: If the evaluation result is low intensity, the speech processing strategies selected and executed include: using the default frame cutting window parameters for speech frame division, maintaining the standard silent duration threshold and the fixed endpoint detection method, and completing the speech segment definition in a conventional manner; If the evaluation result is medium intensity, the speech processing strategies selected and executed include: adjusting the size of the speech frame cutting window to increase the window duration, increasing the frame overlap ratio, and extending the silent duration threshold to alleviate the risk of misjudgment of segmentation caused by rhythm fluctuations; If the evaluation result is high intensity, the speech processing strategies selected and executed include: on the basis of adjusting the frame cutting window and the silent threshold, enabling the cross-segment semantic context retention mechanism and enabling the delayed endpoint judgment mode to avoid the occurrence of speech segment cutting errors and recognition timing disorders at the rhythm jump points; And dynamic regulation is carried out during the execution process, specifically as follows: During the execution of the speech processing strategy, continuously monitor the changing trend of the frame-level spectral energy of the speech signal and the changing state of the silent duration between speech segments. If it is detected during the execution of the speech processing strategy that the rhythm mutation index continuously rises in consecutive frames and exceeds the maximum threshold corresponding to the current evaluation level, immediately switch to a higher-level speech processing strategy, update the speech frame cutting window parameters and the silent duration threshold, and enable the cross-segment semantic preservation function; if it is detected during the processing that the rhythm mutation index continuously decreases in consecutive frames and is lower than the minimum threshold corresponding to the current strategy level, downgrade the current speech processing strategy, restore the default frame division parameters, and turn off the extended processing mechanism to achieve real-time dynamic adjustment of the processing strategy.
[0015] In the above technical solution, the technical effects and advantages provided by the present invention are as follows: 1. By constructing a speech rate detection module and a rhythm evaluation module, the present invention realizes the accurate recognition of the speech rate fluctuation and rhythm mutation of the user during the speech expression process. Especially in the scenario of voice translation with Bluetooth headsets, for the problem of unstable speech rate caused by the user's thinking, pausing or sentence pattern switching, it can monitor and judge the fluctuation characteristics in real time at the frame level of the speech signal. By introducing two composite acoustic indexes, namely the speech energy fluctuation coefficient and the speech spectrum dispersion index, and constructing a rhythm mutation index for intensity quantization evaluation, the solution realizes the mathematical modeling and discrimination ability of the degree of sentence rhythm change at the technical level, provides an accurate decision-making basis for the subsequent processing strategy, and effectively breaks through the limitation in the prior art that the rhythm mutation cannot be perceived and quantified.
[0016] 2. By linking the evaluation result with the speech processing strategy through the speech regulation module, and automatically switching to different combinations of parameters at different levels according to the rhythm mutation intensity, including frame cutting window adjustment, silent duration determination delay, and context semantic preservation mechanism, etc., the present invention realizes the refinement, dynamicization and scene adaptability of speech segmentation processing. Especially in the case of high-intensity mutation, the system not only avoids the problems of sentence cutting and semantic disorder caused by inaccurate frame division under the traditional strategy, but also improves the robustness of endpoint recognition through the delayed segmentation mechanism, makes the recognition flow of the whole speech more smooth, and enhances the fault tolerance ability and translation stability of the system for complex speech structures.
[0017] 3. After the voice recognition is completed, the text verification module of the present invention performs software-level structural verification and semantic reconstruction on the language structure and word order of the recognized text, so that the translation output not only has grammatical compliance, but also maintains logical coherence and language naturalness, further enhancing the user-end perception experience. Without relying on additional hardware devices, the overall solution completes a complete closed-loop processing chain from front-end detection, process regulation to back-end semantic optimization based on the voice acquisition channel of the Bluetooth headset, significantly improving the intelligent response ability of the voice translation system to complex rhythm expressions, and having strong practical value and industrial promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0019] Figure 1 It is a schematic diagram of the modules of a voice translation system based on a Bluetooth headset according to the present invention; Figure 2 It is a system mind map of a voice translation system based on a Bluetooth headset according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Now, the exemplary embodiments will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these exemplary embodiments are provided so that the present disclosure will be more complete and comprehensive, and will fully convey the concept of the exemplary embodiments to those skilled in the art.
[0021] The present invention provides a voice translation system based on a Bluetooth headset as shown in Figure 1 and Figure 2 , which includes a voice acquisition module, a speech rate detection module, a rhythm evaluation module, a voice regulation module, and a text verification module; The voice acquisition module collects the user's voice signal through the Bluetooth headset, preprocesses the collected voice signal, and divides the voice signal into a continuous frame sequence according to a preset frame length and frame shift after preprocessing; The Bluetooth headset collects the user's voice signal through a built-in microphone and usually outputs it as a raw audio stream. Before using the voice signal for subsequent processing, preprocessing is required to improve the signal quality. The first step of preprocessing is usually denoising, which can remove background noise through frequency-domain filtering or adaptive filtering algorithms. Next, pre-emphasis processing is performed, applying a high-pass filter to enhance the high-frequency part of the voice signal, which can compensate for the attenuation of the voice signal in the high-frequency range and improve the clarity and details of the signal. Finally, signal normalization processing is carried out, by adjusting the amplitude range of the audio signal to make it consistent, avoiding the influence of too large or too small signal amplitude on subsequent analysis and processing. Through these steps, preprocessing ensures the quality of the voice signal and provides a stable basis for subsequent analysis and feature extraction.
[0022] After the preprocessing of the voice signal is completed, the next step is to divide the audio signal according to the preset frame length and frame shift. A common approach is to use window function techniques, such as the short-time Fourier transform (STFT), to divide the audio signal into frames of multiple short-time windows. The frame length determines the number of sample points contained in each frame, usually 20 - 40 milliseconds, while the frame shift represents the step size of each window slide, often 10 milliseconds or 20 milliseconds. In this way, the voice signal is cut into multiple overlapping frames, and each frame represents the signal characteristics within a period of time. The division of frames not only helps the subsequent feature extraction work (such as MFCC, spectrogram, etc.) proceed smoothly, but also ensures that the dynamic changes of the voice can be captured within a smaller time window, avoiding the loss of important information due to too large a time window.
[0023] The purpose of doing this is to improve the accuracy of subsequent voice processing modules (such as speech recognition or translation). Through preprocessing, noise and unnecessary frequency components are removed, and the voice signal becomes cleaner and clearer, facilitating subsequent analysis. Frame division cuts the continuous voice signal into manageable small segments, facilitating more accurate feature extraction and analysis by the model in each time period. The information within each frame can be processed independently, which can avoid the loss of information during changes in speech rate or pauses, thus ensuring that sudden changes or long pauses in the voice signal do not lead to recognition errors or inaccurate translations. Therefore, these processing steps are crucial for improving the accuracy and stability of the speech translation system.
[0024] The speech rate detection module establishes a time window based on a continuous frame sequence, analyzes the speech rate change state within the time window, determines the time period when the user has an obvious speech rate fluctuation during expression, and labels this time period as the speech rate fluctuation time period. An obvious speech rate fluctuation means that the speech rate fluctuation amplitude exceeds a preset threshold; In this embodiment, in the speech rate detection module, the size of the time window is set based on the continuous frame sequence, and the time window is divided into several time periods; by analyzing the speech rate change state of each time period, specifically including analyzing the time interval between adjacent frames and the change in speech energy, the speech rate fluctuation amplitude of each time period is calculated; the time period with the speech rate fluctuation amplitude exceeding the preset threshold is determined as the time period when the user has an obvious speech rate fluctuation during expression, and it is marked as the speech rate fluctuation time period; among them, the speech rate fluctuation amplitude is determined by calculating the ratio of the time interval between adjacent frames to the difference in speech energy.
[0025] When implementing the process of "setting the size of the time window based on the continuous frame sequence and dividing the time window into several time periods", first, a fixed time window size needs to be set according to the frame sequence of the speech signal, usually set in milliseconds. This time window size is usually adjusted according to the rhythm of the language and the requirements of speech processing, and the common window size is from 20ms to 40ms. Then, the speech signal is divided according to the set window size. Specifically, each time window contains several frames (usually cut using the preset frame length and frame shift), and it is ensured that there is an overlap between these time periods to avoid information loss. The overlap size is usually set to 10ms to 20ms, that is, the setting of the frame shift, to ensure that the signal changes can be fully captured within each time period. In this way, the continuous signal can be cut, making the signal within each time period consistent and facilitating subsequent analysis. This method not only helps to analyze the change in speech rate in a finer granularity but also can better handle the rapid changes in the speech signal, especially when the user's speech rate fluctuates greatly, ensuring that no key speech change information is missed. When implemented by software, the frame cutting and time window setting are usually achieved through signal processing functions or toolkits to ensure the efficiency and accuracy of this process.
[0026] In order to achieve "by analyzing the speech rate change state in each time period, specifically including analyzing the time interval between adjacent frames and the change in speech energy, and calculating the speech rate fluctuation amplitude in each time period", it is first necessary to conduct a detailed analysis of the frames within each time period. The time interval refers to the time difference between adjacent frames, which is usually determined by the set frame length and frame shift. The software can measure the speed change of the speech signal in this time period by calculating the time interval between adjacent frames in each time period. Secondly, the change in speech energy is usually represented by calculating the energy within each frame, and the change in energy reflects the intensity fluctuation of the speech signal. When implementing the software, a short-time energy calculation method can be adopted. By squaring and averaging the signal amplitude of each frame, and then analyzing the energy fluctuation according to the change in signal intensity within each time period. When calculating the speech rate fluctuation amplitude in each time period, the time interval between adjacent frames and the speech energy difference can be combined to obtain the fluctuation amplitude in each time period. This can be achieved by calculating the ratio of the time interval between adjacent frames and the energy difference, reflecting the intensity of the speech rate fluctuation in this time period. This calculation process helps to identify obvious cases of speech rate fluctuation, especially when the user speaks quickly or slowly. The calculated fluctuation amplitude can provide an accurate basis for subsequent speech processing. This method can be implemented through existing signal processing tools and software libraries (such as MATLAB, NumPy in Python, etc.). Using techniques such as the Fast Fourier Transform (FFT), it can efficiently process and calculate the time interval and speech energy change in each time period.
[0027] To achieve "determine the time period with a speech rate fluctuation amplitude exceeding the preset threshold as the time period when the user has an obvious speech rate fluctuation during expression, and label it as the speech rate fluctuation time period", it is first necessary to set a reasonable speech rate fluctuation amplitude threshold. The setting of this threshold is usually based on experience in actual applications or adjusted through experimental data to ensure that normal speech rate fluctuations can be distinguished from obvious speech rate fluctuations. During the implementation process, the software system will traverse the speech rate fluctuation amplitude of each time period and compare it with the preset threshold. If the speech rate fluctuation amplitude of a certain time period is greater than the threshold, it is considered that there is an obvious speech rate fluctuation in this time period. The software can use conditional judgment statements (such as if statements) to compare the fluctuation amplitude of each time period with the threshold one by one. If the condition is met, the time period is labeled as the time period with an obvious speech rate fluctuation. During the labeling process, the software will record the start and end times of this time period to ensure that the time period with obvious fluctuations can be accurately identified. To further improve the accuracy, the system can also filter out short-term and accidental fluctuations by combining the duration. For example, if the speech rate fluctuation amplitude exceeds the threshold but the duration is too short (e.g., less than 50 ms), this time period can be ignored and not considered as a time period with obvious fluctuations. In this way, the system can effectively distinguish the obvious speech rate fluctuations caused by reasons such as thinking or rapid topic switching, and provide reliable input data for subsequent speech recognition, translation, or other processing. Such a processing method can be implemented in the software through signal processing algorithms to ensure that the system efficiently identifies and labels the speech rate fluctuation time periods, avoids misjudgment, and improves the accuracy and effect of speech processing.
[0028] "An obvious speech rate fluctuation means that the speech rate fluctuation amplitude exceeds the preset threshold" means that when analyzing the speech signal, if the speech rate change amplitude of a certain time period exceeds the preset threshold, it is considered that there is an obvious speech rate fluctuation in this section of speech. This threshold is usually set through experiments or experience to ensure that the system can effectively distinguish the daily normal speech rate fluctuations from the drastic speech rate changes, so as to avoid making wrong responses to some minor and irrelevant fluctuations. As for "the speech rate fluctuation amplitude is determined by calculating the ratio of the time interval between adjacent frames and the speech energy difference", the core of this calculation method is to measure the speech rate fluctuation by observing the changes in the speech signal in adjacent time periods. Specifically, the time interval reflects the time difference between adjacent frames. Usually, the time interval is shorter when the speech rate is faster and longer when the speech rate is slower; while the speech energy difference reflects the change in signal intensity. Usually, when the speech suddenly accelerates or slows down, the change in energy will be more obvious. Therefore, by calculating the ratio of the time interval between adjacent frames and the energy difference, a value reflecting the speech rate fluctuation intensity can be obtained. When this value exceeds the preset threshold, it can be considered that within this time period, the speech rate fluctuation has reached an obvious level. This method can effectively capture the important changes in the speech signal and provide an accurate reference basis for the speech processing system.
[0029] A rhythm evaluation module extracts rhythm feature information from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, analyzes it, and evaluates the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period. In this embodiment, in the rhythm evaluation module, rhythm feature information is extracted from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period and preprocessed after being obtained. Acoustic fluctuation feature information and spectral configuration distribution information are extracted from the preprocessed rhythm feature information and analyzed after being extracted to generate a speech energy fluctuation coefficient and a speech spectrum dispersion index respectively. A rhythm mutation intensity evaluation model is constructed for the generated speech energy fluctuation coefficient and speech spectrum dispersion index, and a rhythm mutation index is generated through weighted summation. A preset rhythm mutation index threshold interval is determined and compared with the generated rhythm mutation index after being determined, and the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period is evaluated according to the comparison result.
[0030] In the process of "extracting rhythm feature information from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period", the speech data in this period can be processed frame by frame through a speech signal analysis software system, converting the continuous speech waveform into a series of frame-structured speech samples, and extracting basic parameters that can reflect the speech rhythm change characteristics in each frame. Specifically, based on the speech signal after frame division, the short-time Fourier transform (STFT) and Mel filter bank can be used to obtain the spectral structure of each frame, and then feature values such as short-time energy, zero-crossing rate, spectral energy distribution, and Mel frequency cepstral coefficients (MFCC) are further extracted and arranged in time series to construct complete rhythm feature information. These features are all directly quantifiable numerical data, and software automation extraction can be realized through an audio signal processing library (such as librosa or MATLAB speech toolset) to ensure that the physical properties of the rhythm change can be captured and described by the system in the digital domain.
[0031] Preprocessing the rhythm feature information extracted above is to improve the accuracy and stability of subsequent parameter calculations and prevent data deviation caused by outliers, noise interference, or dimensional inconsistencies. The preprocessing operations generally include three aspects: First, normalize all the extracted frame-level feature data (such as z-score standardization or min-max scaling) to eliminate the influence between different feature dimensions and enable various data to be compared on the same scale; Second, smooth the short-term abnormal data points that appear in consecutive frames. The moving average method or median filtering can be used to reduce the misleading effect of mutant frames on the judgment of overall fluctuations; Third, use a feature selection algorithm (such as principal component analysis PCA or feature ranking) to compress redundant or weakly correlated dimensions and retain the core features that are most representative of rhythm assessment. These preprocessing steps can be executed chainedly through software functions after the audio feature matrix is completed, ensuring that the generated rhythm feature information has high comparability, high robustness, and a good mathematical structure, facilitating subsequent coefficient calculations and the accurate operation of the evaluation model.
[0032] In the process of implementing "extracting acoustic fluctuation feature information and spectral configuration distribution information from the preprocessed rhythm feature information", it is first necessary to perform targeted splitting and structuring on the frame-level speech feature matrix that has been normalized and noise-reduced. For the acoustic fluctuation feature information, two key parameters reflecting the time-domain fluctuation characteristics of the speech signal, namely the short-time energy value and the zero-crossing rate value, can be selected to construct a numerical pair reflecting the strength change and amplitude fluctuation of the speech between different frames; Specifically, the system can extract the energy value and zero-crossing rate value of each frame according to the frame index and form a two-dimensional feature sequence in chronological order. For the spectral configuration distribution information, based on the extracted frame-level Mel-frequency cepstral coefficients (MFCCs), two numerical features representing its frequency-domain fluctuation characteristics can be further extracted: First, calculate the mean absolute deviation between each frame of MFCC and the average value of the corresponding dimension as the cepstral offset intensity; Second, calculate the variance of the overall distribution of MFCCs to reflect the discrete trend of frequency components. Through the above methods, the rhythm feature information is further separated into two feature information sets with clear structures, definite data dimensions, and independent evaluation values, and both are specific numerical data obtained through mathematical operation logic by the software system at the audio feature matrix level, meeting the input requirements of the subsequent evaluation model.
[0033] In the process of implementing "determining the pre-set rhythm mutation index threshold interval", statistical modeling and interval division of historical speech samples are usually carried out in a data-driven manner to establish a standard discrimination benchmark suitable for evaluation. Specifically, first, a large number of speech sample data containing different degrees of rhythm changes need to be collected during the training phase, and the rhythm feature information consistent with the current process is extracted from each sample. Then, the corresponding speech energy fluctuation coefficient and speech spectrum dispersion index are generated respectively, and the rhythm mutation index is synthesized according to the established weight. Subsequently, these index values can be clustered or sorted according to the actual rhythm state labels of the samples (such as "stable", "slight mutation", "severe mutation"), and statistical analysis is carried out on their distribution intervals. For example, the boundary values of mutation intensity are divided by the percentile method or the Gaussian fitting method. Finally, a set of static threshold intervals can be selected within the statistical range with a high sample coverage rate. For example, 0.0 - 0.4 corresponds to stable, 0.4 - 0.7 is moderate mutation, 0.7 - 1.0 is severe mutation, etc. This interval can be solidified in the system in the form of parameter configuration and can be adaptively adjusted according to the subsequent model operation effect. The whole process can be completed through the data annotation module, clustering algorithm module and parameter management module in the software, without manually setting subjective values, ensuring that the division standard has statistical significance and system stability.
[0034] In this embodiment, the acquisition logic of the speech energy fluctuation coefficient is as follows: Extract the acoustic fluctuation feature information from the pre-processed rhythm feature information, specifically including the energy value and zero-crossing rate of each frame in the user speech signal collected by the Bluetooth headset during the speech rate fluctuation time period, and label them respectively as and , represents the energy value of the th frame in the user speech signal collected by the Bluetooth headset during the speech rate fluctuation time period, represents the zero-crossing rate of the th frame in the user speech signal collected by the Bluetooth headset during the speech rate fluctuation time period, , is a positive integer; In the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, the energy value and zero-crossing rate of each frame can be extracted through audio signal processing algorithms in software, and they belong to the two most commonly used low-order acoustic features. First, the energy value is the short-time energy intensity of each frame of the voice signal, reflecting the overall amplitude of the voice waveform within that frame. The acquisition method is as follows: divide the preprocessed voice signal into a series of equally long frames, perform a square operation on the sampling points in each frame and then sum or average them to obtain the total energy of that frame; this process can be achieved through a sliding window method to ensure frame-to-frame continuity. The larger the energy value, the louder or more impactful the voice signal of that frame, which often appears in stressed, plosive, or rapid onset sounds. Second, the zero-crossing rate (ZCR) is the number of times the waveform of the voice signal crosses the zero axis from positive to negative or negative to positive within a frame period, reflecting the degree of waveform oscillation or irregularity. The specific acquisition method is: compare the consecutive sampling points of each frame signal pairwise, record a zero-crossing event when a sign change is found, and then divide the total number by the number of samples within the frame to obtain the normalized ZCR value. When the ZCR is relatively high, it usually corresponds to speech segments with many high-frequency details such as silent segments, fricatives, and nasalized sounds. When the ZCR is relatively low, it mostly appears in vowels, speech pauses, or low-frequency smooth sentences. Both types of data can be automatically extracted in a software environment using standard audio analysis libraries (such as Librosa in Python, Audio Toolbox in MATLAB, etc.), without manual intervention, and have high repeatability and numerical stability, which are the basic input features for constructing a speech rhythm analysis model.
[0035] Calculate the speech energy fluctuation coefficient, and the specific calculation formula is as follows: In the formula, is the speech energy fluctuation coefficient.
[0036] The design of the calculation formula for this speech energy fluctuation coefficient aims to comprehensively measure the collaborative fluctuation characteristics of the voice signal in terms of amplitude intensity and waveform oscillation frequency during the speech rate fluctuation period, so as to quantify the severity of speech rhythm changes. First, squaring the energy value of each frame is to enhance the weight of high-energy frames, so that frames with a sudden increase in speech intensity during the speech rate fluctuation have a more significant contribution to the overall fluctuation perception. Second, the multiplied by it is a regularized representation of the zero-crossing rate of that frame, It is an index that measures the number of positive and negative alternations of a speech signal within a frame, representing the oscillation frequency of the waveform. The larger its value, the more frequent the speech fluctuations. Adding 1 is used to prevent the occurrence of zero in multiplication. The structure of multiplying the two reflects a key concept: having only strong energy or only high oscillation is not sufficient to indicate a drastic rhythm change. Only when the energy and waveform oscillation increase simultaneously within the same frame does it indicate a significant rhythm fluctuation. Therefore, taking the weighted average of the "energy × oscillation" values for each frame (that is, summing over the frames and then dividing by ), the overall co-fluctuation intensity of the speech rate fluctuation time period is obtained. Finally, it is compressed through the natural logarithm function to smooth extreme values, enhance the resolution ability for medium and low amplitude fluctuations, and at the same time ensure that this coefficient has good numerical stability and comparability. The overall formula not only has a compact mathematical structure, but also starting from the physical characteristics of speech, can effectively capture the comprehensive performance of rhythm fluctuation changes.
[0037] There is a positive correlation between the numerical value of the speech energy fluctuation coefficient and "evaluating the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation time period", that is, the larger this coefficient, the more drastic the change in the energy intensity of the speech signal during this time period, and the more frequent the accompanying waveform oscillation, thus reflecting that there is a greater possibility of rhythm mutation in the speech. Specifically, this coefficient can comprehensively capture the drastic fluctuations of speech in both the amplitude and structure dimensions by multiplying the square of the energy of each frame by the zero-crossing rate of the corresponding frame and taking the weighted average. When the user suddenly slows down or speeds up the speech rate during the expression, or there are obvious discontinuities and breaks in the speech rhythm during the language organization process, there are often syllables with drastic changes in strength and frequent acoustic wave jumps, and at this time the energy fluctuation coefficient will increase significantly. Therefore, the higher the numerical value of this coefficient, the greater the change in the rhythm continuity of the speech, so it can be used to quantify the intensity of the rhythm mutation in the speech and provide a key basis for subsequent translation optimization or recognition strategy adjustment.
[0038] In this embodiment, the acquisition logic of the speech spectrum discrete index is as follows: Extract the spectrum configuration distribution information from the preprocessed rhythm feature information, specifically including: obtaining the spectrum feature information of each frame in the user speech signal collected by the Bluetooth headset during the speech rate fluctuation time period, and constructing a Mel-frequency cepstral coefficient vector based on the spectrum feature information of each frame, expressed as: In the formula, is the Mel-frequency cepstral coefficient vector of the th frame, containing components, represents the Mel-frequency cepstral coefficient corresponding to the th cepstral component in the th frame, , , and are all positive integers; During the speech rate fluctuation period, in order to obtain the spectral feature information of each frame from the user speech signal collected by the Bluetooth headset and construct the Mel Frequency Cepstral Coefficient vector (MFCC), the standard speech signal analysis process can be implemented through software. Specifically, first, the collected speech signal is preprocessed, including pre-emphasis, framing, and windowing, to make the original time-domain signal have good stationarity in the time dimension. Subsequently, the fast Fourier transform (FFT) is performed on each frame of the signal to obtain the spectral energy distribution, thereby forming the power spectrum of the frame. Then, the Mel filter bank is applied to filter the power spectrum. The design of the filter follows the non-linear characteristics of the human auditory system's frequency perception, mapping the frequency domain to the Mel scale, so that the low-frequency part is subdivided and the high-frequency part is compressed to more effectively capture the perceptual characteristics of speech. The output of each Mel filter after filtering represents the energy value of a frequency band. Then, logarithmic transformation is performed on these frequency band energy values, so that the relative differences of high energy values are compressed while the changes of low energy values are retained, enhancing the sensitivity to weak speech signals. Finally, the discrete cosine transform (DCT) is performed on the Mel energy sequence after logarithmic transformation to remove the correlation in the frequency domain and compress the signal energy into the first several coefficients, obtaining a set of compact, low-dimensional, and well-discriminative Mel Frequency Cepstral Coefficient vectors (MFCC). Each component of this vector is a definite real value, reflecting the cepstral characteristics of the speech frame in different frequency bands. Usually, the first g cepstral components are selected for analysis. The entire processing process can be completed in a software manner in a local or embedded system through a standard speech processing library (such as Librosa, Python SpeechFeatures, MATLAB Audio Toolbox, etc.), with high efficiency, versatility, and scalability. The constructed MFCC vector provides a clear and structured numerical basis for subsequent extraction of rhythm-related parameters (such as cepstral shift, cepstral dispersion).
[0039] Based on the Mel Frequency Cepstral Coefficient vector of each frame, calculate the following two types of characteristic values: First, the average offset value of the cepstral feature, according to the formula: In the formula, is the average value of the Mel Frequency Cepstral Coefficient corresponding to the th cepstral component in all frames, used to reflect the global central tendency of this frequency domain feature dimension, represents the average offset value of the cepstral feature of the th frame; Cepstral feature average offset value is an important parameter used to measure the degree to which the spectral structure of the th frame of speech signal deviates from the overall average state. Its essence is the weighted average of the difference between the mel-frequency cepstral coefficient vector of this frame and the mean of the cepstral components in all frames. In the actual construction process, first, for each dimension of the cepstral component and its corresponding average value of this dimension in all frames during the entire speech rate fluctuation time period perform an absolute difference operation to reflect the degree of deviation of the spectral form of this frame in this frequency dimension relative to the global central tendency; then average the differences in all frequency band dimensions to obtain a comprehensive offset index . The larger this value is, the more inconsistent the overall spectral form of this frame is compared with the "central spectral form" of other frames, indicating that there may be phenomena such as frequency reconstruction, pronunciation mode change, or vocalization mode mutation in speech expression. The specific manifestations may be sudden consonant bursts, rapid vocal cord closures, frequency jumps, or pronunciation mouth shape conversions, etc. Therefore, high often appears at the jumps in sentence rhythm, sudden changes in speech rate, or the turning points of syntactic structures. If the value is small, it indicates that the spectral form of this frame tends to be stable and is highly consistent with the average expression spectrum during the speech rate fluctuation time period, representing a continuous and stable state of speech expression. Therefore, this value not only reflects the gap between the local and overall spectral forms but also provides an important spectral offset basis for subsequent judgment of the rhythm mutation trend.
[0040] Second, the cepstral distribution dispersion value, according to the formula: In the formula, is the mean of the mel-frequency cepstral coefficients corresponding to all cepstral components in the th frame, representing the local central tendency of the spectral feature distribution within the frame, is the cepstral distribution dispersion value of the th frame; Cepstral distribution dispersion value is a quantitative index used to describe the fluctuation amplitude of the internal structure of the mel-frequency cepstral coefficients of the th frame. Its essence is the degree of dispersion of each component of the cepstral coefficient vector of this frame around the mean within the frame. Specifically, first calculate the mean of all cepstral components of this frame , and then average the squared deviations of each component from the mean, that is, obtain the cepstral distribution dispersion of this frame. This process is equivalent to calculating the statistical variance of all spectral features of this frame in the frequency domain dimension. Therefore The larger the value, the more significant the difference in the spectral shape of the frame among different frequency bands, the more uneven the spectral energy distribution, and it reflects phenomena such as possible pronunciation mixing, severe harmonic fluctuations, or sudden formant changes in the speech signal. Such a situation is often accompanied by a sudden change in the speech pronunciation manner, such as transitioning from a vowel to an affricate or a compound sound, or structural disturbances such as short pauses, elision, or oral dynamic changes during pronunciation. On the contrary, if the value is small, it indicates that the cepstrum structure of this frame is relatively smooth, and the spectral changes in each frequency band are relatively consistent, meaning that the speech expression of this frame has good continuity and stability in the spectral structure. Generally speaking, is the key basis for identifying whether there is internal spectral diffusion or perturbation in the speech rhythm, and the level of its value can be used to assist in judging whether there is a sudden change trend in the sentence rhythm.
[0041] Calculate the speech spectral dispersion index, and the specific calculation formula is as follows: In the formula, is the speech spectral dispersion index.
[0042] The calculation formula of using this speech spectral dispersion index is to comprehensively measure the instability and overall offset trend of each frame of speech in the spectral structure during the speech rate fluctuation period, so as to sensitively identify the possibility of sudden changes in the sentence rhythm. Among them, represents the discrete degree value of the cepstrum distribution of the th frame, which is used to describe the energy difference degree between each frequency band inside this frame. The larger its value, the more unbalanced the spectral shape of this frame, reflecting local pronunciation perturbation or harmonic structure change; represents the average offset value of the cepstrum feature, which is used to represent the deviation degree between the overall spectral structure of this frame and the average spectral shape during the speech rate fluctuation period. The larger its value, the more obvious the change in the speech state, such as a change in the pronunciation manner or a restructuring of the vocal tract structure. Adding the two can take into account both the "internal spectral fluctuation" and "overall structure offset" dimensions at the same time, forming a composite expression that is more sensitive to the perception of sudden changes in the speech rhythm. Furthermore, through an exponential function, the total fluctuation intensity is non-linearly amplified, so that the frames with significant peaks in contribute more in the overall , thus enhancing the system's ability to identify local rhythm anomalies. Finally, taking the average of the calculation results of all frames not only retains the sensitivity of frame-level mutation recognition but also realizes the stable quantification of the rhythm state of the entire speech expression. This calculation method has advantages such as clear theoretical explanation, stable engineering implementation, and natural numerical trend.
[0043] The speech spectral dispersion index The value of is positively correlated with the intensity of the rhythm mutation in the user's expression during the time period of speech speed fluctuation, that is, The larger the value, the higher the possibility and intensity of the sudden change in the rhythm of the sentence. The index is obtained by adding the discrete value of the cepstral distribution of each frame and the average offset value of the cepstral feature and then amplifying it exponentially. It can simultaneously reflect the changes in the local spectrum structure fluctuation and the overall spectrum contour offset of the speech signal. When it is large, it indicates that the spectrum energy in the frame is unevenly distributed among the frequency bands, and there may be harmonic disturbance, unstable pronunciation or structural diffusion. When it is large, it means that the frame spectrum deviates significantly from the average spectrum of the entire speech segment, which may be accompanied by a change in pronunciation or a sudden change in speech rhythm. The two together constitute the dual criterion for rhythm mutation. By summing them and applying exponential operations, the influence of the mutation frame in the entire speech can be further amplified. The larger the value, the more drastic and unstable the frame-level spectrum structure changes during the speech rate fluctuation period, and the more likely the overall speech rhythm will experience a sharp turn, jump, or discontinuous expression, thus indicating a stronger intensity of sentence rhythm mutation. A smaller value means that the speech expression is continuous, stable, and has a smooth rhythm, with basically no obvious rhythm changes.
[0044] In this embodiment, the generated speech energy fluctuation coefficient and speech spectrum dispersion index A rhythm mutation intensity evaluation model is constructed, and the rhythm mutation index is generated by weighted summation. The specific calculation formula is as follows: In the formula, is the rhythm mutation index, and Speech energy fluctuation coefficient and speech spectrum dispersion index The non-zero weight coefficient of .
[0045] In the process of rhythm mutation intensity assessment, in order to achieve a comprehensive judgment of the rhythm mutation trend in the user's voice expression, the speech energy fluctuation coefficient (EFC) and speech spectrum discrete index (SDI) calculated in the early stage can be weighted and fused by software to construct the rhythm mutation index (RMI). In the specific implementation, the two indicators are first normalized to the same numerical scale (such as linear mapping to the [0,1] interval) to eliminate the scale difference caused by different calculation formulas; then two non-zero weight coefficients are introduced and , multiply them by the normalized EFC and SDI respectively, and then sum the two weighted items to generate the final RMI value. The settings of the two weight coefficients need to satisfy , and and are both not zero to ensure that both indicators make substantial contributions in the evaluation. The values of the weight coefficients can be set according to the training corpus or expert experience. For example, when the spectral shape change is more emphasized, the value of can be appropriately increased, while if more attention is paid to the energy dynamic fluctuation, the proportion of can be increased. This weighted fusion method not only retains the discriminative ability of each of the two indicators but also can adapt to the dynamic adjustment requirements under different speech styles and scenarios, realizing a quantitative, adjustable, and sensitive expression of the intensity of the rhythm mutation in the sentence during the speech rate fluctuation period.
[0046] In this embodiment, a preset rhythm mutation index threshold interval is determined, and after determination, it is compared with the generated rhythm mutation index to evaluate the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period according to the comparison result. The specific comparison and analysis are as follows: If , the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period is low intensity; This situation indicates that during the current speech rate fluctuation period, the fluctuation degree of the user's speech signal in the spectral structure is small, without significant local band energy dispersion and no obvious overall spectral contour shift. This shows that although there are slight speech rate fluctuations in the user's speech expression process, the overall rhythm of the sentence remains coherent, stable, and smooth, without the risk of interfering with speech segmentation, recognition, or translation. Therefore, this type of expression usually does not require special speech processing strategies and can continue to be recognized and translated according to the conventional speech rate and mute threshold processing mechanism, avoiding unnecessary consumption of computing resources and response delays, thereby improving the overall system operation efficiency and response speed.
[0047] If , the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period is medium intensity; This situation indicates that during the current speech rate fluctuation period, the user's speech signal has shown a certain degree of instability in the rhythm structure. Specifically, it is manifested as an increase in the degree of frequency band energy dispersion in some frames, or the overall spectral profile begins to deviate from the mainstream average spectrum, but has not reached a serious perturbation state. Such speech states usually occur at sentence connection points, in the middle of complex sentence patterns, or at semantic turning points, and have the potential risk of rhythm jumps. If processed conventionally, it may lead to inaccurate speech frame division, resulting in recognition delay or semantic breakage. Therefore, this type of moderate-intensity rhythm mutation needs to trigger mild strategy intervention, such as extending the speech recognition endpoint judgment time, adjusting the frame window size, or appropriately delaying the mute determination threshold, to enhance the recognition module's tolerance for rhythm uncertainty and improve translation continuity and accuracy.
[0048] If , during the speech rate fluctuation period, the intensity of the sentence rhythm mutation in the user's expression is high intensity.
[0049] This situation indicates that the speech expression state of the user during the speech rate fluctuation period has undergone a significant mutation. Its frame-level spectral structure shows strong discontinuity and structural jumps, which may include complex phenomena such as sudden pronunciation mode switching, rapid speech segment splicing, sudden change in tone, short-term repetition, syllable swallowing, etc. In this case, the frequency band energy fluctuates violently, and the average spectral profile deviates greatly, which will seriously interfere with the conventional recognition path of the speech processing module. If still processed according to the standard algorithm, it is very easy to cause speech segmentation misalignment, keyword truncation, or misactivation of the recognition model, and ultimately lead to word order confusion, semantic misinterpretation, and sentence meaning interruption in the translated content. Therefore, this level must trigger advanced dynamic intervention strategies, such as adjusting the frame overlap ratio, enabling the context semantic preservation mechanism, delaying the recognition decision, etc., to stabilize the recognition process from multiple dimensions and ensure the system still has robust translation accuracy and usability in high-dynamic speech scenarios.
[0050] The speech regulation module selects and executes corresponding speech processing strategies according to the evaluation results and performs dynamic regulation during the execution process; In this embodiment, in the speech regulation module, according to the evaluation results, the corresponding speech processing strategies are selected and executed, specifically: If the evaluation result is low intensity, the speech processing strategies selected and executed include: dividing speech frames using the default frame cutting window parameters, maintaining the standard mute duration threshold and the fixed endpoint detection method, and delimiting speech segments in a conventional manner; In the case where the evaluation result is of low intensity, the system can perform voice frame division and speech segment definition processing using default parameters in software. The specific method is as follows: First, in the preprocessed voice signal, the voice stream is divided frame by frame according to the established frame length (such as 25 ms) and frame shift (such as 10 ms), without adjusting the window overlap ratio and frame duration, to maintain an even coverage of the voice signal in the time dimension. Then, the standard energy threshold and zero-crossing rate determination rules in the silence detection module are applied to identify the silent frames and voiced frames in the voice frames, and it is judged whether the speech segment terminates by whether the duration of consecutive silent frames exceeds the default silence duration threshold (for example, 200 ms). In the endpoint detection stage, a fixed start-stop recognition logic is executed, that is, the speech segment boundaries are calibrated through the starting point of the voice energy and the ending point of the silence, without performing delay compensation or context extension processing. This method can efficiently process speech paragraphs with stable rhythms at low computational costs, avoid unnecessary investment in computing resources, and at the same time reduce the processing delay. Since under the condition of low-intensity rhythm mutations, the influence of speech rate fluctuations on the accuracy of speech segment cutting is small, the integrity and clarity of the input to the speech recognition module can be ensured without dynamic parameter intervention, thus maintaining the overall operating efficiency of the system and the stability of speech translation.
[0051] If the evaluation result is of medium intensity, the selected and executed speech processing strategies include: adjusting the size of the voice frame cutting window to increase the window duration, increasing the frame overlap ratio, and extending the silence duration threshold to mitigate the risk of misjudgment in segmentation caused by rhythm fluctuations. When the evaluation result is of medium intensity, the system can dynamically adjust the speech processing parameters in software to enhance the adaptability to mild rhythm fluctuations. The specific implementation methods include: First, increase the duration of the voice frame cutting window (such as from the default 25 ms to 35 ms), and at the same time increase the frame overlap ratio (such as from 40% to 60%). By increasing the time coverage range and enhancing the inter-frame correlation, the spectral features have a stronger continuous expression on the time axis, thus slowing down the influence of the speech rate change on the determination of speech segment boundaries. Second, in the silence detection module, extend the silence duration threshold from the default value (such as 200 ms) to a higher value (such as 300 ms), so that when the system detects consecutive silent frames, it will not prematurely misjudge as the termination of the speech segment, reducing the speech segment segmentation errors caused by a slowdown in speech rate or short pauses. The entire process is automatically completed by the speech control module calling the parameter configuration interface before strategy execution, without hardware intervention, and can be adjusted in real time in response to the evaluation result. The core of this strategy is to enhance the fault tolerance of the speech processing system to mild rhythm changes by expanding the time coverage granularity of the frames and the silence buffer time, avoiding incorrect splitting or truncation of sentences, and ensuring the integrity and coherence of the translation semantics.
[0052] If the evaluation result is high intensity, the selected and executed speech processing strategies include: on the basis of adjusting the frame cutting window and the silence threshold, enabling the cross-segment semantic context retention mechanism and enabling the delayed endpoint judgment mode to avoid segment cutting errors and recognition timing errors at the rhythm jump points; In the case of a high-intensity evaluation result, to cope with the drastic rhythm mutations in speech and prevent problems such as incorrect segment cutting and recognition misalignment, the system can implement multi-level advanced speech processing strategies through software. First, on the basis of the adjusted frame cutting window duration (such as extended to 40 ms) and the silence duration threshold (such as increased to 400 ms), further enable the cross-segment semantic context retention mechanism, that is, during the speech recognition stage, continuously retain a certain number of recognized text buffer frames or language segments before and after the current segment (for example, the first 2 seconds and the last 1 second). By constructing a context correlation window, the context of the previous and subsequent sentences is used as one of the input features to participate in semantic determination during translation and recognition, thereby reducing recognition mismatches caused by the sudden intrusion of the sentence start or the interruption of the sentence end; at the same time, enable the delayed endpoint judgment mode, that is, when a possible segment end signal (such as silence timeout) is detected, the speech termination is not immediately confirmed, but a delay confirmation interval (such as 100 ms - 150 ms) is set. During this time period, the speech input changes are continuously monitored. If audio activity is detected again, the termination determination is cancelled, thereby avoiding short pauses caused by rhythm mutations being misrecognized as the end of the sentence. The above mechanism significantly enhances the adaptability of the speech processing system to strong rhythm mutation scenarios without affecting the real-time performance of the system through methods such as dynamic caching, predictive control, and endpoint trigger backshift, ensuring the integrity, context coherence, and semantic accuracy of the translation output.
[0053] And perform dynamic regulation during the execution process, specifically: during the execution of the speech processing strategy, continuously monitor the change trend of the frame-level spectral energy of the speech signal and the change state of the silence duration between speech segments. If it is detected during the execution of the speech processing strategy that the rhythm mutation index continuously rises in consecutive frames and exceeds the maximum threshold corresponding to the current evaluation level, immediately switch to a higher-level speech processing strategy, update the speech frame cutting window parameters and the silence duration threshold, and enable the cross-segment semantic retention function; if it is detected during the processing that the rhythm mutation index continuously drops in consecutive frames and is lower than the minimum threshold corresponding to the current strategy level, downgrade the current speech processing strategy, restore the default frame division parameters, and turn off the extended processing mechanism to achieve real-time dynamic adjustment of the processing strategy.
[0054] During the execution of the speech processing strategy, dynamic regulation can be achieved by constructing a real-time monitoring and strategy level linkage mechanism for the rhythm mutation index (RMI) through software. The specific implementation methods include: within each processing cycle, at a fixed time interval (such as every 100 milliseconds), incrementally update the rhythm features of newly acquired frames during the speech rate fluctuation period, calculate the corresponding speech energy fluctuation coefficient (EFC) and speech spectrum dispersion index (SDI) in real time, and generate the latest rhythm mutation index RMI; the system forms a trend sequence with the RMI values in multiple consecutive time periods (such as 5 consecutive frames), and judges whether there is a significant increase or decrease based on the change trend of this sequence. If it is detected that the RMI continuously rises and its average value exceeds the maximum threshold value corresponding to the current processing strategy, it indicates that the rhythm mutation is intensifying. At this time, the system immediately switches to a higher-level processing strategy through the parameter interface, automatically modifies the current frame cutting window length, frame shift parameter and silence determination threshold, and activates the cross-segment semantic context preservation function to enhance the segmental robustness and semantic continuity of the system under the condition of severe rhythm disturbance; on the contrary, if it is detected that the RMI continuously decreases and its average value is lower than the minimum threshold value corresponding to the current strategy level, the system judges that the current speech has returned to a stable state, automatically downgrades the processing strategy to a lower intensity level, restores the default parameter settings and closes the additional processing module, so as to avoid resource waste and recognition overfitting. The entire regulation process relies on continuous indicator streams and judgment logics to achieve real-time strategy adjustment without interrupting the recognition process, effectively adapting to the rapid changes in the speech expression state, and ensuring the stability and accuracy of the speech translation system in dynamic scenarios.
[0055] The text verification module, after completing speech recognition, verifies and adjusts the language structure and word order of the recognized text to ensure that the translated text conforms to the grammar rules of the target language and maintains semantic coherence.
[0056] After completing speech recognition, the text verification module can automatically verify and make necessary adjustments to the language structure and word order in the recognition result through software, so as to improve the accuracy and coherence of the translated text at the grammar and semantic levels. Specifically, first parse the syntactic structure of the text obtained by speech recognition, and methods such as dependency syntactic analysis or phrase structure analysis based on natural language processing (NLP) can be used to identify language components such as the subject-predicate-object structure, modification relationship, and sentence pattern boundary in the text. At the same time, with the help of part-of-speech tagging, entity recognition and phrase chunking technologies, clarify the grammatical roles played by each word in the sentence, and establish a language rule library based on the grammar norms of the target language or rely on a trained language model, such as a tree-based reordering or a sequence learning model (such as BiLSTM-CRF), to perform pattern recognition and reconstruction correction on problems such as subject-predicate inversion, object omission, and modifier dislocation in the sentence.
[0057] Furthermore, to ensure the rationality of the translation output word order and semantic coherence, the system can also judge and optimize the connection logic between multiple sentence outputs based on the context consistency evaluation and semantic association scoring mechanism. The implementation methods include: modeling the language flow coherence through dimensions such as the lexical cohesion between upper and lower sentences, the consistency of time adverbs, and the subject persistence, identifying and eliminating semantic jumps, abrupt tones, or logical breaks between sentences caused by speech recognition errors. In addition, by combining the target language language model to score the candidate word order schemes, the highest-confidence structure is selected as the final corrected text output. The above process is fully implemented at the software level without manual intervention, which can effectively enhance the language expression quality of the final translation result on the premise of ensuring the automated operation of the system, enabling users to obtain a translation text output that conforms to language habits, has a reasonable structure, and is semantically complete.
[0058] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.
[0059] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0060] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0061] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0062] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the above-described embodiments are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0063] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0064] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0065] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A voice translation system based on Bluetooth headset, characterized in that: It includes a voice acquisition module, a speech speed detection module, a rhythm evaluation module, a voice control module and a text verification module; The voice acquisition module collects user voice signals through Bluetooth headsets, pre-processes the collected voice signals, and divides the voice signals into continuous frame sequences according to preset frame lengths and frame shifts after pre-processing; The speech rate detection module establishes a time window based on a continuous frame sequence, analyzes the speech rate change state within the time window, determines the time period when the user's speech rate fluctuates significantly during expression, and marks the time period as a speech rate fluctuation time period, where the speech rate fluctuation significantly means that the amplitude of the speech rate fluctuation exceeds a preset threshold; The rhythm evaluation module extracts rhythm feature information from the user's voice signal collected by the Bluetooth headset during the speech speed fluctuation period, analyzes it, and evaluates the intensity of the sudden change in the rhythm of the user's expression during the speech speed fluctuation period; The speech control module selects and executes the corresponding speech processing strategy according to the evaluation results, and performs dynamic control during the execution process; The text verification module verifies and adjusts the language structure and word order of the recognized text after completing speech recognition to ensure that the translated text conforms to the grammatical rules of the target language and maintains semantic coherence.
2. A voice translation system based on Bluetooth headset according to claim 1, characterized in that: In the speech rate detection module, the size of the time window is set based on the continuous frame sequence, and the time window is divided into several time periods; by analyzing the speech rate change state of each time period, specifically including analyzing the time interval and speech energy change between adjacent frames, the speech rate fluctuation amplitude of each time period is calculated; the time period in which the speech rate fluctuation amplitude exceeds the preset threshold is determined as the time period in which the user has obvious speech rate fluctuation in expression, and it is marked as the speech rate fluctuation time period; the speech rate fluctuation amplitude is determined by calculating the ratio of the time interval between adjacent frames to the speech energy difference.
3. A voice translation system based on Bluetooth headset according to claim 2, characterized in that: In the rhythm evaluation module, the rhythm feature information in the user voice signal collected by the Bluetooth headset during the speech speed fluctuation period is extracted and preprocessed after acquisition; the acoustic fluctuation feature information and the spectrum configuration distribution information are extracted from the preprocessed rhythm feature information, and analyzed after extraction to generate a speech energy fluctuation coefficient and a speech spectrum discrete index respectively; a rhythm mutation intensity evaluation model is constructed for the generated speech energy fluctuation coefficient and speech spectrum discrete index, and a rhythm mutation index is generated by weighted summation; a pre-set rhythm mutation index threshold range is determined, and after determination, it is compared with the generated rhythm mutation index, and the intensity of the sentence rhythm mutation in the user's expression during the speech speed fluctuation period is evaluated according to the comparison result.
4. A voice translation system based on Bluetooth headset according to claim 3, characterized in that: The logic for obtaining the speech energy fluctuation coefficient is as follows: The acoustic fluctuation feature information is extracted from the preprocessed rhythm feature information, specifically including the energy value and zero-crossing rate of each frame in the user voice signal collected by the Bluetooth headset during the speech rate fluctuation period, and calibrated as and , Indicates the first The energy value of the frame, Indicates the first The zero-crossing rate of the frame, , is a positive integer; Calculate the speech energy fluctuation coefficient. The specific calculation formula is as follows: In the formula, is the speech energy fluctuation coefficient.
5. A voice translation system based on Bluetooth headset according to claim 4, characterized in that: The acquisition logic of the speech spectrum dispersion index is as follows: Extracting spectrum configuration distribution information from the preprocessed rhythm feature information specifically includes: obtaining spectrum feature information of each frame in the user voice signal collected by the Bluetooth headset during the speech rate fluctuation period, and constructing a Mel frequency cepstrum coefficient vector based on the spectrum feature information of each frame, expressed as: In the formula, For the Mel-frequency cepstral coefficient vector of the frame, containing Quantity, Indicates Frame The Mel-frequency cepstral coefficients corresponding to the cepstral components are , , and All are positive integers; Based on the Mel frequency cepstral coefficient vector of each frame, the following two types of feature values are calculated: First, the average offset value of the cepstrum feature is based on the formula: In the formula, For the The average value of the Mel-frequency cepstral coefficients corresponding to the cepstral components in all frames, Indicates The average offset value of the cepstrum feature of the frame; Second, the discreteness value of the cepstrum distribution is based on the formula: In the formula, For the The mean of the Mel-frequency cepstral coefficients corresponding to all cepstral components in the frame, For the The discreteness value of the cepstrum distribution of the frame; Calculate the speech spectrum dispersion index. The specific calculation formula is as follows: In the formula, is the speech spectrum dispersion index.
6. A Bluetooth headset-based speech translation system according to claim 5, characterized in that: Energy fluctuation coefficient of the generated speech and speech spectrum dispersion index A rhythm mutation intensity evaluation model is constructed, and the rhythm mutation index is generated by weighted summation. The specific calculation formula is as follows: In the formula, is the rhythm mutation index, and Speech energy fluctuation coefficient and speech spectrum dispersion index The non-zero weight coefficient of .
7. A Bluetooth headset-based speech translation system according to claim 6, characterized in that: Determine the pre-set rhythm mutation index threshold interval , and after determination, the generated rhythm mutation index The comparison is performed, and the intensity of the sudden change in the rhythm of the user's expression during the time period of speech speed fluctuation is evaluated based on the comparison results. The specific comparison analysis is as follows: like , the intensity of the sudden change in the rhythm of the sentence in the user's expression during the speech speed fluctuation period is low; like ,The intensity of the sudden change in the rhythm of the sentence in the user's expression during the speech rate fluctuation period is medium; like ,The intensity of the sudden change in sentence rhythm in the user's expression during the speech speed fluctuation period is high.
8. A Bluetooth headset-based speech translation system according to claim 7, characterized in that: In the speech control module, the corresponding speech processing strategy is selected and executed according to the evaluation results, specifically: If the evaluation result is low intensity, the speech processing strategy selected and executed includes: using the default frame cutting window parameters to divide the speech frames, maintaining the standard silence duration threshold and fixed endpoint detection method, and completing the speech segment definition in a conventional manner; If the evaluation result is medium intensity, the speech processing strategies selected and executed include: adjusting the speech frame cutting window size to increase the window duration, improve the frame overlap ratio, and extend the silence duration threshold to alleviate the risk of segment misjudgment caused by rhythm fluctuations; If the evaluation result is high intensity, the speech processing strategy selected and executed includes: based on adjusting the frame cutting window and silence threshold, turning on the cross-segment semantic context retention mechanism and enabling the delay endpoint judgment mode to avoid segment cutting errors and recognition timing disorder at the rhythm jump point; And dynamic regulation is carried out during the execution process. Specifically, during the execution of the speech processing strategy, the frame-level spectral energy change trend of the speech signal and the change status of the silence duration between speech segments are continuously monitored. If it is detected that the rhythm mutation index continues to rise in consecutive frames and exceeds the maximum value threshold corresponding to the current evaluation level during the execution of the speech processing strategy, it will immediately switch to a higher-level speech processing strategy, update the speech frame cutting window parameters and the silence duration threshold, and enable the cross-segment semantic preservation function; if it is detected that the rhythm mutation index continues to decrease in consecutive frames and is lower than the minimum value threshold corresponding to the current strategy level during the processing, the current speech processing strategy will be downgraded, the default frame division parameters will be restored, and the extended processing mechanism will be closed to achieve real-time dynamic adjustment of the processing strategy.
Citation Information
Patent Citations
Improved automated voice synthesis employing enhanced prosodic treatment of text, spelling of text and rate of annunciation
CA2594071A1
Voice speed self-adaptive recognition system
CN114067787A
Voice analysis method, device and equipment and storage medium thereof
CN118658466A
Digital human real-time interaction system and digital human real-time interaction method
CN119440254A
Intelligent real-time language synchronous translation system and terminal thereof
CN119920244A
Cited By
Multi-party dialogue translation method and device, electronic equipment and storage medium
CN121393468A
Multi-language adaptive identification method based on AI
CN121708902A
AI-based multilingual adaptive recognition method
CN121708902B