A voice translation system based on a Bluetooth headset

By introducing speech speed detection and rhythm evaluation modules into the Bluetooth headphone voice translation system, the speech processing strategy is dynamically regulated, and the frame cutting error problem caused by speech fluctuations is solved, which improves the accuracy and coherence of speech translation.

CN120164455BActive Publication Date: 2025-07-22VISION INTELLIGENCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510639144.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-22
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing Bluetooth headset voice translation technology cannot dynamically regulate the voice processing strategy when the user's speech speed fluctuates significantly, resulting in frame cutting errors, resulting in speech recognition errors and translation inaccuracy.

Method used

Preprocessing is performed through the voice acquisition module, combining the speech speed detection module and the rhythm evaluation module, the speech speed fluctuation time period is identified, and the speech processing strategy is dynamically adjusted through the speech control module, including the frame cutting window and the mute duration threshold, ensuring the accuracy of speech recognition.

Benefits of technology

It realizes accurate recognition of fluctuations in speech speed and rhythm changes in user speech expression, improves the stability and translation accuracy of the speech translation system, and maintains semantic coherence and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164455B_ABST
    Figure CN120164455B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice translation system based on a Bluetooth headset, which relates to the technical field of Bluetooth headset voice translation. The system includes a voice acquisition module, a speech rate detection module, a rhythm evaluation module, a voice regulation module, and a text verification module. The rhythm evaluation module extracts the rhythm feature information in the user voice signal collected by the Bluetooth headset during the speech rate fluctuation period, analyzes it, and evaluates the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation period. The voice regulation module selects and executes the corresponding voice processing strategy according to the evaluation result, and performs dynamic regulation during the execution process. The present invention solves the problem that the sentence rhythm mutation caused by the speech rate fluctuation cannot be dynamically regulated, realizes the intelligent adaptation and full-process response of the voice processing strategy, optimizes the structure of the translation result, and significantly improves the translation accuracy and semantic coherence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice translation for Bluetooth headsets, and particularly to a voice translation system based on Bluetooth headsets. Background Art

[0002] A voice translation system based on Bluetooth headsets is an intelligent communication system that integrates speech recognition, machine translation, and speech synthesis technologies. Its core lies in the real-time collection, processing, and output of speech through Bluetooth headsets. Users can wear Bluetooth headsets to transmit the spoken language to a mobile terminal or an edge computing device in real time. The system automatically recognizes the speech content, performs cross-language translation, and feeds back the target language result in the form of speech, thus achieving instant communication between two languages or even multiple languages. This system has the characteristics of high portability, hands-free operation, and natural operation, and is particularly suitable for scenarios such as international travel, international business, foreign language teaching, and auxiliary communication for special groups (such as hearing-impaired people). Its significance lies in breaking down language barriers and improving the communication efficiency between people, and it is one of the important directions for the integrated application of artificial intelligence and wearable devices.

[0003] The existing voice translation technology based on Bluetooth headsets usually uses Bluetooth headsets as the front-end speech collection and output devices, and cooperates with intelligent terminals (such as smartphones or dedicated translation devices) to achieve a complete voice translation process. Its voice translation process mainly includes the following key links: First, the user speaks the original language through the Bluetooth headset, and the headset transmits the voice signal to the terminal device in the form of audio data; Second, the built-in speech recognition module (ASR) in the terminal extracts features and performs modeling analysis on the received audio data, and converts it into the corresponding text content; Subsequently, the translation engine (MT) semantically understands the source language text and converts it into the text information of the target language, usually combined with neural network machine translation technology to improve accuracy and naturalness; Then, the speech synthesis module (TTS) converts the target language text into a natural voice signal; Finally, the translated speech is transmitted back to the user through the Bluetooth headset to achieve a complete cross-language interaction from speech to speech. The entire system may also be equipped with pre-processing modules such as speech enhancement, noise reduction, and echo cancellation to ensure high translation accuracy and usability in complex environments.

[0004] The existing technology has the following deficiencies:

[0005] During the process of a user using a Bluetooth headset for voice translation, if there are obvious fluctuations in the speaking speed during their expression, such as the speaking speed suddenly slowing down from a relatively fast pace when organizing language, or quickly resuming expression after a pause, it will cause a rhythm mutation in the voice signal in the time dimension. Such situations are common when the user is speaking while thinking, the content of the expression is complex, or the sentence structure is changed temporarily. It is difficult to maintain a stable time interval and beat structure of the voice. Since the Bluetooth headset only serves as a front-end audio acquisition device and does not have the function of perceiving rhythm changes itself, and the voice processing module at the back end of the system usually segments and processes the audio stream based on fixed frame cutting windows, silence thresholds, and endpoint detection algorithms, frame cutting errors are likely to occur when the speaking speed mutates, such as mis-segmenting a semantically complete speech segment into two segments, or wrongly merging two semantically independent speech segments, resulting in misalignment of the input to the speech recognition model and problems such as syllable mis-cutting and phoneme overlapping. The existing voice translation technology based on Bluetooth headsets cannot dynamically adjust the voice processing strategy according to the intensity of the sentence rhythm mutation when there are obvious fluctuations in the user's speaking speed during expression, and still performs unified frame processing according to the conventional speaking speed mode, which will cause incomplete interception of voice features or mixing of information between upper and lower sentences, and then lead to speech recognition errors, such as disordered word order, missing keywords, and broken semantic structures, seriously affecting the accuracy of the final translation content and the continuity of user communication.

[0006] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0007] The object of the present invention is to provide a voice translation system based on a Bluetooth headset to solve the problems in the above background art.

[0008] To achieve the above object, the present invention provides the following technical solution: A voice translation system based on a Bluetooth headset, comprising a voice acquisition module, a speaking speed detection module, a rhythm evaluation module, a voice regulation module, and a text verification module;

[0009] The voice acquisition module collects the user's voice signal through the Bluetooth headset, preprocesses the collected voice signal, and divides the voice signal into a continuous frame sequence according to a preset frame length and frame shift after preprocessing;

[0010] The speaking speed detection module establishes a time window based on the continuous frame sequence, analyzes the change state of the speaking speed within the time window, determines the time period when there are obvious fluctuations in the user's speaking speed during expression, and marks this time period as the speaking speed fluctuation time period, where the obvious speaking speed fluctuation means that the amplitude of the speaking speed fluctuation exceeds a preset threshold;

[0011] A rhythm evaluation module extracts rhythm feature information from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, analyzes it, and evaluates the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period;

[0012] A voice regulation module selects and executes corresponding voice processing strategies according to the evaluation results and performs dynamic regulation during the execution process;

[0013] A text verification module, after completing speech recognition, verifies and adjusts the language structure and word order of the recognized text to ensure that the translated text conforms to the grammar rules of the target language and maintains semantic coherence.

[0014] Preferably, in the speech rate detection module, the size of the time window is set based on the continuous frame sequence, and the time window is divided into several time periods; by analyzing the speech rate change state of each time period, specifically including analyzing the time interval between adjacent frames and the change in speech energy, the speech rate fluctuation amplitude of each time period is calculated; the time period with the speech rate fluctuation amplitude exceeding the preset threshold is determined as the time period when the user's speech rate fluctuates significantly during the expression, and it is marked as the speech rate fluctuation period; the speech rate fluctuation amplitude is determined by calculating the ratio of the time interval between adjacent frames to the speech energy difference.

[0015] Preferably, in the rhythm evaluation module, rhythm feature information is extracted from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, and preprocessing is performed on it after acquisition; acoustic fluctuation feature information and spectral configuration distribution information are extracted from the preprocessed rhythm feature information, and analysis is performed after extraction to generate a speech energy fluctuation coefficient and a speech spectrum dispersion index respectively; a rhythm mutation intensity evaluation model is constructed for the generated speech energy fluctuation coefficient and speech spectrum dispersion index, and a rhythm mutation index is generated through weighted summation; a preset rhythm mutation index threshold interval is determined, and after determination, it is compared with the generated rhythm mutation index, and the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period is evaluated according to the comparison result.

[0016] Preferably, the acquisition logic of the speech energy fluctuation coefficient is as follows:

[0017] Acoustic fluctuation feature information is extracted from the preprocessed rhythm feature information, specifically including the energy value and zero-crossing rate of each frame in the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, and they are respectively marked as and , represents the energy value of the th frame in the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, represents the Zero-crossing rate of the frame , is a positive integer;

[0018] Calculate the voice energy fluctuation coefficient, and the specific calculation formula is as follows:

[0019]

[0020] In the formula, is the voice energy fluctuation coefficient.

[0021] Preferably, the acquisition logic of the voice spectrum discrete index is as follows:

[0022] Extract the spectrum configuration distribution information from the preprocessed rhythm feature information, specifically including: obtaining the spectrum feature information of each frame in the user voice signal collected by the Bluetooth headset during the speech rate fluctuation time period, and constructing a Mel frequency cepstral coefficient vector based on the spectrum feature information of each frame, expressed as:

[0023]

[0024] In the formula, is the Mel frequency cepstral coefficient vector of the frame, containing components, represents the Mel frequency cepstral coefficient corresponding to the th cepstral component in the th frame, , , and are all positive integers;

[0025] Based on the Mel frequency cepstral coefficient vector of each frame, calculate the following two types of characteristic values:

[0026] First, the average offset value of the cepstral feature, according to the formula:

[0027]

[0028]

[0029] In the formula, is the average value of the Mel frequency cepstral coefficient corresponding to the th cepstral component in all frames, represents the average offset value of the cepstral feature of the th frame;

[0030] Second, the discrete degree value of the cepstral distribution, according to the formula:

[0031]

[0032]

[0033] In the formula, is the mean of the mel-frequency cepstral coefficients corresponding to all cepstral components in the th frame, is the numerical value of the cepstrum distribution dispersion of the th frame;

[0034] Calculate the speech spectrum dispersion index, and the specific calculation formula is as follows:

[0035]

[0036] In the formula, is the speech spectrum dispersion index.

[0037] Preferably, for the generated speech energy fluctuation coefficient and the speech spectrum dispersion index construct a rhythm mutation intensity evaluation model, and generate a rhythm mutation index through weighted summation. The specific calculation formula is as follows:

[0038]

[0039] In the formula, is the rhythm mutation index, and are the non-zero weight coefficients of the speech energy fluctuation coefficient and the speech spectrum dispersion index respectively, and .

[0040] Preferably, determine a preset rhythm mutation index threshold interval , and after determination, compare it with the generated rhythm mutation index , and evaluate the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation period according to the comparison result. The specific comparison analysis is as follows:

[0041] If , the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation period is low intensity;

[0042] If , the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation period is medium intensity;

[0043] If , the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation period is high intensity.

[0044] Preferably, in the speech regulation module, according to the evaluation result, select and execute the corresponding speech processing strategy, specifically:

[0045] If the evaluation result is low intensity, the selected and executed speech processing strategies include: dividing speech frames using the default frame cutting window parameters, maintaining the standard silent duration threshold and the fixed endpoint detection method, and defining speech segments in a conventional manner;

[0046] If the evaluation result is medium intensity, the selected and executed speech processing strategies include: adjusting the size of the speech frame cutting window to increase the window duration, increasing the frame overlap ratio, and extending the silent duration threshold to mitigate the risk of segmentation misjudgment caused by rhythm fluctuations;

[0047] If the evaluation result is high intensity, the selected and executed speech processing strategies include: on the basis of adjusting the frame cutting window and the silent threshold, enabling the cross-segment semantic context retention mechanism and the delayed endpoint judgment mode to avoid speech segment cutting errors and recognition timing disorders at rhythm jump points;

[0048] And dynamic regulation is carried out during the execution, specifically: during the execution of the speech processing strategy, continuously monitor the changing trend of the frame-level spectral energy of the speech signal and the changing state of the silent duration between speech segments. If it is detected during the execution of the speech processing strategy that the rhythm mutation index continuously rises in consecutive frames and exceeds the maximum threshold corresponding to the current evaluation level, immediately switch to a higher-level speech processing strategy, update the speech frame cutting window parameters and the silent duration threshold, and enable the cross-speech-segment semantic retention function; if it is detected during the processing that the rhythm mutation index continuously drops in consecutive frames and is lower than the minimum threshold corresponding to the current strategy level, downgrade the current speech processing strategy, restore the default frame division parameters and turn off the extended processing mechanism to achieve real-time dynamic adjustment of the processing strategy.

[0049] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:

[0050] 1. By constructing a speech speed detection module and a rhythm evaluation module, the present invention realizes the accurate recognition of the speed fluctuation and rhythm mutation of the user during the speech expression process. Especially for the problem of unstable speech speed caused by the user's thinking, pausing or sentence pattern switching in the voice translation scenario of Bluetooth headsets, it can monitor and judge the fluctuation characteristics in real time at the frame level of the speech signal. By introducing two composite acoustic indexes, namely the speech energy fluctuation coefficient and the speech spectrum dispersion index, and constructing a rhythm mutation index for intensity quantification evaluation, the solution realizes the mathematical modeling and discrimination ability of the degree of sentence rhythm change at the technical level, provides an accurate decision-making basis for the subsequent processing strategy, and effectively breaks through the limitation in the prior art that the rhythm mutation cannot be perceived and quantified.

[0051] 2. The present invention links the evaluation result with the speech processing strategy through the speech regulation module, and automatically switches to different levels of parameter combinations according to the intensity of rhythm mutation, including frame cutting window adjustment, mute duration determination delay, and context semantic preservation mechanism, etc., realizing the refinement, dynamicization, and scene adaptability of speech segmentation processing. Especially in the case of high-intensity mutation, the system not only avoids the problems of sentence cutting and semantic disorder caused by inaccurate frame division under the traditional strategy, but also improves the robustness of endpoint recognition through the delay judgment mechanism, making the recognition flow of the whole speech smoother, enhancing the fault tolerance ability of the system for complex speech structures and the translation stability.

[0052] 3. After the speech recognition is completed, the present invention performs software-level structural verification and semantic reconstruction on the language structure and word order of the recognized text through the text verification module, so that the translation output not only has grammatical compliance, but also maintains logical coherence and language naturalness, further improving the user-end perception experience. Without relying on additional hardware devices, the overall solution completes a complete closed-loop processing chain from front-end detection, process regulation to back-end semantic optimization based on the speech acquisition channel of the Bluetooth headset, significantly improving the intelligent response ability of the speech translation system to rhythmically complex expressions, and having strong practical value and industrial promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0054] Figure 1 It is a schematic diagram of the modules of a speech translation system based on a Bluetooth headset according to the present invention;

[0055] Figure 2 It is a system mind map of a speech translation system based on a Bluetooth headset according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0056] Now, the exemplary embodiments will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these exemplary embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the exemplary embodiments to those skilled in the art.

[0057] The present invention provides a Figure 1 and Figure 2 shown speech translation system based on a Bluetooth headset, including a speech acquisition module, a speech speed detection module, a rhythm evaluation module, a speech regulation module, and a text verification module;

[0058] The voice acquisition module collects the user's voice signal through a Bluetooth headset, preprocesses the collected voice signal, and divides the voice signal into a continuous frame sequence according to a preset frame length and frame shift after preprocessing;

[0059] The Bluetooth headset collects the user's voice signal through a built-in microphone and usually outputs it as a raw audio stream. Before using the voice signal for subsequent processing, preprocessing is required to improve the signal quality. The first step in preprocessing is usually denoising, which can remove background noise through frequency domain filtering or adaptive filtering algorithms. Next, pre-emphasis processing is performed, and a high-pass filter is applied to enhance the high-frequency part of the voice signal, which can compensate for the attenuation of the voice signal in the high-frequency range and improve the clarity and details of the signal. Finally, signal normalization processing is performed, and the amplitude range of the audio signal is adjusted to make it consistent, avoiding the influence of too large or too small signal amplitude on subsequent analysis and processing. Through these steps, preprocessing ensures the quality of the voice signal and provides a stable basis for subsequent analysis and feature extraction.

[0060] After the voice signal preprocessing is completed, the next step is to divide the audio signal according to the preset frame length and frame shift. A common approach is to use window function techniques, such as the short-time Fourier transform (STFT), to divide the audio signal into frames of multiple short-time windows. The frame length determines the number of sample points contained in each frame, usually 20 - 40 milliseconds, while the frame shift represents the step size of each window slide, often 10 milliseconds or 20 milliseconds. In this way, the voice signal is cut into multiple overlapping frames, and each frame represents the signal characteristics within a certain time period. The division of frames not only helps the subsequent feature extraction work (such as MFCC, spectrogram, etc.) proceed smoothly but also ensures that the dynamic changes of the voice can be captured within a smaller time window, avoiding the loss of important information due to too large a time window.

[0061] The purpose of doing this is to improve the accuracy of subsequent voice processing modules (such as voice recognition or translation). Through preprocessing, noise and unnecessary frequency components are removed, and the voice signal becomes cleaner and clearer, facilitating subsequent analysis. Frame division cuts the continuous voice signal into manageable small segments, facilitating more accurate feature extraction and analysis by the model in each time period. The information within each frame can be processed independently, which can avoid the loss of information during changes in speech rate or pauses, thus ensuring that sudden changes or long pauses in the voice signal do not lead to recognition errors or inaccurate translations. Therefore, these processing steps are crucial for improving the accuracy and stability of the voice translation system.

[0062] The speech rate detection module establishes a time window based on a continuous frame sequence, analyzes the speech rate change state within the time window, determines the time period when the user's speech rate fluctuates significantly during expression, and calibrates this time period as the speech rate fluctuation time period. A significant speech rate fluctuation means that the speech rate fluctuation amplitude exceeds a preset threshold.

[0063] In this embodiment, in the speech rate detection module, the size of the time window is set based on the continuous frame sequence, and the time window is divided into several time periods. By analyzing the speech rate change state of each time period, specifically including analyzing the time interval between adjacent frames and the change in speech energy, the speech rate fluctuation amplitude of each time period is calculated. The time period with a speech rate fluctuation amplitude exceeding the preset threshold is determined as the time period when the user's speech rate fluctuates significantly during expression, and it is calibrated as the speech rate fluctuation time period. Among them, the speech rate fluctuation amplitude is determined by calculating the ratio of the time interval between adjacent frames to the difference in speech energy.

[0064] When implementing the process of "setting the size of the time window based on the continuous frame sequence and dividing the time window into several time periods", first, a fixed time window size needs to be set according to the frame sequence of the speech signal, usually set in milliseconds. This time window size is usually adjusted according to the rhythm of the language and the requirements of speech processing. A common window size is from 20ms to 40ms. Then, the speech signal is divided according to the set window size. Specifically, each time window contains several frames (usually cut using a preset frame length and frame shift), and it is ensured that there is an overlap between these time periods to avoid information loss. The overlap size is usually set from 10ms to 20ms, that is, the setting of the frame shift, to ensure that the signal changes within each time period can be fully captured. In this way, the continuous signal can be cut, making the signal within each time period consistent and facilitating subsequent analysis. This method not only helps to analyze the change in speech rate in a finer granularity but also can better handle the rapid changes in the speech signal, especially when the user's speech rate fluctuates greatly, ensuring that no key speech change information is missed. When implemented by software, the frame cutting and time window setting are usually achieved through signal processing functions or tool libraries to ensure that this process is efficient and accurate.

[0065] In order to achieve "by analyzing the speech rate change state in each time period, specifically including analyzing the time interval between adjacent frames and the change in speech energy, and calculating the speech rate fluctuation amplitude in each time period", it is first necessary to conduct a detailed analysis of the frames within each time period. The time interval refers to the time difference between adjacent frames, which is usually determined by the set frame length and frame shift. The software can measure the speed change of the speech signal in this time period by calculating the time interval between adjacent frames in each time period. Secondly, the change in speech energy is usually represented by calculating the energy within each frame, and the change in energy reflects the intensity fluctuation of the speech signal. When implementing the software, the short-time energy calculation method can be adopted. By squaring and averaging the signal amplitude of each frame, and then analyzing the energy fluctuation based on the change in signal intensity within each time period. When calculating the speech rate fluctuation amplitude in each time period, the time interval between adjacent frames and the speech energy difference can be combined to obtain the fluctuation amplitude in each time period. This can be achieved by calculating the ratio of the time interval between adjacent frames and the energy difference, reflecting the intensity of the speech rate fluctuation in this time period. This calculation process helps to identify situations where the speech rate fluctuates significantly, especially when the user speaks quickly or slowly. The calculated fluctuation amplitude can provide an accurate basis for subsequent speech processing. This method can be implemented through existing signal processing tools and software libraries (such as MATLAB, NumPy in Python, etc.). Using techniques such as the Fast Fourier Transform (FFT), it is possible to efficiently process and calculate the time interval between adjacent frames and the change in speech energy in each time period.

[0066] In order to achieve "determining the time period with a speech rate fluctuation amplitude exceeding a preset threshold as the time period when the user has an obvious speech rate fluctuation during expression and labeling it as a speech rate fluctuation time period", it is first necessary to set a reasonable threshold for the speech rate fluctuation amplitude. The setting of this threshold is usually based on experience in actual applications or adjusted through experimental data to ensure that normal speech rate fluctuations can be distinguished from obvious speech rate fluctuations. During the implementation process, the software system will traverse the speech rate fluctuation amplitude of each time period and compare it with the preset threshold. If the speech rate fluctuation amplitude of a certain time period is greater than the threshold, it is considered that there is an obvious speech rate fluctuation in this time period. The software can use conditional judgment statements (such as if statements) to compare the fluctuation amplitude of each time period with the threshold one by one. If the condition is met, this time period will be labeled as a time period with an obvious speech rate fluctuation. During the labeling process, the software will record the start and end times of this time period to ensure that the time period with obvious fluctuations can be accurately identified. To further improve the accuracy, the system can also combine the duration to filter out short-term and accidental fluctuations. For example, if the speech rate fluctuation amplitude exceeds the threshold but the duration is too short (e.g., less than 50 ms), this time period can be ignored and not regarded as a time period with obvious fluctuations. In this way, the system can effectively distinguish the obvious speech rate fluctuations caused by reasons such as thinking or rapid topic switching, and provide reliable input data for subsequent speech recognition, translation, or other processing. Such a processing method can be implemented in the software through signal processing algorithms to ensure that the system efficiently identifies and labels the speech rate fluctuation time periods, avoids misjudgment, and improves the accuracy and effect of speech processing.

[0067] "Obvious speech rate fluctuation means that the speech rate fluctuation amplitude exceeds the preset threshold" means that when analyzing the speech signal, if the speech rate change amplitude of a certain time period exceeds the preset threshold, it is considered that there is an obvious speech rate fluctuation in this section of speech. This threshold is usually set through experiments or experience to ensure that the system can effectively distinguish daily normal speech rate fluctuations from drastic speech rate changes, thus avoiding false responses to some minor and irrelevant fluctuations. As for "the speech rate fluctuation amplitude is determined by calculating the ratio of the time interval between adjacent frames and the difference in speech energy", the core of this calculation method is to measure the speech rate fluctuation by observing the changes in the speech signal in adjacent time periods. Specifically, the time interval reflects the time difference between adjacent frames. Usually, the time interval is shorter when the speech rate is faster and longer when the speech rate is slower; while the difference in speech energy reflects the change in signal intensity. Usually, when the speech suddenly accelerates or slows down, the change in energy will be more obvious. Therefore, by calculating the ratio of the time interval between adjacent frames and the difference in energy, a value reflecting the intensity of the speech rate fluctuation can be obtained. When this value exceeds the preset threshold, it can be considered that within this time period, the speech rate fluctuation has reached an obvious level. This method can effectively capture important changes in the speech signal and provide an accurate reference basis for the speech processing system.

[0068] A rhythm evaluation module extracts rhythm feature information from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, analyzes it, and evaluates the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period.

[0069] In this embodiment, in the rhythm evaluation module, rhythm feature information is extracted from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period and preprocessed after acquisition; acoustic fluctuation feature information and spectral configuration distribution information are extracted from the preprocessed rhythm feature information and analyzed after extraction to generate a voice energy fluctuation coefficient and a voice spectrum discrete index respectively; a rhythm mutation intensity evaluation model is constructed for the generated voice energy fluctuation coefficient and voice spectrum discrete index, and a rhythm mutation index is generated through weighted summation; a preset rhythm mutation index threshold interval is determined and compared with the generated rhythm mutation index after determination, and the intensity of the rhythm mutation in the user's expression during the speech rate fluctuation period is evaluated according to the comparison result.

[0070] In the process of "extracting rhythm feature information from the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period", the speech data in this period can be processed frame by frame through a voice signal analysis software system, converting the continuous voice waveform into a series of frame-structured voice samples, and extracting basic parameters that can reflect the voice rhythm change characteristics in each frame. Specifically, based on the voice signal after frame division, the short-time Fourier transform (STFT) and Mel filter bank can be used to obtain the spectral structure of each frame, and then eigenvalue such as short-time energy, zero-crossing rate, spectral energy distribution, Mel-frequency cepstral coefficients (MFCC) are further extracted and arranged in time series to construct complete rhythm feature information. These features are all directly quantifiable numerical data, and software automation extraction can be realized through an audio signal processing library (such as librosa or MATLAB speech toolset) to ensure that the physical properties of rhythm changes can be captured and described by the system in the digital domain.

[0071] Preprocessing the rhythm feature information extracted above is to improve the accuracy and stability of subsequent parameter calculations and prevent data deviation caused by outliers, noise interference, or dimensional inconsistencies. The preprocessing operations generally include three aspects: First, normalize all the extracted frame-level feature data (such as z-score standardization or min-max scaling) to eliminate the influence between different feature dimensions and enable various data to be compared on the same scale; Second, smooth the short-term abnormal data points that appear in consecutive frames. The moving average method or median filtering can be used to reduce the misleading effect of mutation frames on the judgment of overall fluctuations; Third, use feature selection algorithms (such as principal component analysis PCA or feature ranking) to compress redundant or weakly correlated dimensions and retain the core features that are most representative of rhythm assessment. These preprocessing steps can be executed chainedly through software functions after the audio feature matrix is completed to ensure that the generated rhythm feature information has high comparability, high robustness, and a good mathematical structure, facilitating subsequent coefficient calculations and the accurate operation of the evaluation model.

[0072] In the process of "extracting acoustic fluctuation feature information and spectral configuration distribution information from the preprocessed rhythm feature information", it is first necessary to perform targeted splitting and structuring on the frame-level speech feature matrix that has been normalized and noise-reduced. For the acoustic fluctuation feature information, two key parameters that reflect the time-domain fluctuation characteristics of the speech signal, namely the short-time energy value and the zero-crossing rate value, can be selected to construct a numerical pair that reflects the strength change and amplitude fluctuation of the speech between different frames; Specifically, the system can extract the energy value and zero-crossing rate value of each frame according to the frame index and form a two-dimensional feature sequence in chronological order. For the spectral configuration distribution information, based on the extracted frame-level Mel-frequency cepstral coefficients (MFCC), two numerical features representing its frequency-domain fluctuation characteristics can be further extracted: First, calculate the mean absolute deviation between each frame of MFCC and the average value of the corresponding dimension as the cepstral offset intensity; Second, calculate the variance of the overall distribution of MFCC to reflect the discrete trend of frequency components. Through the above methods, the rhythm feature information is further separated into two feature information sets with clear structures, definite data dimensions, and independent evaluation values, and both are specific numerical data obtained through mathematical operation logic by the software system at the audio feature matrix level, meeting the input requirements of the subsequent evaluation model.

[0073] In the process of implementing "determining the preset rhythm mutation index threshold interval", statistical modeling and interval division of historical speech samples are usually carried out in a data-driven manner to establish a standard discrimination benchmark suitable for evaluation. Specifically, first, a large number of speech sample data containing different degrees of rhythm changes need to be collected in the training stage, and the rhythm feature information consistent with the current process is extracted from each sample. Then, the corresponding speech energy fluctuation coefficient and speech spectrum dispersion index are generated respectively, and the rhythm mutation index is synthesized according to the established weight. Subsequently, these index values can be clustered or sorted according to the actual rhythm state labels of the samples (such as "stable", "slight mutation", "severe mutation"), and statistical analysis is carried out on their distribution intervals. For example, the boundary values of mutation intensity are divided by the percentile method or Gaussian fitting method. Finally, a set of static threshold intervals can be selected within the statistical range with a high sample coverage rate. For example, 0.0 - 0.4 corresponds to stable, 0.4 - 0.7 is moderate mutation, 0.7 - 1.0 is severe mutation, etc. This interval can be solidified in the system in a parameter configuration manner and can be adaptively adjusted according to the subsequent model operation effect. The whole process can be completed through the data annotation module, clustering algorithm module and parameter management module in the software, without manually setting subjective values, ensuring that the division standard has statistical significance and system stability.

[0074] In this embodiment, the acquisition logic of the speech energy fluctuation coefficient is as follows:

[0075] Extract the acoustic fluctuation feature information from the preprocessed rhythm feature information, specifically including the energy value and zero-crossing rate of each frame in the user speech signal collected by the Bluetooth headset during the speech rate fluctuation time period, and mark them respectively as and , represents the energy value of the -th frame in the user speech signal collected by the Bluetooth headset during the speech rate fluctuation time period, represents the zero-crossing rate of the -th frame in the user speech signal collected by the Bluetooth headset during the speech rate fluctuation time period, , is a positive integer;

[0076] In the user voice signals collected by the Bluetooth headset during the speech rate fluctuation period, the energy value and zero-crossing rate of each frame can be extracted through audio signal processing algorithms in software, and they belong to the two most commonly used low-order acoustic features. First, the energy value (Energy Value) refers to the short-time energy intensity of each frame of the voice signal, reflecting the overall amplitude size of the voice waveform within that frame. The acquisition method is as follows: divide the preprocessed voice signal into a series of equally long frames, perform a square operation on the sampling points in each frame and then sum or average them to obtain the total energy of that frame; this process can be achieved through a sliding window method to ensure frame-to-frame continuity. The larger the energy value, the louder or more impactful the voice signal of that frame, and it often appears in stressed, explosive, or fast onset sounds. Second, the zero-crossing rate (Zero-Crossing Rate, ZCR) refers to the number of times the waveform of the voice signal crosses the zero axis from positive to negative or negative to positive within a frame period, reflecting the degree of waveform oscillation or irregularity. The specific acquisition method is: compare the consecutive sampling points of each frame signal pairwise, record a zero-crossing event when a sign change is found, and then divide the total number of times by the number of samples within the frame to obtain the normalized ZCR value. When the ZCR is relatively high, it usually corresponds to voice segments with many high-frequency details such as silent segments, fricative sounds, and nasalized sounds; when the ZCR is relatively low, it mostly appears in vowels, voice pauses, or low-frequency smooth sentences. Both types of data can be automatically extracted in a software environment using standard audio analysis libraries (such as Librosa in Python, Audio Toolbox in MATLAB, etc.) without manual intervention, and they have high repeatability and numerical stability, and are the basic input features for constructing a speech rhythm analysis model.

[0077] Calculate the speech energy fluctuation coefficient, and the specific calculation formula is as follows:

[0078]

[0079] In the formula, is the speech energy fluctuation coefficient.

[0080] The design of the calculation formula for this speech energy fluctuation coefficient aims to comprehensively measure the collaborative fluctuation characteristics of the voice signal in two dimensions: amplitude intensity and waveform oscillation frequency during the speech rate fluctuation period, so as to quantify the severity of speech rhythm changes. First, Squaring the energy value of each frame is to enhance the weight of high-energy frames, so that frames with a sudden increase in speech intensity during the speech rate fluctuation period have a more significant contribution to the overall fluctuation perception. Second, the multiplied by it is the regularized representation of the zero-crossing rate of that frame, is an index that measures the number of positive and negative alternations within a frame of a speech signal, representing the oscillation frequency of the waveform. The larger its value, the more frequent the speech fluctuations. Adding 1 is used to prevent zero values in multiplication. The structure of their multiplication embodies a key concept: having only strong energy or only high oscillation is not sufficient to indicate a drastic rhythm change. Only when both the energy and the waveform oscillation increase simultaneously within the same frame does it indicate a significant rhythm fluctuation. Therefore, taking the weighted average of such "energy × oscillation" values for each frame (i.e., summing over the frames and then dividing by ), the overall co - fluctuation intensity of the speech rate fluctuation period is obtained. Finally, it is compressed through the natural logarithm function to smooth extreme values, enhance the discrimination ability for medium - low amplitude fluctuations, and ensure good numerical stability and comparability of this coefficient. The overall formula not only has a compact mathematical structure but also, starting from the physical characteristics of speech, can effectively capture the comprehensive performance of rhythm fluctuations.

[0081] There is a positive correlation between the numerical value of the speech energy fluctuation coefficient and "evaluating the intensity of sudden rhythm changes in the user's expression during the speech rate fluctuation period", that is, the larger this coefficient, the more drastic the change in the energy intensity of the speech signal during this period, and the more frequent the accompanying waveform oscillation, thus reflecting that there is a greater likelihood of sudden rhythm changes in the speech. Specifically, this coefficient can comprehensively capture the drastic fluctuations of speech in both the amplitude and structure dimensions by multiplying the square of the energy of each frame by the zero - crossing rate of the corresponding frame and then taking the weighted average. When the user suddenly slows down or speeds up the speech rate during expression, or there are obvious discontinuities or breaks in the speech rhythm during language organization, there are often syllables with drastic changes in strength and frequent acoustic wave jumps, and at this time, the energy fluctuation coefficient will increase significantly. Therefore, the higher the value of this coefficient, the greater the change in the rhythm continuity of the speech, so it can be used to quantify the intensity of sudden rhythm changes in the speech and provide a key basis for subsequent translation optimization or recognition strategy adjustment.

[0082] In this embodiment, the acquisition logic of the speech spectrum discrete index is as follows:

[0083] Extract the spectral configuration distribution information from the pre - processed rhythm feature information, specifically including: obtaining the spectral feature information of each frame in the user speech signal collected by the Bluetooth headset during the speech rate fluctuation period, and constructing a Mel - Frequency Cepstral Coefficient (MFCC) vector based on the spectral feature information of each frame, expressed as:

[0084]

[0085] In the formula, is the Mel - Frequency Cepstral Coefficient vector of the th frame, which contains components, represents the th frame, and the the Mel-frequency cepstral coefficient corresponding to one cepstral component, , , and are all positive integers;

[0086] During the speech rate fluctuation period, in order to obtain the spectral feature information of each frame from the user voice signal collected by the Bluetooth headset and construct the Mel-frequency cepstral coefficient vector (MFCC), the standard speech signal analysis process can be implemented through software. Specifically, first, the collected speech signal is preprocessed, including pre-emphasis, framing, and windowing, to make the original time-domain signal have good stationarity in the time dimension. Subsequently, the fast Fourier transform (FFT) is performed on each frame of the signal to obtain the spectral energy distribution, thereby forming the power spectrum of this frame. Then, the Mel filter bank is applied to filter the power spectrum. The design of the filter follows the non-linear characteristics of the human auditory system's perception of frequency, mapping the frequency domain to the Mel scale, so that the low-frequency part is subdivided and the high-frequency part is compressed, in order to more effectively capture the perceptual characteristics of speech. The output of each Mel filter after filtering represents the energy value of a frequency band. Then, a logarithmic transformation is performed on these frequency band energy values, so that the relative differences of high energy values are compressed while the changes of low energy values are retained, enhancing the sensitivity to weak speech signals. Finally, a discrete cosine transform (DCT) is performed on the Mel energy sequence after logarithmic transformation to remove the correlation in the frequency domain and compress the signal energy into the first several coefficients, obtaining a set of compact, low-dimensional, and well-discriminative Mel-frequency cepstral coefficient vectors (MFCC). Each component of this vector is a definite real value, reflecting the cepstral characteristics of the speech frame in different frequency bands. Usually, the first g cepstral components are selected for analysis. The entire processing process can be completed in a software manner in a local or embedded system through a standard speech processing library (such as Librosa, Python SpeechFeatures, MATLAB Audio Toolbox, etc.), with high efficiency, versatility, and scalability. The constructed MFCC vector provides a clear and structured numerical basis for subsequent extraction of rhythm-related parameters (such as cepstral shift, cepstral dispersion).

[0087] Based on the Mel-frequency cepstral coefficient vector of each frame, calculate the following two types of characteristic numerical values:

[0088] First, the average offset numerical value of the cepstral feature, according to the formula:

[0089]

[0090]

[0091] In the formula, is the The average of the Mel-frequency cepstral coefficients corresponding to each cepstral component over all frames, which is used to reflect the global central tendency of this frequency-domain feature dimension. Indicates the average offset value of the cepstral features for the

[0092] Frame. The average offset value of the cepstral features is an important parameter for measuring the degree to which the spectral structure of the frame of speech signal deviates from the overall average state. Its essence is the weighted average of the difference between the Mel-frequency cepstral coefficient vector of this frame and the mean of the cepstral components in all frames. In the actual construction process, first, for each dimension of the cepstral component and its corresponding average value of this dimension in all frames within the entire speech rate fluctuation time period an absolute difference operation is performed to reflect the degree of deviation of the spectral form of this frame in this frequency dimension from the global central tendency; subsequently, the differences in all frequency band dimensions are averaged to obtain a comprehensive offset index . The larger this value, the more inconsistent the overall spectral form of this frame is compared with the "central spectral form" of other frames, indicating that there may be phenomena such as frequency reconstruction, change in pronunciation method, or sudden change in vocalization mode in speech expression. Specific manifestations may include sudden consonant bursts, rapid vocal cord closures, frequency jumps, or pronunciation mouth shape conversions. Therefore, high often appears at the junctions of sentence rhythm jumps, rapid changes in speech rate, or syntactic structure transitions. If

[0093] the value is small, it indicates that the spectral form of this frame tends to be stable and is highly consistent with the average expression spectrum during the speech rate fluctuation time period, representing a continuous and stable state of speech expression. Therefore, this value not only reflects the gap between the local and overall spectral forms but also provides an important spectral offset basis for subsequent judgment of the rhythm mutation trend.

[0093] Second, the cepstral distribution dispersion value, according to the formula:

[0094]

[0095]

[0096] where is the mean of the Mel-frequency cepstral coefficients corresponding to all cepstral components in the frame, representing the local central tendency of the spectral feature distribution within the frame, is the cepstral distribution dispersion value for the

[0097] Frame; The cepstral distribution dispersion value The quantitative index of the fluctuation amplitude of the internal structure of the frame Mel frequency cepstral coefficients is essentially the degree of dispersion of the components of the cepstral coefficient vector of the frame around the mean value within the frame. Specifically, first calculate all the cepstral components of the frame The mean , then square the deviations of each component from the mean and average them to get the cepstrum distribution dispersion of the frame This process is equivalent to calculating the statistical variance of all spectral features of the frame in the frequency domain dimension, so The larger the value, the more significant the difference between the spectral forms of the frame in different frequency bands, and the more uneven the spectral energy distribution, which reflects that there may be mixed pronunciation, violent harmonic fluctuations, or sudden changes in formants in the speech signal. This situation is often accompanied by sudden changes in the pronunciation of speech, such as the transition from vowels to affricates or compound sounds, or structural disturbances such as short pauses, swallowing, and changes in oral dynamics in the pronunciation. On the contrary, if A smaller value indicates that the cepstrum structure of the frame is relatively smooth and the spectrum changes of each frequency band are relatively consistent, which means that the speech expression of the frame has good continuity and stability in the spectrum structure. It is the key basis for identifying whether the speech rhythm has internal spectrum diffusion or disturbance. Its value can be used to assist in judging whether there is a mutation trend in the sentence rhythm.

[0098] Calculate the speech spectrum dispersion index. The specific calculation formula is as follows:

[0099]

[0100] In the formula, is the speech spectrum dispersion index.

[0101] The speech spectrum dispersion index The calculation formula is to comprehensively measure the instability of the spectral structure and the overall deviation trend of each frame of speech during the time period of speech rate fluctuation, so as to sensitively identify the possibility of sudden changes in sentence rhythm. Indicates The discreteness value of the frame cepstrum distribution is used to characterize the energy difference between the frequency bands within the frame. The larger the value, the more unbalanced the spectrum of the frame is, reflecting the local pronunciation disturbance or harmonic structure change. Indicates the average offset value of the cepstrum feature, which is used to indicate the degree of deviation between the overall spectrum structure of the frame and the average spectrum shape during the speech rate fluctuation period. The larger the value, the more obvious changes in the speech state, such as changes in pronunciation and reorganization of the vocal tract structure. Adding the two together can take into account both the "internal spectrum fluctuation" and "overall structure offset" dimensions, forming a composite expression that is more sensitive to sudden changes in speech rhythm. Then, the total fluctuation intensity is nonlinearly amplified through an exponential function, so that those in The frames with significant peaks in the overall The contribution is greater, thereby improving the system's ability to identify local rhythm anomalies. Finally, the calculation results of all frames are averaged, which not only retains the sensitivity of frame-level mutation recognition, but also achieves stable quantification of the rhythm state of the entire speech expression. This calculation method has the advantages of clear theoretical explanation, stable engineering implementation, and natural numerical trend.

[0102] Speech Spectral Dispersion Index The value of is positively correlated with the intensity of the rhythm mutation of the user's expression during the time period of speech speed fluctuation, that is, The larger the value, the higher the possibility and intensity of the sudden change in the rhythm of the sentence. The index is obtained by adding the discrete value of the cepstral distribution of each frame and the average offset value of the cepstral feature and then amplifying it exponentially. It can simultaneously reflect the changes in the local spectrum structure fluctuation and the overall spectrum contour offset of the speech signal. When it is large, it indicates that the spectrum energy in the frame is unevenly distributed among the frequency bands, and there may be harmonic disturbance, unstable pronunciation or structural diffusion. When it is large, it means that the frame spectrum deviates significantly from the average spectrum of the entire speech segment, which may be accompanied by a change in pronunciation or a sudden change in speech rhythm. The two together constitute the dual criterion for rhythm mutation. By summing them and applying exponential operations, the influence of the mutation frame in the entire speech can be further amplified. The larger the value, the more drastic and unstable the frame-level spectrum structure changes during the speech rate fluctuation period, and the more likely the overall speech rhythm will experience a sharp turn, jump, or discontinuous expression, thus indicating a stronger intensity of sentence rhythm mutation. A smaller value means that the speech expression is continuous, stable, and has a smooth rhythm, with basically no obvious rhythm changes.

[0103] In this embodiment, the generated speech energy fluctuation coefficient and speech spectrum dispersion index A rhythm mutation intensity evaluation model is constructed, and the rhythm mutation index is generated by weighted summation. The specific calculation formula is as follows:

[0104]

[0105] In the formula, is the rhythm mutation index, and Speech energy fluctuation coefficient and speech spectrum dispersion index The non-zero weight coefficient of .

[0106] During the evaluation of the rhythm mutation intensity, to achieve a comprehensive determination of the rhythm mutation trend in the user's speech expression, the software can be used to perform weighted fusion on the previously calculated speech energy fluctuation coefficient (EFC) and speech spectrum dispersion index (SDI) to construct a rhythm mutation index (RMI). Specifically, during implementation, first, the two indicators are respectively normalized to the same numerical scale (such as linearly mapped to the interval [0,1]) to eliminate the scale differences brought by different calculation formulas; subsequently, two non-zero weight coefficients and are introduced, which are respectively multiplied by the normalized EFC and SDI, and then the two items are summed after weighting to generate the final RMI value. The settings of the two weight coefficients need to satisfy and and are both not zero to ensure that both indicators make substantial contributions in the evaluation. The values of the weight coefficients can be set according to the training corpus or expert experience. For example, when more emphasis is placed on the change in the spectrum shape, the value of can be appropriately increased, while if more attention is paid to the energy dynamic fluctuation, the proportion of can be increased. This weighted fusion method not only retains the discriminant ability of the two indicators respectively but also can adapt to the dynamic adjustment requirements under different speech styles and scenarios, realizing a quantitative, adjustable, and sensitive expression of the rhythm mutation intensity of sentences during the speech rate fluctuation period.

[0107] In this embodiment, a pre-set rhythm mutation index threshold interval is determined, and after determination, it is compared with the generated rhythm mutation index , and the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation period is evaluated according to the comparison result. The specific comparison and analysis are as follows:

[0108] If , the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation period is low intensity;

[0109] This situation indicates that during the current speech rate fluctuation period, the fluctuation degree of the user's speech signal in the spectral structure is small, without significant local band energy dispersion and no obvious overall spectral contour shift. This shows that although there are slight speech rate fluctuations during the user's speech expression process, the sentence rhythm as a whole remains coherent, stable, and smooth, without forming a risk of interfering with speech segmentation, recognition, or translation. Therefore, this type of expression usually does not require special speech processing strategies and can continue to be recognized and translated according to the conventional speech rate and silence threshold processing mechanism, avoiding unnecessary consumption of computing resources and response delays, thereby improving the overall system operation efficiency and response speed.

[0110] If , the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation period is medium intensity;

[0111] This situation indicates that the user's speech signal has shown a certain degree of instability in the rhythm structure during the current speech rate fluctuation period. Specifically, it is manifested as an increase in the degree of discrete frequency band energy in some frames, or the overall spectral profile begins to deviate from the mainstream average spectrum, but it has not reached the state of severe perturbation. Such speech states usually occur at sentence connection points, in the middle of complex sentence patterns, or at semantic turning points, and have the potential risk of rhythm jumps. If processed conventionally, it may lead to inaccurate speech frame division, resulting in recognition delay or semantic breakage. Therefore, this type of moderate-intensity rhythm mutation needs to trigger mild strategy intervention, such as extending the speech recognition end point judgment time, adjusting the frame window size, or appropriately delaying the silence judgment threshold, to enhance the tolerance of the recognition module to rhythm uncertainty and improve translation continuity and accuracy.

[0112] If , during the speech rate fluctuation period, the intensity of the sentence rhythm mutation in the user's expression is high intensity.

[0113] This situation indicates that the speech expression state of the user during the speech rate fluctuation period has undergone a significant mutation. Its frame-level spectral structure shows strong discontinuity and structural jumps, which may include complex phenomena such as sudden pronunciation mode switching, rapid speech segment splicing, sudden change of tone, short-term repetition, syllable swallowing, etc. In this case, the frequency band energy fluctuates violently, and the average spectral profile deviates greatly, which will seriously interfere with the conventional recognition path of the speech processing module. If still processed according to the standard algorithm, it is very easy to cause speech segmentation misalignment, keyword truncation, or misactivation of the recognition model, and ultimately lead to word order disorder, semantic misinterpretation, and sentence meaning interruption in the translated content. Therefore, this level must trigger advanced dynamic intervention strategies, such as adjusting the frame overlap ratio, enabling the context semantic preservation mechanism, delaying the recognition decision, etc., to stabilize the recognition process from multiple dimensions and ensure the system still has robust translation accuracy and usability in high-dynamic speech scenarios.

[0114] The speech regulation module selects and executes corresponding speech processing strategies according to the evaluation results, and performs dynamic regulation during the execution process;

[0115] In this embodiment, in the speech regulation module, according to the evaluation results, the corresponding speech processing strategies are selected and executed, specifically:

[0116] If the evaluation result is low intensity, the speech processing strategies selected and executed include: dividing speech frames using the default frame cutting window parameters, maintaining the standard silence duration threshold and the fixed end point detection method, and defining speech segments in a conventional manner;

[0117] In the case where the evaluation result is of low intensity, the system can perform voice frame partitioning and speech segment delimitation processing using default parameters in software. The specific method is as follows: First, in the preprocessed voice signal, the voice stream is partitioned frame by frame according to a predefined frame length (such as 25 ms) and frame shift (such as 10 ms), without adjusting the window overlap ratio and frame duration, to maintain an even coverage of the voice signal in the time dimension; Then, the standard energy threshold and zero-crossing rate determination rules in the silence detection module are applied to identify the silent frames and voiced frames in the voice frames, and it is determined whether the speech segment terminates by whether the duration of consecutive silent frames exceeds the default silence duration threshold (for example, 200 ms); In the endpoint detection stage, a fixed start-stop recognition logic is executed, that is, the speech segment boundaries are calibrated through the starting point of the sound energy and the ending point of the silence, without performing delay compensation or context extension processing. This method can efficiently process speech paragraphs with stable rhythms at low computational cost, avoid unnecessary investment in computing resources, and at the same time reduce the processing delay. Since under low-intensity and sudden rhythm change conditions, the influence of speech rate fluctuations on the accuracy of speech segment cutting is small, the integrity and clarity of the input to the speech recognition module can be ensured without dynamic parameter intervention, thus maintaining the overall operating efficiency of the system and the stability of speech translation.

[0118] If the evaluation result is of medium intensity, the selected and executed speech processing strategies include: adjusting the size of the voice frame cutting window to increase the window duration, increasing the frame overlap ratio, and extending the silence duration threshold to mitigate the risk of misjudgment in segmentation caused by rhythm fluctuations;

[0119] When the evaluation result is of medium intensity, the system can dynamically adjust the speech processing parameters in software to enhance the adaptability to mild rhythm fluctuations. The specific implementation methods include: First, increase the duration of the voice frame cutting window (such as increasing from the default 25 ms to 35 ms), and at the same time increase the frame overlap ratio (such as increasing from 40% to 60%). By increasing the time coverage range and enhancing the inter-frame correlation, the spectral features have a stronger continuous expression on the time axis, thus slowing down the influence of the speech rate change on the determination of the speech segment boundary; Second, in the silence detection module, extend the silence duration threshold from the default value (such as 200 ms) to a higher value (such as 300 ms), so that when the system detects consecutive silent frames, it will not misjudge as the termination of the speech segment prematurely, reducing the speech segment segmentation errors caused by the slowdown of the speech rate or short pauses; The whole process is automatically completed by the speech regulation module calling the parameter configuration interface before the strategy is executed, without hardware intervention, and can be adjusted in real time in response to the evaluation result. The core of this strategy is to improve the fault tolerance of the speech processing system to mild rhythm changes by expanding the time coverage granularity of the frame and the silence buffer time, avoiding the wrong splitting or truncation of sentences, and ensuring the integrity and coherence of the translation semantics.

[0120] If the evaluation result is high intensity, the selected and executed speech processing strategies include: on the basis of adjusting the frame cutting window and the silence threshold, enabling the cross-segment semantic context retention mechanism and enabling the delayed endpoint judgment mode to avoid segment cutting errors and recognition timing disorders at the rhythm jump points;

[0121] In the case of a high-intensity evaluation result, to cope with the drastic rhythm mutations in speech and prevent problems such as incorrect segment cutting and recognition misalignment, the system can implement multi-level advanced speech processing strategies in software. First, on the basis of the adjusted frame cutting window duration (such as extended to 40 ms) and the silence duration threshold (such as increased to 400 ms), further enable the cross-segment semantic context retention mechanism, that is, during the speech recognition stage, continuously retain a certain number of recognized text buffer frames or language segments before and after the current segment (for example, the first 2 seconds and the last 1 second). By constructing a context association window, the context of the sentences before and after is used as one of the input features to participate in semantic determination during translation and recognition, thereby reducing recognition mismatches caused by the sudden intrusion at the beginning of the sentence or the interruption at the end of the sentence; at the same time, enable the delayed endpoint judgment mode, that is, when a possible segment end signal (such as silence timeout) is detected, do not immediately confirm the termination of the speech, but set a delayed confirmation interval (such as 100 ms - 150 ms), and continue to monitor the change of the speech input during this time period. If audio activity is detected again, cancel the termination determination, thereby avoiding short pauses caused by rhythm mutations from being mistaken for the end of the sentence. The above mechanism significantly enhances the adaptability of the speech processing system to strong rhythm mutation scenarios without affecting the real-time performance of the system through methods such as dynamic caching, predictive control, and endpoint trigger backshift, ensuring the integrity, context coherence, and semantic accuracy of the translation output.

[0122] And perform dynamic regulation during the execution process, specifically: during the execution of the speech processing strategy, continuously monitor the change trend of the frame-level spectral energy of the speech signal and the change state of the silence duration between speech segments. If it is detected during the execution of the speech processing strategy that the rhythm mutation index continuously rises in consecutive frames and exceeds the maximum threshold corresponding to the current evaluation level, immediately switch to a higher-level speech processing strategy, update the speech frame cutting window parameters and the silence duration threshold, and enable the cross-segment semantic retention function; if it is detected during the processing that the rhythm mutation index continuously decreases in consecutive frames and is lower than the minimum threshold corresponding to the current strategy level, downgrade the current speech processing strategy, restore the default frame division parameters and turn off the extended processing mechanism to achieve real-time dynamic adjustment of the processing strategy.

[0123] During the execution of the speech processing strategy, dynamic regulation can be achieved by constructing a real-time monitoring and strategy level linkage mechanism for the rhythm mutation index (RMI) in software. The specific implementation methods include: within each processing cycle, at a fixed time interval (such as every 100 milliseconds), incrementally update the rhythm features of newly acquired frames during the speech rate fluctuation period, calculate the corresponding speech energy fluctuation coefficient (EFC) and speech spectrum dispersion index (SDI) in real time, and generate the latest rhythm mutation index RMI; the system forms a trend sequence with RMI values in multiple consecutive periods (such as 5 consecutive frames), and judges whether there is a significant increase or decrease based on the change trend of this sequence. If it is detected that the RMI continuously rises and its average value exceeds the maximum threshold value corresponding to the current processing strategy, it indicates that the rhythm mutation is intensifying. At this time, the system immediately switches to a higher-level processing strategy through the parameter interface, automatically modifies the current frame cutting window length, frame shift parameter and silence determination threshold, and activates the cross-segment semantic context preservation function to enhance the segmental robustness and semantic continuity of the system under the condition of severe rhythm disturbance; on the contrary, if it is detected that the RMI continuously decreases and its average value is lower than the minimum threshold value corresponding to the current strategy level, the system judges that the current speech has returned to a stable state, automatically downgrades the processing strategy to a lower intensity level, restores the default parameter settings and closes the additional processing module, so as to avoid resource waste and recognition overfitting. The entire regulation process relies on a continuous indicator stream and judgment logic to achieve real-time strategy adjustment without interrupting the recognition process, effectively adapting to the rapid changes in the speech expression state, and ensuring the stability and accuracy of the speech translation system in dynamic scenarios.

[0124] The text verification module, after completing speech recognition, verifies and adjusts the language structure and word order of the recognized text to ensure that the translated text conforms to the grammar rules of the target language and maintains semantic coherence.

[0125] After completing speech recognition, the text verification module can automatically verify and make necessary adjustments to the language structure and word order in the recognition result through software, so as to improve the accuracy and coherence of the translated text at the grammar and semantic levels. Specifically, in the implementation, first parse the syntactic structure of the text obtained by speech recognition. Dependency syntactic analysis or phrase structure analysis methods based on natural language processing (NLP) can be used to identify language components such as the subject-predicate-object structure, modification relationship and sentence pattern boundary in the text. At the same time, with the help of part-of-speech tagging, entity recognition and phrase chunking technologies, clarify the grammatical roles played by each word in the sentence, and establish a language rule library based on the grammar norms of the target language or rely on a trained language model, such as a model based on tree-based reordering or a sequence learning model (such as BiLSTM-CRF), to perform pattern recognition and reconstruction correction on problems such as subject-predicate inversion, object omission, and modifier dislocation in the sentence.

[0126] Furthermore, to ensure the rationality of the word order and the coherence of the semantics in the translation output, the system can also judge and optimize the connection logic between multiple sentence outputs based on the context consistency evaluation and semantic association scoring mechanism. The implementation methods include: modeling the coherence of the language flow through dimensions such as the lexical cohesion between upper and lower sentences, the consistency of temporal adverbs, and the persistence of the subject, identifying and eliminating semantic jumps, abrupt tones, or logical breaks between sentences caused by speech recognition errors. In addition, by combining the target language language model to score the candidate word order schemes, the highest-confidence structure is selected as the final corrected text output. The above process is fully implemented at the software level without manual intervention, which can effectively enhance the language expression quality of the final translation result on the premise of ensuring the automated operation of the system, enabling users to obtain a translation text output that conforms to language habits, has a reasonable structure, and a complete semantics.

[0127] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0128] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0129] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0130] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0131] In several embodiments provided by this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the above-described embodiments are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0132] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0133] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0134] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A voice translation system based on a Bluetooth headset, characterized in that, It includes a voice acquisition module, a speech rate detection module, a rhythm evaluation module, a voice regulation module, and a text verification module; The voice acquisition module collects the user's voice signal through a Bluetooth headset, preprocesses the collected voice signal, and divides the voice signal into a continuous frame sequence according to a preset frame length and frame shift after preprocessing; The speech rate detection module establishes a time window based on the continuous frame sequence, analyzes the speech rate change state within the time window, determines the time period when the user has an obvious speech rate fluctuation during expression, and labels this time period as the speech rate fluctuation time period. The obvious speech rate fluctuation means that the speech rate fluctuation amplitude exceeds a preset threshold; The rhythm evaluation module extracts the rhythm feature information in the user's voice signal collected by the Bluetooth headset during the speech rate fluctuation time period, analyzes it, and evaluates the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation time period; The voice regulation module selects and executes the corresponding voice processing strategy according to the evaluation result, and performs dynamic regulation during the execution process; The text verification module, after completing the speech recognition, verifies and adjusts the language structure and word order of the recognized text to ensure that the translated text conforms to the grammar rules of the target language and maintains semantic coherence.

2. The voice translation system based on a Bluetooth headset according to claim 1, characterized in that, In the speech rate detection module, the size of the time window is set based on the continuous frame sequence, and the time window is divided into several time periods; by analyzing the speech rate change state of each time period, specifically including analyzing the time interval between adjacent frames and the change of speech energy, calculating the speech rate fluctuation amplitude of each time period; determining the time period with the speech rate fluctuation amplitude exceeding the preset threshold as the time period when the user has an obvious speech rate fluctuation during expression, and labeling it as the speech rate fluctuation time period; the speech rate fluctuation amplitude is determined by calculating the ratio of the time interval between adjacent frames and the speech energy difference.

3. The voice translation system based on a Bluetooth headset according to claim 2, wherein, In the rhythm evaluation module, the rhythm feature information in the user's voice signal collected by the Bluetooth headset during the speech rate fluctuation time period is extracted and preprocessed after acquisition; the acoustic fluctuation feature information and the spectral configuration distribution information are extracted from the preprocessed rhythm feature information, and analyzed after extraction to generate the speech energy fluctuation coefficient and the speech spectrum dispersion index respectively; a rhythm mutation intensity evaluation model is constructed for the generated speech energy fluctuation coefficient and speech spectrum dispersion index, and the rhythm mutation index is generated by weighted summation; a preset rhythm mutation index threshold interval is determined, and after determination, it is compared with the generated rhythm mutation index, and the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation time period is evaluated according to the comparison result.

4. The voice translation system based on a Bluetooth headset according to claim 3, wherein The acquisition logic of the speech energy fluctuation coefficient is as follows: Extract the acoustic fluctuation feature information from the preprocessed rhythm feature information, specifically including the energy value and zero-crossing rate of each frame in the user voice signal collected by the Bluetooth headset during the speech rate fluctuation time period, and calibrate them respectively as and , represents the energy value of the th frame in the user voice signal collected by the Bluetooth headset during the speech rate fluctuation time period, represents the zero-crossing rate of the th frame in the user voice signal collected by the Bluetooth headset during the speech rate fluctuation time period, , is a positive integer; Calculate the speech energy fluctuation coefficient, and the specific calculation formula is as follows: In the formula, is the voice energy fluctuation coefficient.

5. The voice translation system based on a Bluetooth headset according to claim 4, wherein The acquisition logic of the speech spectrum dispersion index is as follows: Extract the spectral configuration distribution information from the preprocessed rhythm feature information, specifically including: obtaining the spectral feature information of each frame in the user's voice signal collected by the Bluetooth headset during the speech rate fluctuation time period, and constructing a Mel-frequency cepstral coefficient vector based on the spectral feature information of each frame, expressed as: wherein, is the Mel-frequency cepstral coefficient vector of the th frame, including components, represents the Mel-frequency cepstral coefficient corresponding to the th cepstral component in the th frame, , , and are all positive integers; Based on the Mel-frequency cepstral coefficient vector of each frame, calculate the following two types of characteristic values: First, the average deviation value of the cepstrum feature, according to the formula: wherein, is the average value of the Mel frequency cepstral coefficients corresponding to the -th cepstral component in all frames, represents the average offset value of the cepstral features of the -th frame; Second, the dispersion value of the cepstrum distribution, according to the formula: In the formula, is the mean of the mel-frequency cepstral coefficients corresponding to all cepstral components in the th frame, is the numerical value of the cepstrum distribution dispersion of the th frame; Calculate the discrete index of the speech spectrum. The specific calculation formula is as follows: In the formula, is the discrete index of the voice spectrum.

6. The voice translation system based on a Bluetooth headset according to claim 5, characterized in that, For the generated voice energy fluctuation coefficient and the voice spectrum discrete index Construct a rhythm mutation intensity evaluation model, and generate a rhythm mutation index through weighted summation. The specific calculation formula is as follows: In the formula, is the rhythm mutation index, and are respectively the non-zero weight coefficients of the voice energy fluctuation coefficient and the voice spectrum dispersion index , and .

7. The voice translation system based on a Bluetooth headset according to claim 6, wherein Determine the pre-set rhythm mutation index threshold range , and after determination, compare it with the generated rhythm mutation index , and evaluate the intensity of the sentence rhythm mutation in the user's expression during the speech rate fluctuation period according to the comparison result. The specific comparison analysis is as follows: If , the intensity of the sudden change in the sentence rhythm in the user's expression during the speech rate fluctuation period is low intensity; If , the intensity of the sudden change in the sentence rhythm in the user's expression during the speech rate fluctuation period is medium intensity; If , the intensity of the sudden change in the sentence rhythm in the user's expression during the speech rate fluctuation period is high intensity.

8. The voice translation system based on a Bluetooth headset according to claim 7, characterized in that In the speech control module, according to the evaluation results, select and execute the corresponding speech processing strategies, specifically: If the evaluation result is low intensity, the speech processing strategies selected and executed include: using the default frame cutting window parameters to divide speech frames, maintaining the standard silent duration threshold and the fixed endpoint detection method, and defining speech segments in a conventional manner; If the evaluation result is medium intensity, the speech processing strategies selected and executed include: adjusting the size of the speech frame cutting window to increase the window duration, increasing the frame overlap ratio, and extending the silent duration threshold to mitigate the risk of misjudgment of segmentation caused by rhythm fluctuations; If the evaluation result is high intensity, the speech processing strategies selected and executed include: on the basis of adjusting the frame cutting window and the silent threshold, enabling the cross-segment semantic context retention mechanism and enabling the delayed endpoint judgment mode to avoid the occurrence of speech segment cutting errors and recognition timing disorders at the rhythm jump points; And perform dynamic regulation during the execution process, specifically: during the execution of the speech processing strategy, continuously monitor the change trend of the frame-level spectral energy of the speech signal and the change state of the silent duration between speech segments. If it is detected during the execution of the speech processing strategy that the rhythm mutation index continuously rises in consecutive frames and exceeds the maximum threshold corresponding to the current evaluation level, immediately switch to a higher-level speech processing strategy, update the speech frame cutting window parameters and the silent duration threshold, and enable the cross-speech segment semantic retention function; If it is detected during the processing that the rhythm mutation index continuously decreases in consecutive frames and is lower than the minimum threshold corresponding to the current strategy level, downgrade the current speech processing strategy, restore the default frame division parameters and turn off the extended processing mechanism to achieve real-time dynamic adjustment of the processing strategy.

Citation Information

Patent Citations

  • Voice speed self-adaptive recognition system

    CN114067787A

  • Voice analysis method, device and equipment and storage medium thereof

    CN118658466A