Real-time speech translation method based on multi-modal fusion and intelligent terminal
By using multimodal fusion technology to locate facial feature points and perform mathematical transformations, combined with speech recognition and facial expression features, the problem of lack of emotional information in existing real-time speech translation systems is solved, enabling emotional speech output and improving the accuracy of translation and communication effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU XUNYIDI TECH CO LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-12
AI Technical Summary
Existing real-time speech translation systems struggle to effectively integrate the speaker's visual and emotional information in cross-language communication, resulting in output speech lacking emotional warmth and failing to accurately convey the immediate emotional information of designers, thus impacting communication efficiency.
By locating facial feature points, constructing a temporal trajectory function, performing discrete-time Laplace transform and complex frequency domain feature function reconstruction, and combining speech recognition and facial expression features, multimodal fusion translation is performed to generate an emotional speech stream.
It achieves deep integration of emotional information into speech synthesis, accurately reproduces the tone and emotional fluctuations of the original speaker, improves the semantic fidelity and naturalness of the translation, and meets the needs of real-time response and accurate information transmission in cross-language communication.
Smart Images

Figure CN122024696A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech signal processing technology, and in particular to a real-time speech translation method and intelligent terminal based on multimodal fusion. Background Technology
[0002] In the field of real-time speech translation, existing technologies typically employ a cascaded processing flow combining speech recognition and machine translation, which can provide basic language conversion support for scenarios such as cross-border conferences and business negotiations. However, in communication situations involving emotional transmission and detailed expression, existing systems still have certain limitations. Specifically, most current translation systems rely primarily on audio information for text conversion and speech synthesis, rarely incorporating the speaker's visual emotional information. As a result, the output speech often fails to fully reflect the emotional state of the original speech in terms of intonation and rhythm.
[0003] For example, in remote collaborative design discussions, designers often use facial expressions and gestures to convey their design intentions and emotional inclinations when explaining creative concepts. For instance, they may use a smile to express their approval of a solution or a frown to indicate their doubts about the details. Existing translation devices can usually only translate the text content corresponding to the speech, and the synthesized translated speech is mostly flat in tone and lacks emotional fluctuations, making it difficult to convey the designer's real-time emotional information. This may prevent remote collaborators from accurately perceiving the designer's design attitude and key points, indirectly affecting communication efficiency and collaborative effectiveness. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide a real-time speech translation method and intelligent terminal based on multimodal fusion, which achieves a dynamic balance between real-time performance and translation accuracy, and meets the dual requirements of instant response and accurate information transmission in cross-language communication.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a real-time speech translation method based on multimodal fusion, the method comprising: Step 1: Locate three facial feature points: the center of the left pupil, the center of the right pupil, and the center of the lips. Based on the motion coordinates of these three facial feature points, construct a temporal trajectory function with time as the horizontal axis and normalized motion coordinates as the vertical axis. Step 2: Perform discrete-time Laplace transform on the time-series trajectory function to obtain the complex frequency domain characteristic function, and perform partial fraction expansion and inverse Z-transform on the complex frequency domain characteristic function in the region of convergence to reconstruct the time-domain emotional state vector, and calculate the emotional intensity harmonic factor. Step 3: Input the continuous clean speech data block into the pre-trained streaming speech recognition model for recognition, obtain the text segment at the current time point, and update and maintain the recognition candidate word map based on the text segment; Step 4: Detect whether a preliminary meaningful phrase or clause structure is formed in the candidate word graph. When a preliminary meaningful phrase or clause structure is detected, perform a preliminary translation of the phrase or clause structure to generate candidate translation fragments. Step 5: Semantic boundary judgment is performed on the candidate translation fragments. When the judgment reaches the complete semantic boundary, the pre-trained context translation model is triggered to generate the final optimized translation. The final optimized translation is then used to correct or replace the candidate translation fragments to form the translation text stream. Step 6: Extract the prosodic features of the source language audio stream, combine them with the emotional intensity harmonization factor and the control instructions issued by the adaptive control center, and perform speech synthesis on the translated text stream to generate and play the target language speech stream in real time.
[0006] Secondly, real-time voice translation smart terminals based on multimodal fusion include: System bus, processor, memory, power supply components, network components, display screen, speaker, camera one, camera two, microphone one and microphone two; The memory stores a computer program; the processor executes the computer program to perform the following steps: The microphones 1 and 2 are used to acquire the source language speech signal of at least one speaker; the camera 1 and camera 2 are used to acquire the facial expression and posture image information of the corresponding speaker. The source language speech signal is denoised and speech endpoints are detected, and the processed speech signal is then subjected to streaming speech recognition to obtain the source language recognized text; facial expression and posture image information is analyzed to obtain facial expression feature information to assist semantic understanding. The source language recognized text and facial expression feature information are fused into multimodal features, and machine translation is performed based on the fused features to generate target language translated text. Prosodic features are extracted from the source language speech signal and combined with facial expression features to synthesize speech from the target language translation text, generating the target language speech signal. The target language speech signal is played through the speaker, wherein the processor is also used to monitor the ambient noise intensity in real time during the speech acquisition process and dynamically adjust the noise reduction parameters and pickup directivity parameters of microphone one and microphone two.
[0007] The above-described solution of the present invention has at least the following beneficial effects: By capturing the dynamic trajectories of key facial feature points and combining mathematical transformations to quantify emotional states, emotional information is deeply integrated into the speech synthesis process. This enables the target language speech stream to accurately reproduce the original speaker's tone, emotional fluctuations, and expressive emphasis, overcoming the limitations of stiff and emotionally impersonal translated speech. This makes cross-language communication more realistic and enhances emotional resonance and information transmission integrity for both parties. Integrating core information from both facial expressions and speech modalities, facial emotional features assist semantic understanding and disambiguation, effectively compensating for information loss in complex expression scenarios using a single speech modality. This reduces semantic misinterpretations, especially in scenarios with implicit intentions or emphasis, improving the semantic fidelity and accuracy of the translation and ensuring the translation results better reflect the true meaning of the original speech. A layered processing model is employed, at the sentence level... Candidate translation fragments can be generated even when the input is incomplete, ensuring low-latency response. Simultaneously, precise translation is triggered by semantic boundary judgment, and translation quality is optimized by combining contextual information, achieving a dynamic balance between real-time performance and translation accuracy. This meets the dual needs of instant response and accurate information delivery in cross-language communication. Through adaptive control of the translation process, combined with multimodal information acquisition and processing capabilities, it can adapt to different communication scenarios, such as daily conversations and professional communication. The entire translation process can complete multimodal information acquisition, emotion quantification, layered translation, and emotional voice output without manual intervention. The adaptive control mechanism reduces operational complexity, allowing users to focus on communication itself without worrying about device operation details, thus improving ease of use and natural interaction. Attached Figure Description
[0008] Figure 1 This is a flowchart illustrating a real-time speech translation method based on multimodal fusion provided in an embodiment of the present invention.
[0009] Figure 2 This is a schematic diagram of the hardware structure of a real-time speech translation smart terminal based on multimodal fusion provided in an embodiment of the present invention.
[0010] Figure 3 This is a schematic diagram of the front main structure of a real-time voice translation smart terminal based on multimodal fusion provided in an embodiment of the present invention.
[0011] Figure 4 This is a schematic diagram of the rear main structure of a real-time voice translation smart terminal based on multimodal fusion provided in an embodiment of the present invention.
[0012] Figure 5 This is a schematic diagram of the front peripheral structure of a real-time voice translation smart terminal based on multimodal fusion provided in an embodiment of the present invention.
[0013] Figure 6This is a schematic diagram of the rear peripheral structure of a real-time voice translation smart terminal based on multimodal fusion, provided in an embodiment of the present invention.
[0014] Explanation of reference numerals in the attached drawings: 501, System bus; 502, Processor; 503, Memory; 504, Power supply component; 505, Network component; 506, Display screen; 507, Speaker; 508, Camera 1; 509, Camera 2; 510, Microphone 1; 511, Microphone 2; 512, Main module; 513, Camera module; 514, Main module cable outlet; 515, Main bracket; 516, Peripheral module; 517, Peripheral cable outlet; 518, Peripheral mounting hardware. Detailed Implementation
[0015] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0016] like Figure 1 As shown, embodiments of the present invention propose a real-time speech translation method based on multimodal fusion, the method comprising the following steps: Step 1: Locate three facial feature points: the center of the left pupil, the center of the right pupil, and the center of the lips. Based on the motion coordinates of these three facial feature points, construct a temporal trajectory function with time as the horizontal axis and normalized motion coordinates as the vertical axis. Step 2: Perform discrete-time Laplace transform on the time-series trajectory function to obtain the complex frequency domain characteristic function, and perform partial fraction expansion and inverse Z-transform on the complex frequency domain characteristic function in the region of convergence to reconstruct the time-domain emotional state vector, and calculate the emotional intensity harmonic factor. Step 3: Input the continuous clean speech data block into the pre-trained streaming speech recognition model for recognition, obtain the text segment at the current time point, and update and maintain the recognition candidate word map based on the text segment; Step 4: Detect whether a preliminary meaningful phrase or clause structure is formed in the candidate word graph. When a preliminary meaningful phrase or clause structure is detected, perform a preliminary translation of the phrase or clause structure to generate candidate translation fragments. Step 5: Semantic boundary judgment is performed on the candidate translation fragments. When the judgment reaches the complete semantic boundary, the pre-trained context translation model is triggered to generate the final optimized translation. The final optimized translation is then used to correct or replace the candidate translation fragments to form the translation text stream. Step 6: Extract the prosodic features of the source language audio stream, combine them with the emotional intensity harmonization factor and the control instructions issued by the adaptive control center, and perform speech synthesis on the translated text stream to generate and play the target language speech stream in real time.
[0017] In this embodiment of the invention, by capturing the dynamic trajectory of key facial feature points and combining mathematical transformations to quantify emotional states, emotional information is deeply integrated into the speech synthesis process. This enables the target language speech stream to accurately reproduce the tone, emotional fluctuations, and expressive emphasis of the original speaker, breaking through the limitations of stiff and emotionally impersonal translated speech. This makes cross-language communication more realistic and enhances the emotional resonance and information transmission integrity of both parties. By integrating core information from both facial expressions and speech modalities, facial emotional features assist in semantic understanding and disambiguation, effectively compensating for the information loss of a single speech modality in complex expression scenarios and reducing semantic misinterpretations. Especially in scenarios with implicit intentions or emphasis, this improves the semantic fidelity and accuracy of the translation, making the translation result more closely match the true meaning of the original speech. A layered processing model is employed. This system generates candidate translation fragments even when the sentence is not fully input, ensuring low-latency response. Simultaneously, it triggers accurate translation through semantic boundary judgment and optimizes translation quality by combining contextual information, achieving a dynamic balance between real-time performance and translation accuracy. This meets the dual needs of instant response and accurate information delivery in cross-language communication. Through adaptive control of the translation process, combined with multimodal information acquisition and processing capabilities, it can adapt to different communication scenarios, such as daily conversations and professional communication. The entire translation process can complete multimodal information acquisition, emotion quantification, layered translation, and emotional voice output without manual intervention. The adaptive control mechanism reduces operational complexity, allowing users to focus on communication itself without worrying about device operation details, thus improving ease of use and natural interaction.
[0018] In another preferred embodiment of the present invention, prior to step 1: Step 001 involves real-time capture or reception of the source language audio stream and the speaker's video stream. Specifically, this includes: continuously acquiring source language sound signals from the surrounding environment using the array microphone in the multimodal acquisition module; converting the sound wave signals into electrical signals at a fixed sampling frequency of 16kHz using the array microphone; then converting the electrical signals into 16-bit deep digital audio data using an analog-to-digital converter; each segment of digital audio data is accompanied by a unique timestamp, accurate to the millisecond level, used to mark the time node of data acquisition; and continuously capturing the speaker's image information at a fixed frame rate of 30 frames per second using the low-power CMOS camera in the multimodal acquisition module, converting the optical image... The signal is converted into digital video data, with the image resolution set to 1280×720. A corresponding timestamp is added to each frame of digital video data, and the timestamp is kept synchronized with the audio stream to ensure that the audio stream and video stream are time-aligned. If the source language audio stream and the speaker's video stream come from external smart devices such as mobile phones and computers, this type of data is received through the network communication module of the smart terminal (wireless Wi-Fi / BLE dual-mode or wired network). During the reception process, the data integrity is checked by a cyclic redundancy check algorithm. After confirming that there is no packet loss or damage, the data is temporarily stored in the cache of the core control module, and the storage format is consistent with the locally acquired data.
[0019] Step 002 involves real-time noise reduction and speech endpoint detection on the captured or received source language audio stream to generate continuous clean speech data blocks, and detecting face regions from the captured or received speaker video stream. Specifically, the core control module transmits the source language audio stream and speaker video stream obtained in step 001 to a dedicated preprocessing unit, simultaneously performing audio preprocessing (real-time noise reduction and speech endpoint detection) and video face region detection operations. This ultimately generates continuous clean speech data blocks and locates face regions in the video stream. Specifically, the preprocessing unit first extracts the output signal from the speaker in the smart terminal's voice output module, delaying this signal by 10 milliseconds as a reference signal. An adaptive filtering algorithm is used, targeting the error between the original audio signal and the filtered reference signal. The filtering coefficients are iteratively adjusted using a least mean square algorithm. In each iteration, the error value is multiplied by a step size coefficient of 0.01 to update the filtering coefficients, ensuring that the filtered reference signal is consistent with the echo signal in the audio stream. Then, the original audio signal collected by the array microphone is subtracted... The filtered reference signal achieves precise echo cancellation, retaining only the speaker's voice signal. The echo-cancelled audio signal is then processed by frame segmentation, dividing the continuous audio stream into 20-millisecond audio frames with a 5-millisecond overlap between adjacent frames. A Fast Fourier Transform (FFT) is used to calculate the spectral information of each audio frame. First, the audio frame is multiplied by a Hanning window with an amplitude varying between 0 and 1. The smooth attenuation at both ends of the window function reduces spectral leakage. Then, the FFT converts the time-domain audio frames into frequency-domain data, obtaining the signal amplitude at each frequency point. A preset ambient noise amplitude threshold of -45 dB is used. The signal amplitude at each frequency point is compared to this threshold. Frequency components with amplitudes below -45 dB are identified as ambient noise and their amplitudes are multiplied by a preset attenuation coefficient of 0.3 for suppression. Frequency components with amplitudes above -45 dB are considered valid speech signals, and their original amplitudes are preserved. After filtering the ambient noise, an Inverse Fast Fourier Transform (IFFT) converts the frequency-domain data back to the time-domain audio signal.
[0020] The short-time energy and zero-crossing rate of each audio frame are calculated. The short-time energy is calculated by summing the squared amplitude values of each sample point within the audio frame and dividing by the number of sample points in that frame. The zero-crossing rate is calculated by counting the number of times the amplitude sign changes between adjacent sample points within the audio frame and dividing by the number of sample points in that frame. A short-time energy threshold of 0.02 and a reasonable range for the zero-crossing rate of 0.05 to 0.3 are set. When the short-time energy of three consecutive frames is higher than the energy threshold of 0.02 and the zero-crossing rate is within the reasonable range of 0.05 to 0.3, it is determined as the start of a speech segment. When the short-time energy of five consecutive frames is lower than the energy threshold of 0.02 and the zero-crossing rate exceeds the reasonable range of 0.05 to 0.3, it is determined as the end of a speech segment. Based on the start and end points, the processed audio signal is segmented into continuous clean speech data blocks.
[0021] The process of detecting the face region in the speaker's video stream is as follows: The video stream is sequentially broken down into single-frame images. Each frame is first converted to grayscale. The weighting coefficients 0.299, 0.587, and 0.114 are determined based on the human eye's sensitivity to red, green, and blue colors. These coefficients are used to calculate the grayscale value of each pixel in the image. The calculation logic is: Grayscale value = 0.299 × red channel pixel value + 0.587 × green channel pixel value + 0.114 × blue channel pixel value. This converts the color image to a grayscale image. Subsequently, histogram equalization is performed on the grayscale image. First, the number of pixels at each grayscale level is counted, and then the cumulative distribution function is calculated using the formula: [Formula omitted]. ,in Represents the gray level (value range is 0 to 255). For the first The number of pixels corresponding to each gray level The total number of pixels in the entire grayscale image is used as the starting point. Then, the calculated cumulative distribution function is mapped proportionally to a grayscale range of 0 to 255 to adjust the image's grayscale distribution, improve image contrast, and enhance the distinction between facial features and the background. The preprocessed grayscale image is scanned region by region, with a scanning window size of 20×20 pixels. The window moves in a step of 5 pixels, and each moved window is considered a candidate region. Contour and texture features are extracted from each candidate region. Contour features are obtained by using an edge detection algorithm to acquire the arrangement information of pixels at the region boundaries, and texture features are obtained by calculating the mean and variance of the grayscale values of the pixels within the region. The extracted candidate region features are compared with a preset facial feature template for similarity calculation. The calculation method is as follows: first, the two sets of features are converted into feature vectors, the dot product of the two feature vectors is calculated, and then divided by the product of the magnitudes of the two feature vectors to obtain the cosine similarity. The preset similarity threshold is 0.7. This threshold value is selected by referencing the real-time speech translation intelligent terminal. For terminal applications (such as border inspection and video conferencing), there is a dual requirement for detection accuracy and real-time performance. A threshold of 0.7 is within the industry-common reasonable threshold range of 0.6 to 0.8, which can balance the detection rate of face regions and the false detection rate of non-face regions. Secondly, combined with the computing power of terminal hardware (such as the RKNNNPU processing capability of the AI computing module), this threshold can avoid a large number of non-face candidate regions entering subsequent processing due to an excessively low threshold, resulting in wasted computing power and increased latency. At the same time, it can also prevent the omission of some face regions with slight facial pose deviations or occlusions due to an excessively high threshold, ensuring the continuity of emotional feature extraction. When the cosine similarity between the candidate region feature and the face feature template is higher than the threshold of 0.7, the candidate region is determined to be a face region, and the pixel coordinates of the upper left and lower right corners of the region are recorded to clarify the range of the face region. For candidate regions with a similarity lower than the threshold of 0.7, they are determined to be non-face regions and filtered out, completing the face region detection of a single frame image. Subsequent frames are executed in the same manner.
[0022] In this embodiment, audio preprocessing effectively filters out irrelevant signals such as speaker echoes and environmental noise through echo cancellation and environmental noise suppression. Combined with speech endpoint detection, it accurately segments speech segments, reducing the negative impact of interference factors on recognition performance. Face region detection in the video stream improves image quality through preprocessing such as grayscale conversion and histogram equalization. Then, feature matching accurately locates the face range, eliminating interference factors such as background and irrelevant objects, ensuring the effectiveness of emotional feature extraction. The synchronous capture and preprocessing of audio and video streams ensure the consistency of the two modal data in the time dimension, avoiding fusion deviations caused by data temporal misalignment. From data capture to the generation of clean speech data blocks and the location of face regions, there is no significant delay throughout the process, meeting the core requirement of real-time speech translation for immediate response.
[0023] In a preferred embodiment of the present invention, step 1 includes: Step 100: Based on the face region, locate three facial feature points: the center of the left pupil, the center of the right pupil, and the center of the lips. Specifically, using the face region detected in step 002 as the target area, first perform grayscale enhancement processing on the image of this region. By adjusting the image's grayscale contrast, enhance the texture differences between the eyes, lips, and surrounding skin, providing a clear image foundation for feature point recognition. Then, perform precise localization of the three feature points—the center of the left pupil, the center of the right pupil, and the center of the lips—region by region, as follows: For the localization of the left and right eye regions, within the face region... The candidate eye region is defined (based on facial proportions, the upper and lower boundaries of the candidate eye region are 1 / 5 to 2 / 5 of the face region height, and the left and right boundaries are 1 / 4 to 1 / 2 and 1 / 2 to 3 / 4 of the face region width, respectively). An edge detection algorithm is used to extract the eye contour, and then the Hough circle detection algorithm is used to identify the pupil contour. During Hough circle detection, a reasonable pupil radius range is first set based on the size of the candidate eye region, combined with common facial eye size features; this range specifically ranges from 5 to 20 pixels. Then, a Hough spatial accumulator is constructed, and the circle equation is defined as follows: Here, (a, b) are the coordinates of the center of the circle to be determined, (x, y) are the coordinates of any pixel on the eye contour, and r is any radius value within a preset radius range. For each pixel on the eye contour, all possible center positions (a, b) are traversed within a preset radius range of 5 to 20 pixels. If the coordinates (x, y) of the pixel satisfy the equation of the circle with the corresponding radius r, then the center position (a, b) is incremented by 1 vote. After the traversal, the total number of pixels in the eye contour is counted, and this number is multiplied by 40% to obtain the voting threshold (for example, if the total number of pixels in the eye contour is 100, then the voting threshold is 40). Center positions with a vote count higher than this threshold are selected, and the position with the highest vote count is the center of the pupil contour. Its corresponding x-axis and y-axis coordinates are the pixel coordinates of the center of the left and right pupils. After completing the eye feature point localization, Continue with the lip center localization operation. Define a candidate lip area within the face region (based on facial proportions, the upper and lower boundaries of the candidate lip area are 3 / 5 to 4 / 5 of the face region height, and the left and right boundaries are 1 / 3 to 2 / 3 of the face region width). Highlight lip color features through color space conversion, then use an edge detection algorithm to extract the lip contour. Calculate the geometric center coordinates of the lip contour. The calculation method is as follows: first, count the x-axis coordinates of all pixels on the lip contour, add all x-axis coordinates together to obtain the sum, and then divide by the total number of contour pixels to obtain the x-axis center coordinates. Similarly, count the y-axis coordinates of all pixels on the lip contour, add all y-axis coordinates together to obtain the sum, and then divide by the total number of contour pixels to obtain the y-axis center coordinates. The coordinate pair consisting of the x-axis center coordinates and the y-axis center coordinates is the pixel coordinate of the lip center.
[0024] Step 101: Based on the pixel coordinates of three facial feature points in consecutive video frames, calculate the motion coordinates of each feature point and normalize them to obtain normalized motion coordinates. Specifically, this includes: extracting the pixel coordinates (x-axis coordinates and y-axis coordinates) of each feature point in each frame based on the temporal order of the consecutive video frames; for each feature point, subtracting the x-axis coordinate of the previous frame from the x-axis coordinate of the current frame to obtain the displacement in the x-axis direction; subtracting the y-axis coordinate of the previous frame from the y-axis coordinate of the current frame to obtain the displacement in the y-axis direction; and combining the x-axis and y-axis displacements to form the motion coordinates of the feature point in the current frame relative to the previous frame. Motion coordinates; Calculate the maximum and minimum x-axis displacement and y-axis displacement of all feature points within a consecutive preset number of frames (e.g., 10 frames); for each feature point's motion coordinates, normalize them using the following formulas: Normalized x-axis coordinate = (Current x-axis displacement - Minimum x-axis displacement) / (Maximum x-axis displacement - Minimum x-axis displacement); Normalized y-axis coordinate = (Current y-axis displacement - Minimum y-axis displacement) / (Maximum y-axis displacement - Minimum y-axis displacement); This calculation maps the motion coordinates of all feature points to the range of 0 to 1, resulting in normalized motion coordinates.
[0025] Step 102: Based on the normalized motion coordinates and their corresponding timestamps, construct a temporal trajectory function with time as the horizontal axis and normalized motion coordinates as the vertical axis. Specifically, this includes: extracting the timestamp corresponding to each video frame (consistent with the timestamp of the video stream in Step 001, accurate to the millisecond level); mapping the normalized motion coordinates of each feature point to the timestamp of its corresponding video frame, ensuring that the normalized motion coordinates of each feature point can be associated with a unique time point; after completing the one-to-one mapping between timestamps and normalized motion coordinates, construct two independent temporal trajectory functions, using the time value corresponding to the timestamp as the horizontal axis (independent variable) and the normalized x-axis and y-axis coordinates of each feature point as the vertical axes (dependent variables), respectively reflecting the dynamic changes of the feature points in the x-axis and y-axis directions. The formula for the x-axis temporal trajectory function is as follows: The formula for the time-series trajectory function in the y-axis direction is: In the formula Representing the The timestamp (in milliseconds) corresponding to each video frame. The normalized x-axis coordinates of the feature points at this time point. The normalized y-axis coordinates of the feature points at that time point; the function is represented by a discrete function, with each time point... Corresponding to a unique normalized motion coordinate value This forms complete time-series trajectory data.
[0026] In this embodiment, the localization of key facial feature points focuses on the core area. By combining facial proportion features with image enhancement and contour extraction technologies, the accuracy of feature point localization is ensured, and interference from irrelevant areas on feature extraction is reduced. Motion coordinate calculation accurately reflects the dynamic changes of feature points through displacement statistics. Normalization processing eliminates the scale differences in displacement amounts under different feature points and different scenarios, enabling the data to have a unified analysis standard. The temporal trajectory function associates discrete normalized motion coordinates with timestamps, intuitively presenting the changing patterns of facial feature points over time and completely preserving the dynamic information of facial expressions. The entire process is based on the detected facial region, avoiding indiscriminate processing of the entire image, reducing the amount of data processing, matching the hardware computing power configuration of smart terminals, ensuring the real-time performance of the processing flow, and adapting to the needs of real-time speech translation scenarios.
[0027] In a preferred embodiment of the present invention, step 2 includes: Step 200: Perform a discrete-time Laplace transform operation on the time-series trajectory function to obtain the complex frequency domain characteristic function of the complex frequency variable. Specifically, this includes: defining the complete operational formula for the discrete-time Laplace transform in the x-axis direction as follows: ,in for axial direction complex frequency domain characteristic function, For complex frequency variables, for The time-series trajectory function of the axis is in the first... discrete time points The corresponding normalized motion coordinate values, The natural constant is defined; the complete formula for the discrete-time Laplace transform along the y-axis is: ,That for axial direction complex frequency domain characteristic function, for The time-series trajectory function of the axis is in the first... discrete time points The corresponding normalized motion coordinate values are defined, with other variables defined in the same way as the x-axis. For the time-series trajectory functions of the x-axis and y-axis, the time variables are replaced with complex frequency variables s, and the normalized motion coordinate values corresponding to all discrete time points in each function are iterated. For the normalized motion coordinate values at each time point on the x-axis... According to the formula After completing the calculations, the results at all time points are summed sequentially to obtain the complex frequency domain characteristic function corresponding to the x-axis time-series trajectory function. Similarly, according to the formula Perform calculations at each time point along the y-axis and accumulate the results to obtain the complex frequency domain characteristic function along the y-axis. Finally, the complex frequency domain characteristic functions in the x-axis and y-axis directions are output.
[0028] Step 201: Analyze the expression of the complex frequency domain characteristic function, solve for the range of complex frequency variables that make the complex frequency domain characteristic function sequence converge, and obtain the region of convergence of the complex frequency domain characteristic function. Specifically, this includes: first, locating all poles in the expression of each complex frequency domain characteristic function, i.e., the x-axis complex frequency domain characteristic function. The poles are the complex frequency variable s values whose denominators are equal to zero, satisfying the formula (in for (denominator polynomial); y-axis complex frequency domain characteristic function The poles are the complex frequency variable s values whose denominators are equal to zero, satisfying the formula (in for The denominator polynomial); combined with the time-domain characteristics of the time-series trajectory function, it is determined that the signal types corresponding to the x-axis and y-axis are both right-hand signals (non-zero values only at the current and subsequent time points), and their corresponding regions of convergence are respectively. The region of convergence of the complex frequency domain characteristic function of the x-axis is the set of all complex frequency variables in the complex plane whose real parts are greater than the real parts of their maximum poles, i.e. (in Represents complex frequency variable The real part; for The region of convergence of the y-axis complex frequency domain characteristic function is the set of all complex frequency variables in the complex plane whose real parts are greater than the real parts of its largest pole. (in for (The pole with the largest real part among all poles); verify whether this range of values satisfies the condition that the sequence of characteristic functions in the complex frequency domain is absolutely integrable, i.e., the x-axis satisfies... The y-axis satisfies (The sum of the absolute values of all terms is a finite value), and finally the convergence regions of the two complex frequency domain characteristic functions are determined and recorded.
[0029] Step 202: Within the region of convergence, decompose the complex frequency domain feature function into the sum of multiple first-order fractions, completing partial fraction expansion; perform inverse Z-transform operation on each expanded first-order fraction, and superimpose the inverse transform results of each fraction in the time domain to reconstruct the time-domain sentiment state vector. Specifically, this includes: within the region of convergence determined in step 201, processing each complex frequency domain feature function and reconstructing the time-domain sentiment state vector through inverse Z-transform, as follows: Organize the complex frequency domain feature function into... (in For the numerator polynomial, (where the denominator is a polynomial), factoring the denominator polynomial yields... (n represents the total number of first-order factors after factoring the denominator polynomial, i.e., the total number of poles of the characteristic function in the complex frequency domain,) For the first (each having an extreme point), and then the characteristic function is decomposed into the sum of multiple first-order fractions, i.e. ,in For the first The residual values corresponding to each pole are obtained through the formula. The calculation yields (by multiplying the characteristic function by the corresponding denominator factor and then substituting it into the extreme point value to solve for the limit); for each first-order fraction Perform the inverse Z-transform operation separately, based on the basic formula of the inverse Z-transform. (in is the inverse Z-transform operator, used to convert complex frequency domain (Z-domain) signals into discrete time domain signals; z is the complex variable of the Z-transform, which is the core variable in Z-domain analysis; The constant term of the complex variable z in the denominator of the first-order fraction corresponds to the pole of the complex frequency domain characteristic function in this step. ; The numerator of the first-order fraction corresponds to the pole in this step. Corresponding retention value , For unit step sequence, (For discrete-time indexing), convert each first-order fraction in the complex frequency domain into a discrete time-domain sequence; then, sum the inverse Z-transform results of all first-order fractions after decomposing the same feature function in the complex frequency domain at the corresponding time points to obtain the time-domain sequence corresponding to that feature function. Combine the time-domain sequences along the x-axis and y-axis into The complete temporal sentiment state vector is obtained by reconstruction, where This is a temporal emotion state vector, used to comprehensively represent the emotional dynamics reflected by the movement of feature points along the x and y axes. It is the vector transpose symbol.
[0030] Step 203: Perform numerical differentiation on the time-series trajectory function, calculate its first and second derivative values at each sampling time point, detect the location where the sign of the second derivative value changes, determine the location as the inflection point of the time-series trajectory function, and extract the absolute value of the first derivative value at each inflection point as the instantaneous change intensity. Specifically, this includes: calling the two time-series trajectory functions constructed in step 102 again, calculating the derivative through numerical differentiation and detecting inflection points, and extracting the instantaneous change intensity, as follows: for each time-series trajectory function... Select two adjacent sampling time points in chronological order. and Through formula Calculate the first derivative value at the next time point (by dividing the difference in normalized motion coordinates between adjacent time points by the time interval), iterate through all sampling time points, and complete the calculation of the first derivative value for the entire sequence; adopt the same adjacent time point traversal method as the first derivative, and use the formula... Calculate the second derivative value at the next time point (by dividing the difference of the first derivative values at adjacent time points by the time interval), and complete the calculation of the second derivative values for the entire sequence; perform sign determination on the second derivative value at each sampling time point one by one, and compare it with the current time point. Compared to the previous time point The sign of the second derivative, if it satisfies If the sign changes, the current time point is determined as an inflection point of the time-series trajectory function, and the first derivative value corresponding to each inflection point is extracted. ( (The inflection point is the time point), and the absolute value of this value is used to obtain the instantaneous change intensity, i.e. .
[0031] Step 204: Based on the number of all detected inflection points and the average value of instantaneous change intensity, an inflection point modulation factor is calculated. Based on the denominator pole values, corresponding numerator residue values, and inflection point modulation factor of each first-order fraction after partial fraction expansion, an emotion intensity harmonization factor is calculated through weighted fusion. Specifically, this includes: calculating the inflection point modulation factor and the emotion intensity harmonization factor sequentially based on the partial fraction expansion results of step 202 and the inflection point detection data of step 203. The role of the emotion intensity harmonization factor is to quantify the degree of emotional fluctuation of the speaker, providing emotional parameter support for subsequent speech synthesis, so that the synthesized speech can accurately reproduce the tone and emotion of the original speaker. The role of weighted fusion is to highlight the importance of the emotional features corresponding to different first-order fractions, making the emotion intensity assessment more in line with the actual expression situation. Specifically, let the total number of all inflection points detected in step 203 be M, and the sum of the instantaneous change intensity of all inflection points be... Where u is a discrete index used to distinguish different inflection points, and its value ranges from 1 to M. For the first The instantaneous change intensity at each inflection point, and the total number of sampling points for the time-series trajectory function are... First, calculate the average value of the instantaneous change intensity, that is... Then, the inflection point modulation factor is calculated, i.e. Extract the pole values of the denominator of each first-order fraction after partial fraction expansion in step 202. and the corresponding molecular retention values Calculate the absolute value of each pole. As the weighting coefficients of this first-order fraction, through the formula Calculate the harmonic components of each first-order fraction (where (This is the inflection point modulation factor), and then through the formula (in The total number of first-order fractions after partial fraction expansion in step 202 (i.e., the total number of poles of the characteristic function in the complex frequency domain) is added together with the harmonic components of all first-order fractions to obtain the final emotional intensity harmonic factor.
[0032] In this embodiment, the discrete-time Laplace transform converts the dynamic trajectory of feature points in the time domain to the complex frequency domain, highlighting the frequency characteristics of trajectory changes and providing richer dimensional support for the extraction of emotion-related features. Solving the convergence domain ensures the effectiveness of the feature function in the complex frequency domain and the accuracy of subsequent calculations, avoiding distortion of analysis results due to improper values of dependent variables. The combination of partial fraction expansion and inverse Z-transform achieves accurate reconstruction of complex frequency domain features into the time domain, preserving dynamic details related to emotion states and improving the reliability of the time-domain emotion state vector. Numerical differentiation and inflection point detection can capture abrupt changes in the trajectory of feature points, and the intensity of instantaneous changes intuitively reflects the intensity of emotional expression, providing a key basis for the quantification of emotion intensity. The calculation of the inflection point modulation factor and the emotion intensity harmonic factor integrates the frequency and abrupt change characteristics of the trajectory, realizing the comprehensive utilization of multi-dimensional features, making the assessment of emotion intensity more comprehensive and in line with actual expression.
[0033] In a preferred embodiment of the present invention, step 3 includes: Step 300: Continuous clean speech data blocks are input into a pre-trained streaming speech recognition model in chronological order. The streaming speech recognition model processes the continuous clean speech data blocks and outputs one or more candidate words at the current time point and the probability corresponding to each candidate word. Specifically, this includes: building a streaming speech recognition model using a Transformer-based encoder architecture, introducing a chunk-wise block processing mechanism to divide long speech sequences into fixed-length continuous speech blocks for parallel processing, while preserving the contextual dependency information of adjacent blocks; the model includes an input layer, a convolutional feature extraction layer, a Transformer encoder layer, a fully connected layer, and an output layer. The convolutional feature extraction layer is used to extract Mel-spectral features of the speech, and the Transformer encoder layer captures the temporal correlation of the speech sequence through a multi-head attention mechanism. The fully connected layer, in conjunction with the Softmax activation function, outputs the probability distribution of word terms. A multilingual speech-text alignment dataset is selected as the training data, covering diverse scenarios such as daily conversations, business negotiations, and border inspections. The data includes the speech of speakers with different accents, genders, and ages, and background noise data of varying intensities is added to improve the model's anti-interference ability. The training data is divided into training, validation, and test sets in an 8:1:1 ratio. The cross-entropy loss function is used as the optimization objective, and the stochastic gradient descent algorithm is used to iteratively train the model. In each iteration, the speech data is converted into Mel-spectral features and input to the model. The model outputs a predicted word term sequence. By calculating the cross-entropy loss value between the predicted sequence and the real text sequence, the model parameters are updated through backpropagation. The training stops when the speech recognition accuracy on the validation set no longer improves, and the final pre-trained model parameters are saved.
[0034] The continuous clean speech data blocks generated in step 1 are sequentially input into the trained streaming speech recognition model. For each input speech data block, the model first extracts Mel-spectral features, then captures the temporal correlation of the features through the encoder layer. The processed feature vector is then input into a fully connected layer. The fully connected layer maps the feature vector to a score vector in the word dictionary dimension through a linear transformation. This score vector is then input into a Softmax activation function to calculate the probability of each word. A fixed probability threshold of 0.3 is set. This threshold value is determined based on the model training and validation results and the actual scenario requirements. On the one hand, a threshold of 0.3 can effectively filter out words with high confidence, preventing low-probability (e.g., less than 0.3) suspicious words from entering subsequent processing, thus reducing the impact of misidentification on the translation results. On the one hand, this threshold can balance the quantity and quality of candidate words. It will not result in too few or even zero candidate words due to an excessively high threshold (such as 0.5 and above), affecting the integrity of the translation and the ability to resolve ambiguities. On the other hand, it will not introduce a large number of low-confidence words due to an excessively low threshold (such as 0.1 and below), increasing the maintenance complexity and computing power consumption of the candidate word map. It is suitable for the real-time processing needs of diverse scenarios such as daily conversations, business negotiations, and border inspections. All words with a probability value greater than the threshold are selected as candidate words at the current time point. If the number of candidate words after selection is zero, the word with the highest probability value is selected as the candidate word. At the same time, the probability value corresponding to each candidate word is output. This probability value is the confidence of the current word among all possible words. The sum of the probabilities of all candidate words is 1.
[0035] Step 301: At the initial moment, starting from the first candidate word, initialize a candidate word graph containing nodes and directed edges. Nodes represent candidate words, and directed edges represent the transition relationships between candidate words, weighted by the probabilities corresponding to the candidate words. Specifically, at the initial moment, based on the first candidate word output in step 300, initialize a candidate word graph containing nodes and directed edges. Each node in the candidate word graph stores the core information of a single candidate word, including word content and the probability corresponding to that word. Using the first candidate word output in step 300 as the initial node, create a corresponding node instance and store the content and probability of that candidate word in the node attributes. The initial candidate word graph only contains the aforementioned initial node. Since there are no subsequent candidate words, no directed edges are established across nodes. At this time, the path in the word graph is only a single path starting from the starting point (initial node), and the cumulative probability value of this path is the probability of the candidate word corresponding to the initial node, thus completing the initial construction of the candidate word graph.
[0036] Step 302: At each subsequent time point, new candidate words are added as new nodes to the candidate word graph. Using the probability corresponding to the candidate word as weight, directed edges are established from the optional nodes of the previous time point to the current new node. Simultaneously, the cumulative probability values of each path from the starting point to the current new node are updated, completing the incremental update and maintenance of the candidate word graph. Specifically, this includes: obtaining all new candidate words output in step 300 at the current time point; creating a corresponding node instance for each new candidate word; storing the word content and corresponding probability in the node attributes; and adding all these new nodes to the existing candidate word graph. For each new node, all optional nodes generated from the previous time point (i.e., nodes added at the previous time point) are processed. Starting from the candidate word nodes in the graph, directed edges are established pointing to the current new node. The weight of each directed edge is set as the probability of the candidate word corresponding to the current new node, which is used to represent the possibility of transitioning from the previous word to the current word. For each path from the starting point through the optional nodes at the previous time point to the current new node, its cumulative probability value is calculated by multiplying the cumulative probability value of the path corresponding to the optional nodes at the previous time point by the weight of the directed edge pointing to the current new node (i.e., the probability of the current new candidate word). All new nodes and their corresponding associated paths are traversed to update the cumulative probability value of each path, thereby realizing the incremental update and maintenance of the candidate word graph and ensuring that the word graph reflects all possible speech recognition paths and their corresponding confidence levels in real time.
[0037] In this embodiment, the pre-trained streaming speech recognition model, after being trained on multi-scene data, possesses strong scene adaptability and anti-interference capabilities. It can accurately extract speech features and output reliable candidate words and probabilities. The streaming processing mechanism is compatible with the time-sequential input method, meeting the low-latency requirements of real-time speech translation and responding instantly to the recognition requests of each segment of speech data. The recognition candidate word graph retains all possible recognition paths through the structure of nodes and directed edges. The incremental update method ensures the continuity of recognition and reflects the confidence of the path through the cumulative probability, thereby improving the accuracy of speech recognition. The word graph maintenance process retains multi-path information, which can effectively address the ambiguity problem in speech recognition and provide rich candidate evidence for context-aware incremental translation, helping to improve the completeness and accuracy of translation.
[0038] In a preferred embodiment of the present invention, step 4 includes: Step 400: After updating the candidate word graph, starting from the new node corresponding to the current time point, backtrack along the directed edges to obtain several candidate paths ending at that node. Specifically, this includes: identifying each new node added in step 302 at the current time point as the backtracking endpoint; each new node corresponds to a candidate word element output at the current time point and is associated with multiple potential recognition paths extending from the initial node; starting from each new node, tracing back the preceding nodes layer by layer along the reverse direction of the directed edges, preserving the integrity of the path during the tracing process, i.e., each path starts from the initial node, passes through the optional nodes at each intermediate time point, and finally terminates at the current new node; extracting all complete paths obtained from the backtracking, sorting them according to the cumulative probability value of the path (the cumulative probability value of the path is the product of the weights of all directed edges on the path, i.e., the product of the probabilities of the candidate words corresponding to each node), and selecting the top 3 to 5 paths with the highest cumulative probability values as candidate paths, which ensures the confidence of the path and avoids increasing the subsequent processing overhead due to too many paths. The number of paths selected can be dynamically adjusted according to the real-time computing power.
[0039] Step 401: Perform part-of-speech tagging and shallow dependency analysis on the lexical sequence corresponding to each candidate path. When the analysis results determine that the lexical sequence of a candidate path conforms to the preset phrase or clause grammatical template, it is determined that a preliminary meaningful phrase or clause structure has been formed. Specifically, this includes: based on the general terminology database and professional terminology database stored locally on the smart terminal, and combined with the grammatical rules, word form change features, and common collocation habits of the source language, tagging each lexical sequence of each candidate path. The tagging categories include basic parts of speech such as nouns, verbs, adjectives, prepositions, and auxiliary words, as well as professional domain-specific parts of speech (such as the specific noun "document type" in border inspection scenarios and the part of speech "contract terminology" in business negotiation scenarios), improving the accuracy of part-of-speech tagging in special scenarios; based on the tagged part-of-speech results, lightweight dependency syntax is used. The analysis algorithm analyzes the core dependency relationships between lexical units, including subject-predicate, verb-object, modifier-head, and prepositional phrase relationships, to clarify the grammatical structure framework of the lexical unit sequence. For example, it identifies verb + noun verb-object structures, adjective + noun modifier-head structures, and noun + verb + object subject-predicate-object structures. It pre-sets grammatical templates for phrases and clauses suitable for diverse scenarios, covering commonly used phrase templates in daily conversations (such as greeting templates and request templates), core sentence templates in business negotiations (such as quotation templates and cooperation intention templates), and special sentence templates in border inspections (such as document verification templates and information inquiry templates). The grammatical structure framework of each path is compared with the pre-set templates. If the structure matches completely or the core components match (such as containing subject-predicate-object core components), it is determined that the lexical unit sequence of that path forms a preliminary meaningful phrase or clause structure.
[0040] Step 402 involves extracting the source language lexical sequences corresponding to phrase or clause structures, converting them into coherent source language text, and performing a preliminary real-time translation of the converted source language text to generate candidate translation fragments. Specifically, this includes: extracting the lexical sequences corresponding to candidate paths forming phrase or clause structures, strictly arranging the lexical sequences according to their temporal order within the path; subsequently refining the word order according to the grammatical conventions of the source language. For example, Chinese follows the basic word order of attributive + subject + adverbial + predicate + attributive + object + complement, while English follows the core word order of subject + predicate + object. If the word order does not conform to the basic word order, it is adjusted to the correct position according to grammatical rules. Then, based on the semantic relationship between words, it is determined whether to add connecting elements. If there is a lack of prepositions between nouns and verbs, a particle is needed to assist in the expression between verbs and objects, or a conjunction is needed to connect phrases, then the corresponding particles, prepositions, or conjunctions are added. During the text conversion process, the locally stored corpus is searched to match common expressions with similar semantics to the current word combination. The conversion results are optimized by referring to their sentence structure and word usage habits to ensure that the final source language text is coherent, natural, and conforms to the norms of language expression.
[0041] The converted source language text is input into the translation processing unit, which is built based on multilingual semantic correspondence rules, sentence transformation norms, and domain adaptation logic, supporting rapid translation across scenarios. First, the source language text undergoes refined preprocessing, splitting the text according to grammatical structures such as subject-predicate, verb-object, modifier-head, and prepositional phrases, clearly distinguishing phrase boundaries, and marking core semantic carriers of sentences (such as nouns acting as subjects and verbs acting as predicates) as core words, and marking limiting or supplementary components (such as adjectives modifying nouns and adverbs modifying verbs) as modifiers. Simultaneously, by comparing with general and specialized terminology databases, the specialized terminology in the text is accurately marked. Industry-specific terms (such as document verification in border inspection scenarios and cooperation intentions in business negotiation scenarios) and fixed collocations (such as reaching a consensus and submitting materials); then, the general terminology database stored locally on the smart terminal and the synchronized professional terminology database are accessed, combined with the current dialogue domain information obtained by the system status monitoring module, to prioritize the screening of professional terminology entries in the corresponding domain. First, completely identical terms are matched to achieve precise correspondence. If no completely identical terms are found, semantically equivalent terms are matched to ensure no deviation between the source language terms and the target language terms; then, based on the explicit sentence structure correspondences in the multilingual parallel corpus, semantic matching rules are used to detect... The target language sentence template, which is structurally identical to the source language text, is used in conjunction with previous part-of-speech tagging results to map the core words and modifiers of the source language to designated positions such as subject, predicate, object, and attributive in the target language sentence structure. For example, the adjective + noun modifier structure in the source language corresponds to the standard adjective + noun or noun + adjective structure in the target language (adjusted according to target language conventions), completing the basic sentence structure conversion. Then, the target language text is optimized through semantic logic supplementation rules. First, the semantic relationships in the text (cause and effect, contrast, progression, etc.) are analyzed, and if conjunctions are lacking, corresponding target language conjunctions are added. Connecting words are then adjusted according to the word order of modifiers in the target language (e.g., if a prepositional modifier in Chinese needs to be placed after the verb in English), ensuring grammatical coherence. Finally, based on four dimensions—accuracy of core word translation, fit of sentence structure, precision of terminology correspondence, and completeness of semantic logic—the matching degree between each candidate translation and the source language semantics is comprehensively calculated, generating 2 to 3 candidate translation fragments that conform to the expression habits of the target language. These fragments are then sorted from high to low according to their matching degree scores. The translation process maintains streaming characteristics, and the sorted candidate translation fragments are output immediately after the text conversion is completed, adapting to the low latency requirements of real-time speech translation.
[0042] In this embodiment, the candidate path backtracking and filtering mechanism filters high-confidence paths based on cumulative probability values, reducing the interference of low-quality paths on subsequent processing. This provides a reliable lexical sequence foundation for grammatical analysis and translation, improving the accuracy of the initial translation. Combined with multi-scenario preset grammatical templates, meaningful structures are accurately identified through part-of-speech tagging and shallow dependency analysis, avoiding invalid translation of lexical sequences that have not formed complete semantics, thus balancing the real-time performance and effectiveness of translation. The initial translation is performed in conjunction with scenario information, generating candidate translation fragments in advance, reducing end-to-end translation latency, and adapting to scenarios with high real-time requirements such as border inspection and business negotiations. The generated candidate translation fragments provide a reference for subsequent accurate translation, ensuring real-time output and laying the foundation for improving the final translation quality.
[0043] In a preferred embodiment of the present invention, step 5 includes: Step 500: Based on the path with the highest cumulative probability from the starting point to the current node in the candidate word graph, obtain the word sequence corresponding to the path. Specifically, this includes: traversing all complete paths extending from the initial node to the current node in the candidate word graph. The cumulative probability value of each path is the product of the weights of all directed edges on the path. That is, multiply the candidate word probabilities corresponding to each node in the path to obtain the overall confidence of the path (for example, if a path contains three nodes with word probabilities of 0.8, 0.7, and 0.9 respectively, the cumulative probability is 0.8 × 0.7 × 0.9 = 0.504); sorting all paths from high to low according to their cumulative probability values, selecting the path at the top of the list as the optimal path, which has the highest confidence and is closest to the real speech recognition result; extracting the candidate words stored in all nodes on the optimal path and arranging them according to the chronological order of the nodes in the path to form a complete word sequence.
[0044] Step 501: Using a pre-trained lightweight punctuation prediction model, perform real-time punctuation prediction on the word sequence to obtain the predicted probability of sentence termination characters. Simultaneously, extract the semantic embedding vectors of the word sequence and obtain the historical semantic embedding vectors corresponding to cached historical word sequences. Calculate the cosine similarity between the semantic embedding vector of the current word sequence and the semantic embedding vectors of the historical word sequences. When a sentence termination character appears in the punctuation prediction result and the cosine similarity is lower than a preset threshold, it is determined that a complete semantic boundary has been reached. When a complete semantic boundary is reached, convert the word sequence into complete source language recognition text. Specifically, this includes: building a punctuation prediction model using a lightweight CNN+LSTM architecture. The model includes an input layer, a convolutional feature extraction layer, an LSTM temporal modeling layer, a fully connected layer, and an output layer to ensure real-time prediction performance. The convolutional feature extraction layer... The model extracts local features from word sequences using layer 1, captures temporal relationships between words using LSTM layers, and outputs predicted probabilities for various punctuation marks using fully connected layers and the Softmax activation function. It selects multilingual text corpora with punctuation covering diverse scenarios such as daily conversations, business negotiations, and border inspections as training data. The corpus includes text fragments corresponding to common sentence-ending characters such as periods, question marks, and exclamation marks, as well as intermediate punctuation marks such as commas and semicolons. The training data is divided into training, validation, and test sets in an 8:1:1 ratio. The model is iteratively trained using the Adam optimization algorithm with cross-entropy loss as the optimization objective. In each iteration, the word sequence is converted into word embedding vectors and input into the model. The loss value between the predicted punctuation distribution and the actual punctuation labels is calculated. Backpropagation updates the model parameters. The model is iterated until the punctuation prediction accuracy on the validation set no longer improves. The model parameters are then saved to local storage on the terminal.The word sequence extracted in step 500 is input into the trained lightweight punctuation prediction model. The model first extracts local features of the word sequence through a convolutional feature extraction layer, then captures the temporal correlation between words through an LSTM temporal modeling layer and integrates them into a global feature vector. Subsequently, the global feature vector is input into a fully connected layer and mapped to a score vector of the punctuation category dimension (including all preset punctuation types such as period, question mark, exclamation mark, comma, and semicolon) through a linear transformation. Finally, the score vector is input into the Softmax activation function for probability normalization calculation. Specifically, for the score value corresponding to each punctuation mark in the score vector, the natural exponent value of the score value is first calculated, and then the natural exponent value is used... Dividing by the sum of the natural index values of all punctuation marks, we obtain the predicted probability of each punctuation mark at the end of the current word sequence. Of particular interest are the predicted probabilities of sentence-ending markers such as periods, question marks, and exclamation marks. These markers are core symbols in natural language that indicate the semantic and grammatical completion of sentences. Their predicted probabilities directly correspond to the model's assessment of the completeness of the current word sequence. Higher probabilities indicate that the model determines the current word sequence has semantically expressed a complete meaning and grammatically conforms to sentence structure, thus meeting the conditions for forming a complete sentence by ending with a punctuation mark. Lower probabilities indicate that the current word sequence is more likely a fragment of a sentence, requiring additional words to form a complete semantic unit.
[0045] Based on a general terminology database, a specialized terminology database, and a multilingual parallel corpus stored locally on smart terminals, each word is assigned a pre-labeled fixed-dimensional semantic feature vector. This vector contains semantic information, part-of-speech information, and domain adaptation information of the word. For the current word sequence, the semantic feature vector corresponding to each word in the sequence is first extracted. Then, the semantic embedding vector of the entire word sequence is calculated using the mean pooling method. That is, the semantic feature vectors of all words are averaged element by element. Specifically, for each dimension of the vector, the values of all words in that dimension are added together, and the sum is divided by the total number of words in the word sequence to obtain the final value of that dimension. The system combines the semantic representation vectors of the current word sequence to form a cosine similarity value. During the translation process, the system caches the semantic embedding vectors corresponding to the processed historical word sequences in real time, forming a historical semantic vector library to ensure that subsequent similarity calculations can be quickly accessed. The similarity between the current semantic embedding vector and the historical semantic embedding vector is calculated using a formula. The calculation process is as follows: first, the dot product of the two vectors is calculated (corresponding elements are multiplied and then summed); then, the magnitude of each vector is calculated separately (each element is squared, the sum is summed, and the square root is taken); finally, the dot product result is divided by the product of the two magnitudes to obtain the cosine similarity value. This value ranges from -1 to 1. The closer it is to 1, the more similar the semantics are; the closer it is to -1, the greater the semantic difference is. A preset threshold of 0.6 for cosine similarity is set. This threshold is determined based on multi-scenario semantic boundary detection experiments, which can effectively distinguish between topic continuation and transition, and avoid misjudgment or omission. When the predicted probability of the sentence end mark in the punctuation prediction result is greater than 0.5, and the cosine similarity between the current semantic embedding vector and the historical semantic embedding vector is less than 0.6, the current word sequence is determined to have formed a complete semantic unit and reached the complete semantic boundary. If the above two conditions are not met simultaneously, the current word sequence and semantic embedding vector are cached, and a re-judgment is made after subsequent word additions. When the complete semantic boundary is determined to have been reached, the word sequence is refined and worded according to the grammatical rules of the source language, for example, in Chinese... The system follows the core word order of modifier + subject + adverbial + predicate + modifier + object + complement, while English follows the basic word order of subject + predicate + object. Any word combinations deviating from the standard word order are corrected one by one to ensure grammatical logic. Necessary connecting elements are then added based on the semantic relationships between word elements. If a preposition is needed between a noun and a verb, an auxiliary word between a verb and an object, or a conjunction is needed between phrases, then appropriate auxiliary words, prepositions, and conjunctions are added accordingly. Simultaneously, the sentence-ending marker (period, question mark, or exclamation mark) with the highest predicted probability is selected and added to the end of the word sequence. Combined with a locally stored corpus, the system optimizes the fluency of expression, ultimately forming a complete, coherent source language recognition text that conforms to the source language's expression norms.
[0046] Step 502 involves inputting the complete source language recognition text, cached historical dialogue text, pre-defined domain knowledge base information, and reconstructed temporal sentiment state vector into a pre-trained context translation model. This pre-trained model performs contextual semantic fusion and disambiguation translation to generate the final optimized translation. Specifically, this includes building a context translation model based on the Transformer architecture, introducing an emotional attention mechanism and a domain adaptation module. The model comprises an encoder, a decoder, a context fusion layer, an emotional interaction layer, and a domain knowledge base interface. The encoder captures the semantic features and temporal relationships of the source language text, the decoder generates the target language translation, the context fusion layer fuses historical dialogue information, and the emotional interaction layer adapts the temporal sentiment state vector. The system uses state vectors and a domain knowledge base interface that supports calls to general and specialized terminology libraries. Training data is selected from multilingual parallel corpora, dialogue history corpora, domain-specific terminology corpora, and sentiment-annotated speech-text alignment corpora, covering diverse scenarios. The training data is divided into training, validation, and test sets in an 8:1:1 ratio. A joint optimization objective is used, employing bilingual alignment loss, sentiment matching loss, and domain adaptation loss. The Adam optimization algorithm is used for iterative model training. In each iteration, source language text, historical dialogue text, domain terminology, and sentiment vectors are input into the model. The loss values between the generated translation and the actual translation are calculated, and backpropagation updates the model parameters. Once the semantic fidelity no longer improves on the validation set, the model parameters are saved and deployed locally.
[0047] The complete source language recognition text generated in step 501, the historical dialogue text cached by the system, the preset domain knowledge base information (including a general terminology library and a corresponding scenario-specific terminology library), and the temporal sentiment state vector reconstructed in step 202 are all input into the context translation model. The model calls information from each dimension through a dedicated interface. First, it performs word segmentation and semantic annotation on the complete source language recognition text, extracts key information and marks context associations on the historical dialogue text, performs scene matching and terminology filtering on the domain knowledge base information, and analyzes the sentiment intensity and sentiment type of the temporal sentiment state vector to ensure the completeness, timeliness, and parsability of the input data. The model initiates a cross-text semantic association process through the context fusion layer, matching the referential words and omitted components in the current complete source language recognition text with the corresponding content in the historical dialogue text one by one, clarifying the specific referents of referential words such as "this," "the above," etc., and supplementing the omitted subjects, objects, and other core components in the sentence to resolve semantic ambiguity caused by missing context. Then, the model deeply analyzes the temporal sentiment state vector through the sentiment interaction layer to extract the sentiment intensity value and sentiment. Type labels (e.g., positive, neutral, negative) are used to adjust the tone of the translation based on emotional intensity. The translation matches the corresponding target language expression habits based on the emotional type (e.g., negative emotions correspond to euphemistic and humble sentence structures, while positive emotions correspond to brisk and enthusiastic sentence structures), ensuring the translation is not only semantically accurate but also closely matches the speaker's emotional inclination. Simultaneously, a precise terminology matching process is executed through a domain knowledge base interface. First, potential professional terms are located in the current complete source language recognition text based on word length and domain features. Then, a full comparison is performed word-by-word and meaning-by-meaning with the corresponding scenario's professional terminology database. Completely identical terms are directly matched... A one-to-one mapping between the source and target languages is established to achieve accurate conversion. If no perfect match is found, the core semantics of the terms are broken down (subject, action, attribute), compared with a library of commonly used expressions in the domain, and the expression with the highest semantic equivalence and scenario adaptability is selected as the replacement to avoid terminological bias and ensure the professionalism of the translation in scenarios such as business negotiations and border inspections. Subsequently, the decoder initiates the semantic disambiguation and translation generation process. For polysemous words, the referential relationship is clarified in combination with the context, professional semantics are locked in with domain information, and the emphasis of the expression is adjusted with emotional information. For ambiguous sentences, the optimal interpretation is selected through logical decomposition and scenario comparison. Then, according to the grammatical rules and word order habits of the target language, sentence-by-sentence sentence-by-sentence sentence-by-sentence conversion (such as adjusting the position of Chinese-English attributives), accurate vocabulary replacement, and supplementation of connecting components that adapt to causal, adversative, and other logical connections are completed. Finally, three rounds of optimization are performed: checking for grammatical issues such as subject-verb agreement and tense, verifying that no core semantics are omitted, and adjusting the fluency of the sentence to fit the expression habits of the target language, ultimately generating a final optimized translation that is semantically accurate, emotionally appropriate, professional, and naturally coherent.
[0048] Step 503: Check whether the source language phrase or clause structure corresponding to the generated candidate translation fragment is contained in the complete source language recognition text. If it is contained, replace the candidate translation fragment with the corresponding part of the final optimized translation to form a coherent translation text flow; otherwise, insert the final optimized translation as a new translation fragment into the translation text flow. Specifically, this includes: accurately extracting the complete word sequence and word arrangement order of the source language phrase or clause corresponding to each candidate translation fragment generated in step 402, and using it as the structure to be matched; using the word sequence of the complete source language recognition text generated in step 501 as a benchmark, matching each word of the structure to be matched in sequence to confirm whether there is a continuous word sequence in the complete source language recognition text that has the same number of words, consistent word content, and completely matching word order as the structure to be matched; if it is determined that the structure to be matched is contained in the complete source language recognition text, then locate the continuous word sequence. The starting and ending positions of the word sequence in the complete source language recognition text are then used to find the corresponding translation segments in the final optimized translation. These segments are then extracted and used to replace the original candidate translation segments, ensuring that the replaced translation content completely corresponds to the complete semantics and improving translation accuracy. If it is determined that the structure to be matched is not contained in the complete source language recognition text, that is, the phrase or clause belongs to an incomplete semantic unit, the temporal position of the candidate translation segment in the generated translation text stream is determined, and the final optimized translation is inserted after that position to maintain the semantic coherence and temporal progression of the translation text stream. After completing the replacement and insertion operations of all candidate translation segments, all translation segments are traversed and reordered according to the chronological order of the source language speech acquisition time corresponding to each segment. Duplicate translation content is removed to ensure that the temporal order of the translation text stream is completely consistent with the input temporal order of the source language speech, forming a coherent, accurate, and non-redundant final translation text stream.
[0049] In this embodiment, the optimal path selection mechanism selects the most reliable word sequence based on cumulative probability values, reducing the impact of low-confidence recognition results on translation accuracy; the dual judgment mechanism of punctuation prediction and semantic similarity calculation identifies complete semantic boundaries, avoiding semantic incompleteness caused by premature translation or delay caused by late translation, balancing translation real-time performance and completeness; the context translation model integrates complete source language text, historical dialogue, domain knowledge, and sentiment state vectors, effectively resolving ambiguity issues, improving the accuracy of professional terminology translation, and the adaptability of the translation to the speaker's emotions, making the translation more in line with actual communication scenarios; the replacement and insertion strategy of the translation text stream ensures seamless connection between the generated candidate translation fragments and the final optimized translation, avoiding translation repetition or breakage, improving the coherence and readability of the translation text stream, and ensuring the user's communication experience.
[0050] In a preferred embodiment of the present invention, step 6 includes: Step 600: Extract the fundamental frequency F0 trajectory, phoneme duration sequence, and energy envelope from the source language audio stream as prosodic features. Specifically, this includes: preprocessing the source language audio stream by first performing a pre-emphasis operation, boosting high-frequency components and filtering low-frequency noise using a first-order high-pass filter; then using frame-by-frame windowing to divide the audio stream into several frames with a frame length of 20 milliseconds and a frame shift of 10 milliseconds, and adding a Hamming window to each frame to reduce spectral leakage. ( The autocorrelation function is calculated using the intra-frame sampling point index (ranging from 0 to L-1, where L is the total number of sampling points in a single frame). The formula for the autocorrelation function is: ,in The delay step is defined as the number of delay steps, ranging from 0 to L-1. The delay step corresponding to the peak value of the autocorrelation function is found, and the fundamental frequency (F0) of the frame is calculated by the ratio of the sampling frequency to the delay step. The fundamental frequency (F0) values of all frames are arranged in chronological order to form a continuous fundamental frequency (F0) trajectory, which visually reflects the pitch variation pattern of the source language speech. The preprocessed audio stream is segmented into phonemes, and combined with the word sequence generated in step 300, the audio stream is mapped to the corresponding phoneme units. The number of audio frames corresponding to each phoneme unit is calculated, and the duration of a single phoneme is calculated using the formula: phoneme duration = number of frames corresponding to the phoneme × frame shift + frame length. Time; the duration of all phonemes is arranged in order of their position in the word sequence to form a phoneme duration sequence, reflecting the speech rate and rhythm changes of the source language speech; short-time energy is calculated for each frame of audio after windowing, and the short-time energy is calculated as the sum of the squares of the amplitudes of all audio sampling points in the frame; the short-time energy values of all frames are arranged in chronological order to form an energy envelope; to reduce noise interference, the energy envelope is smoothed, and the mean energy of each frame is calculated using the moving average method, with the window size set to 5 frames, finally obtaining the smoothed energy envelope, reflecting the volume changes of the source language speech.
[0051] Step 601: The translated text stream, prosodic features, emotional intensity harmonization factor, and control instructions issued by the adaptive control center are input into the speech synthesis module. The speech synthesis module uses the emotional intensity harmonization factor to modulate the emotional color intensity of the fundamental frequency F0 trajectory and energy envelope. Based on the control instructions, the corresponding acoustic model parameters are selected, and the modulated prosodic features are fused to synthesize the translated text stream into a target language speech stream. Specifically, this includes: inputting the translated text stream, the prosodic features extracted in step 600 (including the fundamental frequency F0 trajectory, phoneme duration sequence, and energy envelope), the emotional intensity harmonization factor, and the control instructions issued by the adaptive control center into the speech synthesis module to complete the synthesis of the target language speech stream. The specific process is as follows: The adaptive control center is an independent software decision module that continuously analyzes the current communication scenario. The system's real-time performance status and historical dialogue information are used to dynamically generate control commands. The emotional intensity harmonization factor is calculated by normalizing the temporal emotional state vector reconstructed in step 202, with a value ranging from 0 to 1. The larger the value, the stronger the emotional color. Based on this factor, the fundamental frequency F0 trajectory and energy envelope are modulated. The modulation calculation formula is as follows: Modulated fundamental frequency F0 value = original fundamental frequency F0 value × (1 + emotional intensity harmonization factor), Modulated energy envelope value = original energy envelope value × (1 + emotional intensity harmonization factor). During the modulation process, the phoneme duration sequence remains unchanged to ensure that the speech rate of the synthesized speech is consistent with the rhythm of the source language speech. Through modulation, the pitch and volume of the synthesized speech are matched with the speaker's emotional tendency. For example, positive emotions correspond to higher fundamental frequencies and energy, while negative emotions correspond to relatively flat fundamental frequencies and energy.
[0052] The speech synthesis module receives control commands from the adaptive control center, which include parameters such as scene type and speech style. A pre-defined acoustic model parameter library is constructed according to scene categories. Each scene entry contains core parameters with clearly defined values, and these parameters are finely adjusted for different styles. For example, in business negotiation scenarios, the fundamental frequency dynamic range is 80 to 220 Hz (suitable for both male and female voices), the speech rate baseline range is 4 to 5 syllables / second, the formant frequencies are (F1: 500 to 700 Hz, F2: 1500 to 1800 Hz, F3: 2500 to 2800 Hz), the spectral envelope coefficient is set to order 16 with a value range of 0.1 to 0.8, and the voiced / unvoiced threshold is 0.3 (autocorrelation function value). A narrow fundamental frequency fluctuation range, a smooth speech rate range, and a regular spectral envelope are set to suit formal and rigorous tones. In everyday conversation scenarios, the fundamental frequency dynamic range is 60 to 250 Hz, and the speech rate baseline range is 5 to 6.5 syllables / second. The speech rate is set at 4.5 to 5.5 syllables / second, formant frequencies (F1: 450 to 750 Hz, F2: 1400 to 1900 Hz, F3: 2400 to 2900 Hz), spectral envelope coefficients are set to order 16 with values ranging from 0.05 to 0.9, and voiced / unvoiced threshold is set to 0.25. The fundamental frequency fluctuation limit is relaxed, and the conversational pause interval and formant parameters are optimized to match the rhythm of natural communication. For border inspection scenarios, the fundamental frequency dynamic range is 70 to 230 Hz, the speech rate baseline is 4.5 to 5.5 syllables / second, formant frequencies (F1: 550 to 720 Hz, F2: 1600 to 1850 Hz, F3: 2600 to 2850 Hz), spectral envelope coefficients are set to order 16 with values ranging from 0.15 to 0.85, and voiced / unvoiced threshold is set to 0.32. The formant clarity coefficient is improved, and a speech rate baseline slightly slower than that of daily conversation is set to enhance pronunciation recognition and ensure that the synthesized speech is suitable for the current usage scenario.
[0053] The modulated fundamental frequency F0 trajectory, energy envelope, and unmodulated phoneme duration sequence are deeply fused with the translated text stream. First, prosodic features are precisely mapped to each word and syllable according to the grammatical structure and semantic logic of the translated word units, establishing a one-to-one correspondence. The speech synthesis module, based on the selected acoustic model parameters, converts them into acoustic feature parameters according to the following steps: First, Mel-spectral coefficient transformation. The fused signal is windowed in frames (25ms frame length, 10ms frame shift), and the spectrum is obtained through Fast Fourier Transform. The Mel spectrum is extracted using a 24-dimensional Mel filter bank. After taking the logarithm of the Mel spectrum, a Discrete Cosine Transform is performed, and the first 12 order coefficients are taken as Mel-spectral coefficients. Simultaneously, the first and second order difference coefficients are calculated. The process involves three steps: first, filling in the timing information; second, calculating the fundamental frequency deviation value by subtracting the fundamental frequency reference value of the acoustic model corresponding to the scene from the fundamental frequency F0 value of each frame after modulation, which is used to correct the tone of the synthesized speech; and third, converting the spectrum smoothing parameters by using a 5-frame moving average method to smooth the spectrum amplitude based on the spectrum envelope coefficient, with the weight decreasing according to the inter-frame distance (0.4 for the center frame, 0.25 for each of the two adjacent frames, and 0.05 for each of the edge frames), to obtain the spectrum smoothing parameters. Then, the acoustic feature parameters are reconstructed using a vocoder to generate the target language speech waveform corresponding to the syllables. All speech waveforms are spliced together in time order, and the spectral abrupt changes at the splicing points are corrected by linear interpolation to form a continuous and fluent target language speech stream.
[0054] Step 602 involves real-time playback of the target language speech stream via an audio playback device. Specifically, this includes: transmitting the target language speech stream generated by the speech synthesis module to the buffer of the audio playback device in the form of streaming data. The buffer capacity is set to 2 seconds of speech data, employing a first-in, first-out (FIFO) storage mechanism to ensure that the speech data is output sequentially according to the synthesis timeline, avoiding playback errors; the audio playback device reads the speech data from the buffer, converts the digital signal into an analog audio signal via a digital-to-analog converter, amplifies the signal, and then plays it through output components such as speakers; during playback, the buffer data volume is monitored every 10 milliseconds, with a buffer threshold set to 500 milliseconds of speech data. When the buffer data volume falls below this threshold, the speech synthesis module is immediately triggered to generate subsequent speech data and supplement the buffer, ensuring smooth and uninterrupted playback and maintaining the continuity of the streaming output; the playback device supports real-time adjustment of volume and speech rate, with a volume adjustment range covering 20 to 80 decibels and a speech rate adjustment range of 0.8 to 1.2 times the base speech rate. Users can adjust the output parameters according to their actual needs to ensure that the playback effect of the target language speech stream meets the usage requirements of different scenarios.
[0055] In this embodiment, the accurate extraction of prosodic features fully preserves the core information of the source language speech, such as pitch, rhythm, and volume, providing a high-quality prosodic reference for the synthesis of the target language speech stream and ensuring the naturalness and expressiveness of the synthesized speech. Prosodic feature modulation based on emotional intensity harmonization factors ensures a high degree of match between the emotional color of the synthesized speech and the emotional tendency of the source language speaker, enhancing the realism and approachability of the voice interaction. The acoustic model parameter selection mechanism combined with control commands achieves precise adaptation of the synthesized speech style to the usage scenario, meeting the voice output needs in diverse scenarios. The collaborative mechanism of streaming synthesis and real-time playback ensures the continuity and low latency of the voice output, achieving seamless conversion from the translated text stream to the speech stream and improving the interactive efficiency in real-time translation scenarios.
[0056] like Figures 2 to 6 As shown, embodiments of the present invention also provide a real-time speech translation smart terminal based on multimodal fusion, comprising: System bus 501, processor 502, memory 503, power supply component 504, network component 505, display screen 506, speaker 507, camera one 508, camera two 509, microphone one 510 and microphone two 511; The memory 503 stores a computer program; the processor 502 is used to execute the computer program to perform the following steps: The microphone 510 and microphone 511 are used to collect the source language speech signal of at least one speaker; the camera 508 and camera 509 are used to collect the facial expression and posture image information of the corresponding speaker. The source language speech signal is denoised and speech endpoints are detected, and the processed speech signal is then subjected to streaming speech recognition to obtain the source language recognized text; facial expression and posture image information is analyzed to obtain facial expression feature information to assist semantic understanding. The source language recognized text and facial expression feature information are fused into multimodal features, and machine translation is performed based on the fused features to generate target language translated text. Prosodic features are extracted from the source language speech signal and combined with facial expression features to synthesize speech from the target language translation text, generating the target language speech signal. The target language speech signal is played through the speaker 507. The processor 502 is also used to monitor the ambient noise intensity in real time during the speech acquisition process and dynamically adjust the noise reduction parameters and pickup directivity parameters of the microphone 1 510 and the microphone 2 511.
[0057] like Figures 2 to 6 As shown, embodiments of the present invention also provide a real-time speech translation smart terminal based on multimodal fusion, further comprising: The main module 512 is used to integrate and house the processor 502, memory 503, power supply component 504, network component 505, display screen 506, and speaker 507; The camera module 513 is mounted on the main module 512 in a pitch-adjustable manner via a damping pivot, and integrates the camera 508 and microphone 510 for collecting audio and video information in the first direction. A main body cable outlet 514 is provided in the main body module 512 for passing through a connecting cable; The main support 515 is detachably connected to the main module 512 and is used to support the main module 512 on a plane. The peripheral module 516 integrates the second camera 509 and the second microphone 511, which are used to collect audio and video information in the second direction. A peripheral cable outlet 517 is provided in the peripheral module 516 for passing a connecting cable. The peripheral module 516 is electrically connected to the processor 502 in the main module 512 via the cable through the peripheral cable outlet 517 and the main body cable outlet 514. The peripheral fastener 518 is used to fix the peripheral module 516 to the external support structure.
[0058] It should be noted that the real-time voice translation smart terminal is the real-time voice translation smart terminal corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0059] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A real-time speech translation method based on multimodal fusion, characterized in that, The method includes: Step 1: Locate three facial feature points: the center of the left pupil, the center of the right pupil, and the center of the lips. Based on the motion coordinates of these three facial feature points, construct a temporal trajectory function with time as the horizontal axis and normalized motion coordinates as the vertical axis. Step 2: Perform discrete-time Laplace transform on the time-series trajectory function to obtain the complex frequency domain characteristic function, and perform partial fraction expansion and inverse Z-transform on the complex frequency domain characteristic function in the region of convergence to reconstruct the time-domain emotional state vector, and calculate the emotional intensity harmonic factor. Step 3: Input the continuous clean speech data block into the pre-trained streaming speech recognition model for recognition, obtain the text segment at the current time point, and update and maintain the recognition candidate word map based on the text segment; Step 4: Detect whether a preliminary meaningful phrase or clause structure is formed in the candidate word graph. When a preliminary meaningful phrase or clause structure is detected, perform a preliminary translation of the phrase or clause structure to generate candidate translation fragments. Step 5: Semantic boundary judgment is performed on the candidate translation fragments. When the judgment reaches the complete semantic boundary, the pre-trained context translation model is triggered to generate the final optimized translation. The final optimized translation is then used to correct or replace the candidate translation fragments to form the translation text stream. Step 6: Extract the prosodic features of the source language audio stream, combine them with the emotional intensity harmonization factor and the control instructions issued by the adaptive control center, and perform speech synthesis on the translated text stream to generate and play the target language speech stream in real time.
2. The real-time speech translation method based on multimodal fusion according to claim 1, characterized in that, Before step 1: Capture or receive source language audio streams and speaker video streams in real time; Real-time noise reduction and speech endpoint detection are performed on the captured or received source language audio stream to generate continuous clean speech data blocks, and face regions are detected from the captured or received speaker video stream.
3. The real-time speech translation method based on multimodal fusion according to claim 2, characterized in that, Step 1 includes: Based on the face region, three facial feature points are located: the center of the left pupil, the center of the right pupil, and the center of the lips. Based on the pixel coordinates of three facial feature points in consecutive video frames, the motion coordinates of each feature point are calculated, and the motion coordinates are normalized to obtain normalized motion coordinates. Based on normalized motion coordinates and corresponding timestamps, a time-series trajectory function is constructed with time as the horizontal axis and normalized motion coordinates as the vertical axis.
4. The real-time speech translation method based on multimodal fusion according to claim 3, characterized in that, Step 2 includes: Performing a discrete-time Laplace transform on the time-series trajectory function yields a complex frequency domain characteristic function for the complex frequency variable; Analyze the expression of the complex frequency domain characteristic function, solve for the range of complex frequency variables that make the complex frequency domain characteristic function sequence converge, and obtain the region of convergence of the complex frequency domain characteristic function; Within the region of convergence, the complex frequency domain feature function is decomposed into the sum of multiple first-order fractions to complete the partial fraction expansion; the inverse Z-transform operation is performed on each of the expanded first-order fractions, and the inverse transform results of each fraction are superimposed in the time domain to reconstruct the time-domain sentiment state vector. Numerical differentiation is performed on the time-series trajectory function to calculate its first and second derivative values at each sampling time point. The location where the sign of the second derivative value changes is detected, and this location is determined as the inflection point of the time-series trajectory function. The absolute value of the first derivative value at each inflection point is extracted as the instantaneous change intensity. Based on the number of all detected inflection points and the average value of instantaneous change intensity, an inflection point modulation factor is calculated; based on the denominator pole value, the corresponding numerator residual value, and the inflection point modulation factor of each first-order fraction after partial fraction expansion, an emotion intensity harmonization factor is obtained through weighted fusion calculation.
5. The real-time speech translation method based on multimodal fusion according to claim 4, characterized in that, Step 3 includes: A continuous block of clean speech data is input into a pre-trained streaming speech recognition model in chronological order. The streaming speech recognition model processes the continuous block of clean speech data and outputs one or more candidate words at the current time point and the probability corresponding to each candidate word. At the initial moment, starting from the first candidate word, a candidate word graph containing nodes and directed edges is initialized, where nodes represent candidate words and directed edges represent the transition relationship between candidate words and are weighted by the probability corresponding to the candidate word. At each subsequent time point, new candidate words are added as new nodes to the candidate word graph. The probability corresponding to the candidate word is used as the weight to establish a directed edge from the optional node at the previous time point to the current new node. At the same time, the cumulative probability value of each path from the starting point to the current new node is updated to complete the incremental update and maintenance of the candidate word graph.
6. The real-time speech translation method based on multimodal fusion according to claim 5, characterized in that, Step 4 includes: After updating the candidate word graph, starting from the new node corresponding to the current time point, backtrack along the directed edges to identify the candidate word graph and obtain several candidate paths ending at that node. For each candidate path, the word sequence is tagged with part of speech and shallow dependency analysis is performed. When the analysis results determine that the word sequence of a candidate path conforms to the preset phrase or clause grammatical template, it is determined that a preliminary meaningful phrase or clause structure has been formed. Extract the source language word sequence corresponding to the phrase or clause structure, convert it into coherent source language text, and perform a preliminary translation of the converted source language text in real time to generate candidate translation fragments.
7. The real-time speech translation method based on multimodal fusion according to claim 6, characterized in that, Step 5 includes: Based on the path with the highest cumulative probability from the starting point to the current node in the candidate word graph, obtain the word sequence corresponding to that path; A pre-trained lightweight punctuation prediction model is used to predict punctuation marks in real time on word sequences, obtaining the predicted probability of sentence termination marks. Simultaneously, semantic embedding vectors of word sequences are extracted, and cached historical semantic embedding vectors corresponding to historical word sequences are obtained. The cosine similarity between the semantic embedding vector of the current word sequence and the semantic embedding vectors of historical word sequences is calculated. When a sentence termination mark appears in the punctuation prediction result and the cosine similarity is below a preset threshold, it is determined that a complete semantic boundary has been reached. When a complete semantic boundary is reached, the word sequence is converted into complete source language recognition text. The complete source language recognition text, cached historical dialogue text, preset domain knowledge base information, and reconstructed temporal sentiment state vector are input into a pre-trained context translation model. The pre-trained context translation model performs context semantic fusion and disambiguation translation to generate the final optimized translation. Check whether the source language phrase or clause structure corresponding to the generated candidate translation fragment is contained in the complete source language recognized text. If it is contained, use the corresponding part of the final optimized translation to replace the candidate translation fragment to form a coherent translation text flow; otherwise, insert the final optimized translation as a new translation fragment into the translation text flow.
8. The real-time speech translation method based on multimodal fusion according to claim 7, characterized in that, Step 6 includes: Extract the fundamental frequency F0 trajectory, phoneme duration sequence, and energy envelope from the source language audio stream as prosodic features; The translated text stream, prosodic features, emotional intensity harmonization factor, and control instructions issued by the adaptive control center are input into the speech synthesis module. The speech synthesis module uses the emotional intensity harmonization factor to modulate the emotional color intensity of the fundamental frequency F0 trajectory and energy envelope. According to the control instructions, the corresponding acoustic model parameters are selected, and the modulated prosodic features are fused to synthesize the translated text stream into the target language speech stream. The target language audio stream is played in real time through an audio playback device.
9. A real-time speech translation intelligent terminal based on multimodal fusion, wherein the real-time speech translation intelligent terminal implements the method as described in any one of claims 1 to 8, characterized in that, include: System bus, processor, memory, power supply components, network components, display screen, speaker, camera one, camera two, microphone one and microphone two; The memory stores a computer program; the processor executes the computer program to perform the following steps: The microphones 1 and 2 are used to acquire the source language speech signal of at least one speaker; the camera 1 and camera 2 are used to acquire the facial expression and posture image information of the corresponding speaker. The source language speech signal is subjected to noise reduction and speech endpoint detection processing, and the processed speech signal is subjected to streaming speech recognition to obtain the source language recognized text. Facial expression and pose image information is analyzed to obtain facial expression feature information for assisting semantic understanding; The source language recognized text and facial expression feature information are fused into multimodal features, and machine translation is performed based on the fused features to generate target language translated text. Prosodic features are extracted from the source language speech signal and combined with facial expression features to synthesize speech from the target language translation text, generating the target language speech signal. The target language speech signal is played through the speaker, wherein the processor is also used to monitor the ambient noise intensity in real time during the speech acquisition process and dynamically adjust the noise reduction parameters and pickup directivity parameters of microphone one and microphone two.
10. The real-time speech translation intelligent terminal based on multimodal fusion according to claim 9, characterized in that, Also includes: The main module is used to integrate and house the processor, memory, power supply components, network components, display screen, and speaker; The camera module is mounted on the main module in a pitch-adjustable manner via a damping hinge, and integrates the camera and microphone for collecting audio and video information in a first direction. A main body cable outlet hole is provided in the main body module for threading connecting cables; A main support frame, detachably connected to the main module, is used to support and place the main module on a plane. The peripheral module integrates the second camera and the second microphone for collecting audio and video information in the second direction; A peripheral cable outlet is provided on the peripheral module for passing a connecting cable through it. The peripheral module is electrically connected to the processor in the main module via the cable through the peripheral cable outlet and the main body cable outlet. Peripheral fasteners are used to fix the peripheral module to an external support structure.