Intelligent real-time language translation system, method and terminal

By combining multimodal data collection and sentiment semantic fusion with cultural adaptation and custom rules, the problem of lack of emotion and insufficient personalization in existing translation systems has been solved, achieving accurate transmission of emotional intent and personalized translation to meet diverse scenario needs.

CN122113942APending Publication Date: 2026-05-29深圳市原上科技技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511296881.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing language translation systems lack the ability to integrate multimodal information and cannot capture the speaker's emotional inclination, resulting in semantically accurate but emotionally deficient translations. Furthermore, they fail to adapt to users' personalized needs, leading to insufficient cultural adaptation and risks of privacy breaches.

Method used

The system employs a multimodal acquisition module to simultaneously acquire speech and facial expression data, constructs a personalized voiceprint model, and combines an emotion-semantic fusion module and a cultural adaptation module to achieve accurate emotion transmission and cultural habit adaptation. Furthermore, it generates speech with the original speaker's timbre through a custom language rule engine and a real-time voiceprint cloning engine, supporting personalized translation and privacy protection.

Benefits of technology

It achieves accurate reproduction of emotional intent in cross-language communication, meets personalized translation needs, avoids communication misunderstandings, has real-time and privacy protection capabilities, and is suitable for diverse scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113942A_ABST
    Figure CN122113942A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent real-time language translation system and method and a terminal, and belongs to the technical field of language translation, and comprises a multi-modal acquisition module, a voiceprint feature extraction module, an emotional and semantic fusion module, a dynamic context module, a cross-language translation engine, a self-defined language rule engine, a real-time voiceprint cloning engine, an emotional migration module, a culture adaptation module and a multi-modal output module. The system supports multi-person dialogue scenes and can be continuously iteratively improved through an adaptive optimization module. The system effectively solves the problems of loss of emotional information, single tone, lack of personalized customization and cultural inadaptation in the prior art, and realizes natural, accurate and personalized intelligent cross-language communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of language translation technology, and more specifically, to an intelligent real-time language translation system, method, and terminal. Background Technology

[0002] With the deepening of globalization, cross-language communication has become a core requirement in daily life, business activities, cultural dissemination, and international cooperation. While traditional language translation tools have achieved initial breakthroughs at the basic semantic conversion level, they still have significant limitations in practical applications. Current mainstream translation systems mostly use single text or speech as input sources, lacking the ability to integrate multimodal information. They struggle to capture the emotional inclinations conveyed by speakers through facial expressions and tone changes, often resulting in translations that are "semantically accurate but emotionally lacking," failing to accurately reproduce the emotional intent in communication and thus affecting communication effectiveness. Furthermore, existing systems generally adopt generic translation strategies, failing to fully consider users' personalized needs. For example, they cannot adjust translation logic based on user identity, usage scenario, and domain attributes. In professional dialogues or communications containing private information or special coded language, this can easily lead to translation errors in terminology or risks of privacy leaks.

[0003] In the speech output stage, traditional systems often use fixed synthesized speech libraries, resulting in significant differences between the generated target language speech and the original speaker's vocal characteristics. This not only reduces the naturalness of communication but may also lead to misunderstandings due to mismatches between speech style and context. Furthermore, different languages ​​correspond to different cultural backgrounds. Existing translation systems often focus only on linguistic conversion, neglecting the cultural adaptability of emotional expression patterns and semantic habits. This results in translations that sound awkward in the target language's cultural context and may even lead to ambiguity due to cultural differences.

[0004] From a technical optimization perspective, existing translation systems largely operate independently, lacking dynamic collaboration mechanisms. For example, contextual information cannot be fed back to the translation engine in a timely manner, and emotional features cannot be deeply integrated with semantic content, resulting in overall translation accuracy and real-time performance that are insufficient to meet the demands of complex scenarios. Furthermore, system parameter optimization relies heavily on pre-set training data and cannot adaptively adjust based on real-time user feedback, making it difficult to adapt to different user habits and diverse communication scenarios. Therefore, developing an intelligent real-time language translation system that can integrate multimodal information, achieve personalized translation, ensure accurate emotional transmission, and possess cultural adaptability has become a key direction for addressing the current pain points in cross-language communication. Summary of the Invention

[0005] The purpose of this invention is to provide an intelligent real-time language translation system, method, and terminal to solve the above-mentioned problems.

[0006] To achieve the above objectives, an embodiment of the present invention provides the following technical solution: An intelligent real-time language translation system includes: A multimodal acquisition module is used to simultaneously acquire the user's voice signals and facial expression data; The voiceprint feature extraction module is used to extract the speaker's unique acoustic features from the speech signal, and to build and dynamically update a personalized voiceprint model. The emotion-semantic fusion module is used to extract emotional features (including emotion type, intensity, and expression pattern) from speech and facial expressions, and fuse them with semantic content to form an emotion-semantic representation. The dynamic context module is used to build a context model that includes dialogue history, user characteristics, scene information, and domain labels; A cross-language translation engine used to accurately translate the semantics of a source language into a target language while preserving sentiment marker information; A custom language rules engine is used to automatically replace specific words or expressions according to user-defined rules, supporting code word mapping, privacy protection, and sentiment protection; The real-time voiceprint cloning engine generates basic speech with the original speaker's timbre characteristics based on the target language text and personalized voiceprint model. The emotion transfer module is used to transfer emotional features from the source language to the cloned speech of the target language, so that the emotional expression is consistent with the voiceprint features. The cultural adaptation module is used to adjust emotional expression and semantic representation according to the target language and cultural background. The multimodal output module is used to output target language speech and text that integrates voiceprint features, emotional information, and custom rule application results; The adaptive optimization module is used to optimize the parameters of each module based on user feedback.

[0007] As a further improvement of the present invention, the custom language rule engine includes: The rule management unit supports user-defined rules such as keyword replacement, sensitive word filtering, code word mapping, and domain terminology conversion, and allows users to set the applicable scenarios, objects, sentiment associations, and priorities of the rules; The rule parsing unit is used to parse the triggering conditions and execution actions of rules; The context matching unit is used to determine the applicability of rules based on the current dialogue scenario, the audience, the topic, and the emotional state. The rule execution unit is used to execute the matching rules in the translation process to ensure that the replacement results are consistent with the semantics of the context. The rule conflict resolution unit coordinates multiple rules based on priority, applicability, and user preference when they match simultaneously.

[0008] As a further improvement of the present invention, the voiceprint feature extraction module includes: The acoustic feature extraction unit is used to extract personalized voiceprint features such as spectral features, fundamental frequency curves, formants, and speech rate and rhythm. The voiceprint model building unit constructs an initial model based on a small number of speech samples (≤30 seconds) and dynamically optimizes it through continuous learning; The voiceprint similarity assessment unit is used to quantitatively assess the similarity between cloned speech and the original speaker's speech; The multi-speaker voiceprint management unit is used to store and differentiate the voiceprint models of multiple speakers, and supports real-time switching.

[0009] As a further improvement of the present invention, the emotion transfer module includes: The emotional feature parameterization unit transforms emotional expression into quantifiable acoustic parameters; A cross-linguistic emotion mapping unit establishes an acoustic parameter mapping relationship between the source language and the target language in terms of emotion expression; The emotion-voiceprint coordination unit ensures that the transferred emotion features and the cloned voiceprint features are acoustically consistent. The emotion intensity adaptive unit fine-tunes the emotion intensity according to the emotion expression habits of the target language.

[0010] As a further improvement of the present invention, the emotion semantic fusion module includes: The voice emotion analysis unit is used to identify the type and intensity of emotions through changes in tone, speech rate, and volume. The facial expression recognition unit is used to identify the emotional tendencies corresponding to micro-expressions by analyzing facial feature points. The emotional conflict processing unit is used to assign weights based on contextual information when there is a difference between voice emotion and facial expression emotion. The sentiment-semantic binding unit is used to establish the association between sentiment features and semantic content.

[0011] A smart real-time language translation method includes the following steps: S1. Synchronously collect the user's voice signal and facial expression data, and perform preprocessing; S2. Extract speaker voiceprint features from speech signals and construct or update personalized voiceprint models; S3. Analyze the emotional features and semantic content in speech and facial expressions, and construct an emotion-semantic fusion representation; S4. Based on dialogue history, user characteristics, and scenario information, construct and update a dynamic context model; S5. Translate the source language semantics into target language text while preserving sentiment markers; S6. Call the custom language rule engine, match applicable rules according to the current context, and perform automatic replacement on the target language text; S7. Generate basic cloned speech based on the replaced target language text and personalized voiceprint model; S8. Transfer the emotional features in the source language to the cloned speech of the target language to achieve coordinated expression of emotion and voiceprint; S9. Adapt the translation results to the target language and cultural context; S10. Output the target language speech and text, which integrate the results of fusing voiceprint features, emotional information, and custom rule applications; S11. Collect user feedback and optimize the parameters and models of each module.

[0012] As a further improvement of the present invention, the process of applying the custom rule in step S6 includes: S61. Parse the target language text and identify keywords or expressions that may trigger the rules; S62. Analyze the current contextual information, including scene type, dialogue partners, topic, and emotional state; S63. Match the custom rule that best fits the current context; S64. When a rule conflict is detected, the rule is filtered according to the preset priority and contextual adaptability; S65. Execute the selected rules to replace the target language text; S66. Check the semantic coherence of the replacement result with the context, and make fine adjustments if necessary.

[0013] As a further improvement of the present invention, the emotional transfer process in step S8 includes: S81. Parameterize the emotional expression of the source language speech into acoustic features such as pitch, speech rate, and volume; S82. Based on cross-linguistic sentiment mapping rules, the source language sentiment parameters are converted into sentiment parameters suitable for the target language; S83. Adjust the acoustic parameters of the cloned speech to naturally integrate emotional expression with voiceprint features; S84. Enhance emotionally intense segments to ensure the accurate transmission of key emotional information; S85. Adjust the intensity and manner of emotional expression based on the target language and cultural background.

[0014] As a further improvement of the present invention, in a multi-person dialogue scenario, the following steps are also included: S01. Real-time identification of different speakers, calling the corresponding voiceprint model and custom rule set; S02. Maintain an independent emotional expression pattern library and context sub-model for each speaker; S03. Dynamically adjust translation, voiceprint cloning, and emotion transfer strategies based on speaker identity and dialogue context; S04. Maintain the distinctiveness of the speaker's voiceprint features and the personalization of emotional expression in the translation results; S05. Record the history of emotional interactions and rule applications in multi-person dialogues to support subsequent communication.

[0015] A smart real-time language translation terminal includes: processor; The memory contains computer programs, a user-defined language rule base, and a personalized voiceprint model library; A multimodal acquisition device, including a microphone array and a camera, is used to simultaneously acquire speech and facial expressions; Voiceprint-emotion processing chip, used to accelerate voiceprint feature extraction and emotion transfer calculation; An input interaction device for users to define and manage custom language rules; Audio output device for playing translated speech that incorporates voiceprints, emotions, and custom rules; The communication module is used for data interaction with the cloud.

[0016] Compared with the prior art, the advantages of this invention are: (1) This invention acquires speech and facial expression data simultaneously through a multimodal acquisition module, constructs an emotion-semantic representation by combining an emotion-semantic fusion module, and then achieves accurate transmission of emotion and adaptation of cultural habits through an emotion transfer module and a cultural adaptation module. This makes the speech and text of the target language not only highly consistent with the source language in semantics, but also restores the speaker's emotional intentions. At the same time, it fits the cultural background of the target language, effectively avoiding communication misunderstandings caused by emotional deviations or cultural conflicts, and making cross-language communication closer to the real native language communication experience.

[0017] (2) This invention constructs and dynamically updates a personalized voiceprint model through a voiceprint feature extraction module, and generates speech with a unique user timbre by combining it with a real-time voiceprint cloning engine, thus realizing personalized speech output. The custom language rule engine supports users to define rules such as keyword replacement, sensitive word filtering, and code word mapping according to their needs, which can not only meet the terminology conversion needs of professional fields, but also effectively protect privacy information and avoid leakage of sensitive content. This highly personalized design enables the system to adapt to diverse scenarios such as business, medical, and daily communication, and provide customized translation services for different users.

[0018] (3) This invention accelerates key computational processes through a voiceprint-emotion processing chip, and combines model lightweighting and algorithm optimization to control the latency of the entire process, including translation, voiceprint cloning, and emotion transfer, within a few hundred milliseconds, fully meeting the needs of real-time dialogue. In terms of continuous optimization, the adaptive optimization module collects user feedback and usage behavior data, and uses reinforcement learning algorithms to dynamically adjust the parameters of each module, enabling the system to continuously optimize translation accuracy, emotion transmission effect, and rule matching accuracy as the user's usage time increases. At the same time, the communication module supports data interaction with the cloud, and can obtain the latest model parameters and knowledge base updates in real time, ensuring that the system continuously improves in terms of language coverage and cultural adaptability, providing users with long-term, stable, and continuously optimized translation services. Attached Figure Description

[0019] Figure 1 This is a block diagram of the overall system architecture of the present invention; Figure 2 This is a flowchart illustrating the workflow of the custom rule engine in this invention. Figure 3 This is a flowchart of the voiceprint cloning and emotion transfer process of the present invention; Figure 4 This is a flowchart of the multi-person dialogue scenario processing of the present invention; Figure 5 This is a block diagram of the emotion and semantic fusion module of the present invention. Detailed Implementation

[0020] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0021] Example 1: An intelligent real-time language translation system, specifically comprising: 1. Multimodal Acquisition Module: Employs a multi-device collaborative acquisition architecture. At the hardware level, it integrates a high-sensitivity microphone array and a high-definition camera. The microphone array uses a 4-channel omnidirectional microphone, supporting a 16kHz sampling rate and 24-bit quantization depth, effectively suppressing environmental noise and focusing the speaker's voice signal using beamforming technology to reduce background interference. The camera is a 1080P resolution, 30fps RGB camera equipped with autofocus to ensure clear facial feature capture. At the software level, a synchronous acquisition control algorithm is developed, achieving millisecond-level alignment between the voice signal and facial image data through a timestamp synchronization mechanism, avoiding subsequent sentiment analysis deviations due to data asynchrony. Simultaneously, the acquired raw data undergoes preprocessing, including noise reduction, pre-emphasis, and frame-by-frame windowing of the voice signal, as well as face detection, key point localization, and image enhancement processing of the facial image, providing high-quality input data for subsequent modules.

[0022] 2. Voiceprint Feature Extraction Module: The acoustic feature extraction unit uses Mel-frequency cepstral coefficients (MFCC) and linear predictive cepstral coefficients (LPCC) to extract spectral features. It obtains the fundamental frequency curve through a pitch detection algorithm (such as the YIN algorithm), identifies formant parameters using a formant extraction algorithm (such as the LPC method), and analyzes speech rate and rhythm features using a dynamic time warping (DTW) algorithm to form a multi-dimensional acoustic feature vector. The voiceprint model construction unit employs a lightweight deep neural network (such as an architecture combining convolutional neural networks (CNN) and recurrent neural networks (RNN). Initial model training can be completed with only ≤30 seconds of speech samples. The model uses an incremental learning strategy, continuously receiving new user speech data and dynamically updating network parameters to achieve real-time optimization of the voiceprint model, ensuring its adaptability to changes in user timbre. The voiceprint similarity evaluation unit uses a method combining cosine similarity and dynamic time warping (DTW) to quantitatively calculate the matching degree between the cloned speech and the original speaker's speech in the acoustic feature vector. A similarity threshold (such as 90%) is set. When the similarity of the cloned speech falls below the threshold, the model optimization process is triggered. The multi-speaker voiceprint management unit constructs a distributed voiceprint model library. It adopts a method of binding user identity identifiers (such as account, device ID) with voiceprint models, supports parallel storage and fast retrieval of multi-user voiceprint models, and in multi-person dialogue scenarios, it identifies the current speaker's identity through a real-time voiceprint matching algorithm and automatically switches to the corresponding voiceprint model.

[0023] 3. Emotional Semantic Fusion Module: The speech emotion analysis unit employs a deep learning model (such as a CNN-LSTM network) to extract features from the preprocessed speech signal, including short-time energy, zero-crossing rate, and fundamental frequency change rate. It identifies emotion types (such as joy, anger, sadness, and neutrality) using a trained classification model and quantifies emotion intensity based on an emotion intensity regression model (e.g., a 0-100 score scale). Simultaneously, it analyzes speech rate and intonation variation patterns to determine emotional expression patterns. The facial expression recognition unit uses facial feature point detection technology (such as a 68-point feature point model) to extract motion parameters from key facial regions (such as eyebrows, eyes, and mouth). Combined with a classification model trained on an emotional expression database, it identifies the emotional tendencies corresponding to micro-expressions. Temporal analysis algorithms capture dynamic changes in facial expressions, improving the accuracy of emotion recognition. The emotional conflict processing unit constructs a contextual weight allocation model. When there is a difference between vocal emotion and facial expression emotion (e.g., the vocal expression is neutral but the facial expression shows joy), it combines the dialogue history and scene information provided by the dynamic context module to calculate the credibility weight of the two emotional information. For example, in formal business scenarios, vocal emotion has a higher weight, while in everyday communication scenarios, facial expression emotion has a higher weight, ultimately fusing them to form a unified emotional feature. The emotion-semantic binding unit uses an attention mechanism to associate emotional features (emotion type, intensity, expression pattern) with semantic content at the word and sentence levels. For example, for sentences expressing "anger," relevant emotional markers are enhanced in the semantic representation to form an emotion-semantic fusion representation.

[0024] 4. Dynamic Context Module: Employing a hierarchical context model architecture, this module constructs contextual information across four dimensions. The dialogue history dimension uses Dialogue State Tracking (DST) technology to record the historical text, speech, and emotional information of the current dialogue. A Recurrent Neural Network (RNN) is used to encode this historical information, forming a historical context vector. The user feature dimension integrates user language preferences (e.g., common expressions, domain background), identity information (e.g., age, occupation), and usage habits (e.g., common target languages, rule settings) through a user profiling system to construct a user feature vector. The scene information dimension uses scene recognition algorithms, combined with device sensor data (e.g., location information, time information) and dialogue topic analysis results, to determine the current scene type (e.g., business meetings, daily communication, medical consultations), forming a scene feature vector. The domain label dimension uses keyword matching and domain classification models to identify the professional domains involved in the dialogue (e.g., law, medicine, finance) and assign corresponding domain labels. Contextual information from various dimensions is integrated into a unified contextual model through feature fusion algorithms (such as attention fusion and splicing fusion). This model has the ability to be dynamically updated and can adjust contextual parameters in real time according to new dialogue content, user behavior, and scene changes, providing accurate contextual support for subsequent translation and rule matching.

[0025] 5. Cross-language translation engine: Built on the Transformer architecture, it adopts an encoder-decoder structure and adds a sentiment marker preservation mechanism to traditional semantic translation. The encoder encodes the sentiment-semantic fusion representation and captures the correlation between semantic content and sentiment features through a multi-head attention mechanism, ensuring that sentiment information is not lost during the encoding process. The decoder introduces a sentiment constraint module when generating target language text. Based on the sentiment markers transmitted by the encoder, it adjusts the word selection and sentence structure of the generated text. For example, when expressing "joy," it prioritizes words with positive connotations, and when expressing "sadness," it uses sentence structures that conform to the sentiment expression habits of the target language. At the same time, the engine optimizes the training strategy for low-resource languages, using transfer learning to initialize the low-resource language model using the pre-trained model parameters of high-resource languages, reducing the impact of insufficient low-resource language data on translation accuracy. In terms of real-time optimization, it uses model quantization and pruning techniques to reduce the computational complexity of the model and ensure that the translation latency is controlled within a few hundred milliseconds to meet the needs of real-time communication.

[0026] 6. Custom Language Rule Engine: The rule management unit provides a visual rule configuration interface, allowing users to define rule types (keyword replacement, sensitive word filtering, code word mapping, domain terminology conversion) through graphical operations. Users can also set the applicable scenarios (e.g., business scenarios, private scenarios), applicable objects (e.g., specific contacts, all objects), sentiment associations (e.g., triggering only under specific sentiments), and priorities (levels 1-10, with higher levels indicating higher priority). Rule data is stored in a structured database, supporting the addition, modification, deletion, and export of rules. The rule parsing unit uses a syntax analyzer to convert user-defined rules into machine-executable logical expressions, parsing the triggering conditions (e.g., containing specific keywords, satisfying specific contexts) and execution actions (e.g., replacing with specified words, deleting sensitive words, converting to domain terminology), generating a rule parsing tree. The context matching unit, based on the context model provided by the dynamic context module, calculates the matching degree between the current context and the applicable conditions of each rule. A weighted scoring algorithm is used, considering factors such as scenario adaptability, object adaptability, and sentiment adaptability, to generate a rule matching score. The rule execution unit invokes rules in descending order of matching scores to process the target language text. During execution, a semantic coherence check mechanism is introduced, using a language model to calculate the semantic similarity between the text before and after replacement, ensuring the replacement result does not deviate from the original semantics. The rule conflict resolution unit constructs a conflict detection model. When multiple rules simultaneously meet the triggering conditions, they are first filtered based on rule priority. If priorities are the same, the unit coordinates with contextual matching and user preferences (such as user historical selection records) to select the optimal rule for execution. Simultaneously, the conflict resolution result is recorded, providing data support for subsequent rule optimization.

[0027] It's worth noting that several key configuration dimensions can be added, introducing an "emotional threshold trigger" mechanism that allows users to set the emotional intensity range for rules to take effect, such as "automatically replace aggressive words when an anger intensity of ≥70 is detected." A "time dimension constraint" function can also be developed to support rules activating by time period; for example, business users can set "activate formal language conversion rules from 9:00 to 18:00 on weekdays" to meet communication needs at different times.

[0028] A new "Semantic Relevance" parameter configuration has been added, allowing users to set keyword contextual conditions that must be met for a rule to be triggered. For example, defining "trigger the financial terminology conversion rule only when 'cost' and 'budget' appear simultaneously and the semantic relevance is ≥85%" avoids false triggers caused by isolated keywords. A "Language Style Mapping" function has also been introduced, supporting the conversion of specific styles of the source language (such as humorous, serious, or euphemistic) into corresponding styles in the target language through rules, ensuring consistency in tone and style during cross-language communication.

[0029] 7. Real-time Voiceprint Cloning Engine: Built on end-to-end speech synthesis technology, its core is a personalized speech synthesis model. During model training, the voiceprint feature vectors provided by the personalized voiceprint model are used as constraints. Combined with a large amount of general speech data and a small amount of user speech data, a transfer learning method is employed to train the voiceprint cloning model. The model architecture uses Transformer-TTS. The encoder converts the target language text into text feature vectors, the decoder combines the voiceprint feature vectors to generate a speech spectrogram, and then the vocoder converts the spectrogram into a waveform speech signal to form the basic cloned speech. To improve real-time performance, the engine employs lightweight model technology, performing layer pruning and parameter quantization on the Transformer model, while optimizing the inference process to achieve rapid conversion from text input to speech output. For speech quality optimization, a speech enhancement module is introduced to reduce noise and distortion in the synthesized speech, improving its naturalness and clarity.

[0030] 8. Emotion Transfer Module: The emotion feature parameterization unit converts the emotional expression of the source language speech into quantifiable acoustic parameters, including pitch (mean fundamental frequency, standard deviation), speech rate (syllables / second), volume (mean short-time energy), and prosody (amplitude of pitch variation). These parameters are accurately calculated using an emotion parameter extraction algorithm. The cross-language emotion mapping unit, based on a multilingual emotion acoustic database, constructs a mapping table of emotion parameters between the source and target languages. For example, when expressing "anger," the mean fundamental frequency in Chinese may be higher than in English. The mapping table converts the fundamental frequency parameters of the source language into those that conform to the emotional expression habits of the target language, ensuring that emotion is not distorted during cross-language transmission. The emotion-voiceprint coordination unit constructs an acoustic parameter coordination model, analyzes the voiceprint features of the cloned speech (such as formant frequencies and timbre features), and adjusts the emotion parameters to ensure that the emotional expression and voiceprint features are acoustically consistent. For example, for a voiceprint with a soft timbre, the volume parameter is appropriately reduced when expressing "anger" to avoid the speech sounding harsh. The emotional intensity adaptive unit combines the emotional and cultural characteristics of the target language and fine-tunes the emotional parameters through an emotional intensity adjustment algorithm. For example, when expressing the emotion of "joy", the emotional intensity of Latin-based languages ​​is usually higher than that of East Asian languages. The unit will adjust parameters such as speech rate and pitch variation according to the cultural background of the target language to make the emotional expression more in line with the cognitive habits of the target language users.

[0031] 9. Cultural Adaptation Module: Built upon a cultural semantic database and context-aware technology, this module focuses on the cultural adaptation of emotional expression and semantic representation. Regarding emotional expression cultural adaptation, it establishes an emotion-culture mapping model by collecting data on emotional expression habits in different languages. For example, when expressing "apology," Japanese often uses euphemisms, while English often uses direct sentences. The module adjusts the sentence structure and emotional intensity of the translated text based on the cultural characteristics of the target language, while also optimizing the parameters of the emotion transfer module to ensure that the emotional expression of the target language speech conforms to local cultural habits. Regarding semantic representation cultural adaptation, a cultural semantic knowledge base is constructed, including taboo words, slang, and common expressions from different cultures. When the translated text contains content that may cause cultural misunderstandings, the module automatically replaces it with expressions that conform to the target language culture. Simultaneously, the module combines scene information and domain tags from the dynamic context module to adjust the formality and cultural adaptation strategies of the semantic expression for different scenarios, ensuring that the translation results are natural and appropriate in the target cultural environment.

[0032] 10. Multimodal Output Module: Employing a multi-channel output architecture, this module supports simultaneous output of target language speech and text. The speech output stage integrates a high-performance audio processing chip to optimize the sound quality of the target language speech by incorporating voiceprint features and emotional information. This includes noise reduction, equalizer adjustment, and adaptive volume control, ensuring consistent playback quality across different devices (such as headphones and speakers) and environments. It also supports flexible configuration of the speech output format, allowing users to select different sampling rates and bit rates according to their needs. The text output stage features a real-time text display module that supports synchronized scrolling of text and speech. The text retains emotional markers (such as color and symbols indicating emotional type and intensity) for intuitive understanding of the emotional content. Customizable text formatting is also supported, allowing users to set font, size, color, and other display parameters to meet individual preferences. The module also features output synchronization control, using a timestamp calibration mechanism to ensure precise synchronization between speech playback and text display, preventing audio-text desynchronization.

[0033] 11. Adaptive Optimization Module: Built on user feedback and data-driven principles, this module enables dynamic optimization of all system modules. First, it establishes a user feedback collection mechanism. Through a simple interactive interface (such as star ratings, tag selection, and text input), it collects user feedback on translation results (semantic accuracy, emotional delivery, cultural adaptation), speech quality (naturalness, similarity), and the effectiveness of rule application. Simultaneously, it records user behavior data (such as rule trigger frequency, manual content modification, and target language selection). Then, using reinforcement learning algorithms and user satisfaction as the reward function, it optimizes the parameters of each module. For example, based on user feedback on emotional delivery, it adjusts the weight allocation model of the emotional semantic fusion module and the parameter mapping table of the emotional transfer module; based on user feedback on rule application, it optimizes the rule matching algorithm and conflict resolution strategy of the custom language rule engine; and based on user feedback on translation accuracy, it updates the model parameters of the cross-language translation engine. Furthermore, the module employs an incremental learning strategy, periodically analyzing collected user feedback and usage data to dynamically adjust the operating logic of each module, ensuring continuous system performance optimization and better meeting user needs.

[0034] Example 2: A smart real-time language translation method, comprising the following steps: Step S1: Data Acquisition and Preprocessing First, the microphone array and camera of the multimodal acquisition module are activated. The microphone array uses beamforming technology to focus on the speaker, acquiring a speech signal with a 16kHz sampling rate and 24-bit quantization depth. The camera acquires facial image data at 1080P resolution and 30 frames per second. A timestamp synchronization algorithm is used to align the speech signal and facial image data at the millisecond level to ensure data synchronization. Then, the speech signal is preprocessed: an adaptive noise suppression algorithm is used to remove environmental noise, pre-emphasis processing is used to enhance the energy of high-frequency signals, and Hanning windows are used to segment the speech signal into frames (frame length 20-30ms, frame shift 50%) to reduce inter-frame redundancy. The image data is preprocessed: the MTCNN algorithm is used for face detection to locate the facial region, a 68-point feature point detection algorithm is used to extract key facial feature points, and image enhancement algorithms (such as histogram equalization) are used to improve the clarity of facial features, providing high-quality data for subsequent modules.

[0035] Step S2: Voiceprint Feature Extraction and Model Construction The acoustic feature extraction unit extracts features from the preprocessed speech signal, calculating Mel-frequency cepstral coefficients (MFCC, typically 12-16 dimensions) and linear predictive cepstral coefficients (LPCC). It extracts the fundamental frequency curve using the YIN algorithm, identifies the first 3-5 formant parameters using the LPC method, and analyzes speech rate and rhythm features using the Dynamic Time Warping (DTW) algorithm, forming a multi-dimensional acoustic feature vector. If the user is using the service for the first time, the voiceprint model construction unit uses a lightweight CNN-RNN network, training an initial personalized voiceprint model with speech feature vectors ≤30 seconds long as training data. If the user is using the service again, the model employs an incremental learning strategy, inputting newly extracted acoustic feature vectors into the existing model to update network parameters, achieving dynamic optimization of the voiceprint model. The voiceprint similarity evaluation unit calculates the similarity between the currently extracted voiceprint features and the existing model. If the similarity is lower than a preset threshold (e.g., 90%), a model retraining process is triggered to ensure model accuracy.

[0036] Step S3: Construction of Emotion-Semantic Fusion Representation The speech sentiment analysis unit inputs the preprocessed speech signal into a CNN-LSTM model, extracting features such as short-time energy, zero-crossing rate, and fundamental frequency change rate. The model outputs emotion type (joy, anger, sadness, neutral, etc.) and emotion intensity (0-100 points), while simultaneously analyzing speech rate and intonation variation patterns to determine emotion expression patterns. The facial expression recognition unit processes the preprocessed facial image data, extracting motion parameters of eyebrows, eyes, and mouth using a 68-point feature point model. This data is input into an emotional expression classification model to identify the emotional tendency corresponding to micro-expressions. A temporal analysis algorithm is combined to capture dynamic changes in facial expressions and obtain emotional features. The emotion conflict processing unit calls upon the initial context information (such as scene type and dialogue topic) from the dynamic context module to calculate the credibility weights of speech emotion and facial emotion. For example, in a medical consultation scenario, the speech emotion weight is set to 0.7, and the facial emotion weight is set to 0.3, fusing them to form a unified emotional feature. The emotion-semantic binding unit uses an attention mechanism to associate emotional features with the semantic content output by the semantic analysis module, adding emotional markers to the semantic representation to form an emotion-semantic fusion representation.

[0037] Step S4: Construction and Update of Dynamic Context Model The dialogue history dimension records the historical text, voice, and emotional information of the current dialogue using Dialogue State Tracking (DST) technology, and encodes the historical information using an RNN to generate a historical context vector. The user feature dimension retrieves user language preferences, identity information, and usage habits from the user profiling system to construct a user feature vector. The scene information dimension obtains location information from the device's GPS and time information from the system time, and combines this with dialogue topic analysis results (e.g., identifying the current discussion content as "business negotiation" through keywords) to determine the scene type and generate a scene feature vector. The domain label dimension identifies the professional field involved in the dialogue using a domain classification model (e.g., a classification model based on support vector machines) and assigns corresponding domain labels (e.g., "finance" or "law"). Subsequently, an attention fusion algorithm integrates the historical context vector, user feature vector, scene feature vector, and domain label vector into a unified context model. The context parameters are updated in real time based on newly generated dialogue content, user actions (e.g., switching target languages), and scene changes (e.g., moving from indoors to outdoors) to ensure that the context model remains consistent with the current communication scene, providing accurate support for subsequent translation, rule matching, and other processes.

[0038] Step S5: Cross-linguistic semantic translation and sentiment marker preservation The cross-language translation engine receives information from the sentiment-semantic fusion representation and the output of the dynamic context model. First, the encoder encodes the sentiment-semantic fusion representation, with its multi-head attention mechanism focusing on the association between semantic content and sentiment markers to ensure no sentiment information is lost during encoding. When generating target language text, the decoder adjusts the translation strategy based on scene information and domain labels from the dynamic context model. For example, in a "medical consultation" scenario, it uses precise translations for professional medical terms, while in a "daily communication" scenario, it uses more colloquial expressions. Simultaneously, the sentiment constraint module in the decoder preserves sentiment-related features in the target language text based on the sentiment markers transmitted by the encoder. For instance, for sentences expressing "anger" in the source language, the target language text uses words and phrases with strong emotional inclinations and embeds implicit sentiment markers (such as using specific symbols to indicate emotion type and intensity) to provide a basis for subsequent sentiment transfer. After translation, the engine performs semantic accuracy verification on the target language text, correcting translation deviations by comparing it with a domain terminology database and a common expression database to ensure the accuracy of the translation results.

[0039] Step S6: Applying Custom Language Rules Step S6 follows the logical flow of "parsing-matching-execution-verification", specifically including: S61. The rule parsing unit performs word segmentation and keyword extraction on the target language text, identifies words or expressions that may trigger custom rules, such as extracting words like "confidential project" and "personal mobile phone number" from the text, and determines whether they are related to the user's preset rules. S62. The context matching unit calls information such as scene type (e.g., "business meeting"), dialogue object (e.g., "external partner"), and emotional state (e.g., "neutral") from the dynamic context model to analyze the matching degree between the current context and the applicable conditions of each custom rule. It uses a weighted scoring algorithm to calculate the matching score of each rule. For example, the user presets the rule "In the business meeting scenario, replace 'confidential project' with 'specific cooperation matter'" and it has a high matching score in the current scenario. S63. Select suitable rules from high to low matching scores. If there are multiple rules with high matching scores, proceed to the rule conflict resolution stage. S64. The rule conflict resolution unit first checks the rule priority and selects the rule with higher priority. If the priorities are the same, it further filters the rules by combining user preferences (such as users' history of preferring privacy-protecting rules) and contextual suitability, and finally determines the only rule to be executed. S65. The rule execution unit performs replacement processing on the target language text according to the selected rules, such as replacing "confidential project" with "specific cooperation matters" and replacing the identified personal mobile phone number with "[privacy information]"; S66. After execution, the semantic similarity between the text before and after replacement is calculated by the language model. If the similarity is lower than the preset threshold (such as 95%), the replacement result is fine-tuned to ensure that the replaced text is semantically coherent with the context and does not change the original core semantics and sentiment.

[0040] Step S7: Basic cloned speech generation The real-time voiceprint cloning engine receives target language text that has undergone rule-based replacement and a personalized voiceprint model. First, it converts the target language text into a text feature vector. The text preprocessing module performs word segmentation, part-of-speech tagging, and prosodic prediction to determine the pronunciation rhythm and pauses. Then, the encoder in the engine fuses the text feature vector with the voiceprint feature vector output by the personalized voiceprint model to generate a fused feature vector. The decoder generates a speech spectrogram based on the fused feature vector. The spectrogram generation process strictly follows the timbre features (such as formant frequencies and fundamental frequency range) in the voiceprint model to ensure that the generated speech possesses the timbre characteristics of the original speaker. Finally, the vocoder converts the spectrogram into a waveform speech signal to form the basic cloned speech. Simultaneously, the speech enhancement module performs noise reduction processing on the basic cloned speech, removing noise and distortion generated during synthesis to improve the naturalness and clarity of the speech, ensuring that the speech quality meets the needs of real-time communication.

[0041] Step S8: Emotional Transfer Step S8 revolves around "parameterization transformation - coordination - adaptation", and specifically includes: S81. The Emotional Feature Parameterization Unit extracts acoustic parameters from the source language speech and calculates parameters such as pitch (mean fundamental frequency, standard deviation of fundamental frequency), speech rate (syllables per second), volume (mean short-time energy), and prosody (amplitude of pitch variation). For example, when the source language speech expresses the emotion of "joy", the mean fundamental frequency is 250Hz, the speech rate is 5 syllables / second, and the mean volume is -20dB. S82. The cross-language emotion mapping unit calls the mapping table in the multilingual emotion acoustic database to convert the source language emotion parameters into the corresponding emotion parameters of the target language. For example, when the target language is English, the fundamental frequency average value corresponding to the emotion of "joy" is usually 230Hz and the speech rate is 6 syllables / second. This unit will adjust the fundamental frequency parameter of the source language from 250Hz to 230Hz and the speech rate from 5 syllables / second to 6 syllables / second. S83. The Emotion-Voiceprint Coordination Unit analyzes the voiceprint features of the basic cloned speech. For example, the formant frequency of the cloned speech is low (which manifests as a thick timbre). If the adjusted emotion parameters are applied directly, the speech may be stiff. This unit will further fine-tune the emotion parameters, adjusting the average volume from -20dB to -22dB to make the emotion expression more coordinated with the voiceprint features. S84. Enhance emotionally dense segments (such as sentences in the text with strong emotional tendencies) by increasing the range of pitch changes and accelerating the rate of speech changes to highlight key emotional information and ensure that the audience can clearly perceive the core emotion. S85. Based on the target language cultural background information provided by the cultural adaptation module, the emotional parameters are fine-tuned. For example, when the target language is Japanese, the expression of "joy" is usually more subtle. This unit will adjust the speech rate from 6 syllables / second to 5.5 syllables / second to reduce the intensity of the emotional expression and make it more in line with the emotional expression habits in Japanese culture.

[0042] Step S9: Cultural Adaptation Adjustment The cultural adaptation module first utilizes a cultural semantic knowledge base to perform cultural compatibility testing on the target language speech and text after emotion transfer, identifying content that may cause cultural conflict. Regarding semantic expression, if the text contains taboo words in the target language culture (such as expressions related to pork in an Arabic context), the module automatically replaces them with neutral words that conform to local cultural habits. If the text contains slang or idiomatic expressions, the module adjusts the expression according to regional differences in the target language (such as British versus American English) to ensure the translation is closer to the language habits of local users. In terms of emotional expression, the module further optimizes the emotional parameters of the speech based on the emotional expression preferences of the target language culture. For example, in German culture, emotional expression in business communication is usually restrained; the module reduces the amplitude of volume changes and the rate of speech rate changes to make the emotional expression more in line with the atmosphere of German business culture. Simultaneously, the module adjusts the sentence structure of the text, such as converting commonly used active sentences in Chinese to passive sentences commonly used in Japanese, improving the naturalness and acceptability of the translation in the target language culture.

[0043] Step S10: Multimodal output The multimodal output module first performs final sound quality optimization on the target language speech. It adjusts the frequency response of the speech using an equalizer to make the sound more balanced across different frequency bands. Through a volume adaptive control algorithm, it adjusts the speech playback volume based on the current ambient noise level (e.g., by detecting ambient noise through a microphone array) to ensure that users can hear the speech clearly even in noisy environments. For text output, the module converts the text and implicit sentiment markers into a visual form. For example, it uses red to mark "anger" and blue to mark "sadness," and uses the number of asterisks to indicate the intensity of the emotion (e.g., "anger"). (This indicates a high level of emotion) and supports synchronized scrolling of text and speech. When the speech plays a certain sentence, the corresponding sentence in the text will be highlighted, allowing users to obtain both audio and text information simultaneously. In addition, the module supports customizable output formats. Users can select the audio playback device (such as headphones or speakers), and the font and color of the text display through the settings interface to meet personalized usage needs.

[0044] Step S11: Adaptive Optimization The adaptive optimization module collects user feedback through an interactive interface. Users can rate the translation results by star rating (1-5 stars) for aspects such as semantic accuracy, emotional delivery, cultural adaptability, and voice quality. They can also provide specific suggestions for improvement via text input (e.g., "The emotional expression here is too strong" or "The terminology translation is inaccurate"). Simultaneously, the module automatically records user behavior data, including rule triggering frequency (e.g., the number of times a privacy protection rule is triggered), manually modified content (e.g., changing "cooperation matters" to "cooperation projects" in the translated text), and target language selection preferences. The module employs a reinforcement learning algorithm, using user satisfaction as the core reward function to optimize the parameters of each module: if user feedback indicates poor emotional delivery, the module adjusts the weighting ratio of voice and facial expression in the emotional semantic fusion module, as well as the parameter mapping table in the emotion transfer module; if users frequently manually modify the translation results of a certain type of terminology, the module updates the domain terminology database in the cross-language translation engine to improve the translation accuracy of that type of terminology. In addition, the module will periodically analyze the collected data to identify changes in user habits (such as users frequently switching to Spanish recently) and dynamically adjust the default parameters of each module (such as setting the default target language to Spanish) to ensure that the system performance continues to adapt to user needs.

[0045] Steps S01-S05: Adaptation to multi-person dialogue scenarios In multi-person dialogue scenarios, the system will perform the following additional steps: S01. The voiceprint feature extraction module uses a real-time voiceprint matching algorithm to identify the speaker in the collected speech signal, determine the identity of the current speaker, and call the corresponding personalized voiceprint model and custom rule set from the multi-speaker voiceprint management unit (e.g., user A's rule set focuses on privacy protection, while user B's rule set focuses on domain terminology conversion). S02. The dynamic context module maintains an independent emotional expression pattern library (recording user A's commonly used emotional expression parameters) and context sub-model (recording user A's dialogue history and language preferences) for each speaker, ensuring that each speaker's personalized needs can be met; S03. The cross-language translation engine, voiceprint cloning engine, and emotion transfer module adjust strategies based on the speaker's identity. For example, a more formal translation style and restrained emotion transfer strategy are adopted for user A (business person), while a more colloquial translation style and lively emotion transfer strategy are adopted for user B (daily communication user). S04. In the voice output stage, by adjusting the distinguishability parameters of the voiceprint features (such as increasing the fundamental frequency difference of cloned voices of different speakers), it is ensured that listeners can clearly distinguish the voices of different speakers; in the text output stage, by adding speaker identifiers (such as "[User A]:") before the text, the ownership of the text is clarified, while retaining the emotional markers of each speaker to ensure the personalization of emotional expression; S05. The system records the history of emotional interactions in multi-person dialogues (e.g., after user A expresses "dissatisfaction", user B's emotional response is "appeasement") and the application of rules (e.g., a certain rule is frequently triggered in user A's dialogue). This data is stored in the user profile system and the rule optimization database to provide more accurate personalized support for subsequent communication. For example, in the next multi-person dialogue, the system can optimize the emotional transfer strategy in advance based on the historical emotional interaction pattern.

[0046] Example 3: A smart real-time language translation terminal specifically includes the following structure: 1. Processor: Employs a high-performance multi-core processor, specifically the Qualcomm Snapdragon 8 Gen3. This processor is based on a 4nm process and features one 3.3GHz Cortex-X4 super core, three 3.2GHz Cortex-A720 large cores, and four 2.0GHz Cortex-A520 small cores. Its core function is to act as the control center of the terminal, responsible for scheduling the operation of various modules, including receiving data from the multimodal acquisition device, retrieving programs and data from memory, controlling the voiceprint-emotion processing chip for calculations, and driving the multimodal output module for data output. In terms of operating principles, the processor establishes communication with other hardware modules through a bus interface, receives interrupt requests and data requests from each module, and allocates computing resources according to preset program logic. For example, during real-time translation, the processor prioritizes allocating computing resources to the cross-language translation engine and the voiceprint cloning engine to ensure the real-time performance of translation and speech synthesis. Simultaneously, the processor runs the operating system and applications, providing an interactive interface for the user and processing user commands (such as setting custom rules and selecting the target language).

[0047] 2. Memory: A combined RAM+ROM+SSD architecture is adopted: RAM uses LPDDR5X with a capacity of 16GB, primarily used for temporary storage of data during system operation (such as collected voice signals, generated translated text, and contextual model parameters). Its high-speed read / write characteristics ensure efficient data transmission between modules, avoiding real-time performance impacted by data latency. ROM uses UFS4.0 with a capacity of 512GB, used to store computer programs (such as the operating system, drivers for each module, and model files for the cross-language translation engine), user-defined language rule bases (stored in a structured database format, supporting fast querying and modification), and personalized voiceprint model libraries (storing each user's voiceprint feature vector and model parameters). The memory works by connecting to the processor via the address and data buses. The processor reads or writes data from the memory by sending address and control signals. For the user-defined language rule base and personalized voiceprint model library, the memory employs a partitioned storage strategy, isolating them from the system program partitions, and also supports data encryption to ensure user privacy is not leaked.

[0048] 3. Multimodal Acquisition Device: Composed of a microphone array and a camera. The microphone array uses Knowles' SPK0641HT4H-B, a 4-channel omnidirectional microphone supporting a 16kHz sampling rate and 24-bit quantization depth. Its built-in noise suppression algorithm effectively filters ambient noise, and beamforming technology focuses the speaker's voice signal within a 3-meter range, reducing interference from other directions. The camera uses Sony's IMX586, a 1 / 1.7-inch CMOS sensor with a resolution of 48 megapixels, supporting a 30fps frame rate and autofocus. Its built-in face detection algorithm quickly locates facial areas in the image, ensuring accurate acquisition of facial expression data. The working principle of the multimodal acquisition device is as follows: the microphone array simultaneously acquires voice signals through four microphones, converts the analog signals into digital signals, and then transmits them to the processor through the I2S interface; the camera captures facial images through a CMOS sensor, converts the image signals into digital signals, and then transmits them to the processor through the MIPI interface; at the same time, the device's built-in synchronization clock module adds timestamps to the voice and image signals to ensure that the processor can achieve accurate data alignment.

[0049] 4. Voiceprint-Emotion Processing Chip: Rockchip's RK3588 is selected. This chip integrates a dedicated Neural Processing Unit (NPU) with a computing power of up to 6 TOPS, specifically designed to accelerate voiceprint feature extraction and emotion transfer calculation. Its core functions include: rapidly extracting acoustic features (such as MFCC and fundamental frequency curve) from speech signals, achieving a speed improvement of 3-5 times compared to using processor-only calculations; accelerating the calculation process of emotion feature parameterization and cross-language emotion mapping, ensuring that the latency of the emotion transfer stage is controlled within 100 milliseconds; and supporting multi-task parallel processing, capable of simultaneously handling voiceprint feature extraction and emotion calculation tasks for multiple speakers. In terms of working principle, the voiceprint-emotion processing chip is connected to the processor through the PCIe interface. After the processor transmits the collected voice data to the chip, the chip calls the built-in acceleration algorithm (such as the optimized CNN-LSTM model) to perform calculations and feeds back the calculation results (such as voiceprint feature vectors and emotion parameters) to the processor. The processor then distributes the data to the corresponding modules (such as the voiceprint feature extraction module and the emotion transfer module). Hardware acceleration significantly reduces the processor's computational burden and improves the overall real-time performance of the system.

[0050] 5. Input Interaction Device: Composed of a touchscreen and physical buttons: The touchscreen uses a 7-inch AMOLED screen from BOE (Beijing Oriental Electronics Co., Ltd.), with a resolution of 2560×1440. It supports multi-touch and handwriting input. Users can operate the interactive interface through the touchscreen, such as defining custom rules (clicking the "Add Rule" button, entering keywords and replacement content), selecting the target language (clicking from the language list), and submitting user feedback (clicking the star rating or entering text). The physical buttons include a power button, volume control buttons, and an emergency pause button. The power button is used to turn the terminal on / off, the volume control buttons are used to control the volume of the output voice, and the emergency pause button is used to quickly pause the translation output in an emergency. The working principle of the input interaction device is as follows: The touchscreen detects the user's touch operation through capacitive sensing technology, converts the touch coordinates into electrical signals, and transmits them to the processor through the I2C interface; the physical buttons generate electrical signals through the on / off state of mechanical contacts, which are transmitted to the processor through the GPIO interface; after receiving these signals, the processor parses the user's operation instructions and calls the corresponding program module to execute the operation (such as opening the rule management interface or adjusting volume parameters).

[0051] 6. Audio Output Device: Composed of a speaker and a headphone jack. The speaker is a Harman stereo speaker with a frequency response range of 20Hz-20kHz, a maximum output power of 2W×2, and supports Dolby audio, providing clear and full voice playback. The headphone jack is a standard 3.5mm jack, supporting external wired headphones and Bluetooth 5.3 protocol, allowing connection to wireless Bluetooth headphones to meet users' audio output needs in different scenarios (such as using headphones in quiet environments to avoid disturbing others). The working principle of the audio output device is as follows: The processor transmits the target language speech data (digital signal) after emotional transfer and cultural adaptation to the audio codec (such as Texas Instruments' TLV320AIC3254). The audio codec converts the digital signal into an analog signal and transmits it to the speaker or to external headphones via the headphone jack. Simultaneously, the audio codec supports volume control, adjusting the amplitude of the output analog signal according to the volume adjustment commands sent by the processor to increase or decrease the volume.

[0052] 7. Communication Module: Supports dual-mode Wi-Fi 6 and 5G communication. The Wi-Fi chip uses Qualcomm's QCA6391, supporting both 2.4GHz and 5GHz bands with a maximum transmission rate of 1.2Gbps, suitable for stable indoor network environments. The 5G chip uses Qualcomm's SDX75, supporting Sub-6GHz and millimeter-wave bands with a maximum download rate of 3.8Gbps, suitable for outdoor mobile scenarios. Its core functions include: data interaction with the cloud server, uploading updated personalized voiceprint model data, custom rule modification records, and usage behavior data; downloading optimized translation model parameters, cultural semantic knowledge base update packages, and multilingual emotional acoustic database upgrade files from the cloud; in multi-person dialogue scenarios, it supports real-time data synchronization between multiple terminals, such as synchronizing speaker voiceprint data and emotional features collected by one terminal to other participating terminals to ensure consistency in translation results across terminals; and it supports remote firmware upgrades, allowing users to receive and install system firmware updates directly through the communication module without manual operation, improving system performance and security.

[0053] The communication module works as follows: In Wi-Fi mode, it receives radio frequency signals from the wireless router via an antenna. After demodulation and decoding by the Wi-Fi chip, the signals are converted into digital signals and then transmitted to the processor via the PCIe interface. In 5G mode, the 5G chip receives radio frequency signals from the base station via an antenna. After baseband processing and signal demodulation, the signals are converted into digital signals and transmitted to the processor via the PCIe interface. The processor parses and verifies the received cloud data. If it is model parameters or knowledge base update files, it stores them in the corresponding partition of the memory and triggers relevant modules (such as the cross-language translation engine and cultural adaptation module) to load the updated resources. If it is multi-terminal synchronized data, it distributes it to the corresponding modules (such as the multi-speaker voiceprint management unit and dynamic context module) to ensure that the data of each terminal is consistent. At the same time, the communication module supports encrypted data transmission, using the SSL / TLS protocol to encrypt the data uploaded to the cloud to prevent data from being stolen or tampered with during transmission and to protect user privacy and security.

[0054] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and not restrictive.

[0055] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style of the specification is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. An intelligent real-time language translation system, characterized in that: include: A multimodal acquisition module is used to simultaneously acquire the user's voice signals and facial expression data; The voiceprint feature extraction module is used to extract the speaker's unique acoustic features from the speech signal, and to build and dynamically update a personalized voiceprint model. The emotion-semantic fusion module is used to extract emotional features from speech and facial expressions and fuse them with semantic content to form an emotion-semantic representation; The dynamic context module is used to build a context model that includes dialogue history, user characteristics, scene information, and domain labels; A cross-language translation engine used to accurately translate the semantics of a source language into a target language while preserving sentiment marker information; A custom language rules engine is used to automatically replace specific words or expressions according to user-defined rules, supporting code word mapping, privacy protection, and sentiment protection; The real-time voiceprint cloning engine generates basic speech with the original speaker's timbre characteristics based on the target language text and personalized voiceprint model. The emotion transfer module is used to transfer emotional features from the source language to the cloned speech of the target language, so that the emotional expression is consistent with the voiceprint features. The cultural adaptation module is used to adjust emotional expression and semantic representation according to the target language and cultural background. The multimodal output module is used to output target language speech and text that integrates voiceprint features, emotional information, and custom rule application results; The adaptive optimization module is used to optimize the parameters of each module based on user feedback.

2. The intelligent real-time language translation system according to claim 1, characterized in that: The custom language rule engine includes: The rule management unit supports user-defined rules such as keyword replacement, sensitive word filtering, code word mapping, and domain terminology conversion, and allows users to set the applicable scenarios, objects, sentiment associations, and priorities of the rules; The rule parsing unit is used to parse the triggering conditions and execution actions of rules; The context matching unit is used to determine the applicability of rules based on the current dialogue scenario, the audience, the topic, and the emotional state. The rule execution unit is used to execute the matching rules in the translation process to ensure that the replacement results are consistent with the semantics of the context. The rule conflict resolution unit coordinates multiple rules based on priority, applicability, and user preference when they match simultaneously.

3. The intelligent real-time language translation system according to claim 1, characterized in that: The voiceprint feature extraction module includes: The acoustic feature extraction unit is used to extract personalized voiceprint features such as spectral features, fundamental frequency curves, formants, and speech rate and rhythm. The voiceprint model building unit constructs an initial model based on a small number of speech samples and dynamically optimizes it through continuous learning. The voiceprint similarity assessment unit is used to quantitatively assess the similarity between cloned speech and the original speaker's speech; The multi-speaker voiceprint management unit is used to store and differentiate the voiceprint models of multiple speakers, and supports real-time switching.

4. The intelligent real-time language translation system according to claim 1, characterized in that: The emotion transfer module includes: The emotional feature parameterization unit transforms emotional expression into quantifiable acoustic parameters; A cross-linguistic emotion mapping unit establishes an acoustic parameter mapping relationship between the source language and the target language in terms of emotion expression; The emotion-voiceprint coordination unit ensures that the transferred emotion features and the cloned voiceprint features are acoustically consistent. The emotion intensity adaptive unit fine-tunes the emotion intensity according to the emotion expression habits of the target language.

5. The intelligent real-time language translation system according to claim 1, characterized in that: The sentiment semantic fusion module includes: The voice emotion analysis unit is used to identify the type and intensity of emotions through changes in tone, speech rate, and volume. The facial expression recognition unit is used to identify the emotional tendencies corresponding to micro-expressions by analyzing facial feature points. The emotional conflict processing unit is used to assign weights based on contextual information when there is a difference between voice emotion and facial expression emotion. The sentiment-semantic binding unit is used to establish the association between sentiment features and semantic content.

6. An intelligent real-time language translation method, characterized in that: Includes the following steps: S1. Synchronously collect the user's voice signal and facial expression data, and perform preprocessing; S2. Extract speaker voiceprint features from speech signals and construct or update personalized voiceprint models; S3. Analyze the emotional features and semantic content in speech and facial expressions, and construct an emotion-semantic fusion representation; S4. Based on dialogue history, user characteristics, and scenario information, construct and update a dynamic context model; S5. Translate the source language semantics into target language text while preserving sentiment markers; S6. Call the custom language rule engine, match applicable rules according to the current context, and perform automatic replacement on the target language text; S7. Generate basic cloned speech based on the replaced target language text and personalized voiceprint model; S8. Transfer the emotional features in the source language to the cloned speech of the target language to achieve coordinated expression of emotion and voiceprint; S9. Adapt the translation results to the target language and cultural context; S10. Output the target language speech and text, which integrate the results of fusing voiceprint features, emotional information, and custom rule applications; S11. Collect user feedback and optimize the parameters and models of each module.

7. The intelligent real-time language translation method according to claim 6, characterized in that: The process of applying custom rules in step S6 includes: S61. Parse the target language text and identify keywords or expressions that may trigger the rules; S62. Analyze the current contextual information, including scene type, dialogue partners, topic, and emotional state; S63. Match the custom rule that best fits the current context; S64. When a rule conflict is detected, the rule is filtered according to the preset priority and contextual adaptability; S65. Execute the selected rules to replace the target language text; S66. Check the semantic coherence of the replacement result with the context, and make fine adjustments if necessary.

8. The intelligent real-time language translation method according to claim 6, characterized in that: The emotional transfer process in step S8 includes: S81. Parameterize the emotional expression of the source language speech into acoustic features such as pitch, speech rate, and volume; S82. Based on cross-linguistic sentiment mapping rules, the source language sentiment parameters are converted into sentiment parameters suitable for the target language; S83. Adjust the acoustic parameters of the cloned speech to naturally integrate emotional expression with voiceprint features; S84. Enhance emotionally intense segments to ensure the accurate transmission of key emotional information; S85. Adjust the intensity and manner of emotional expression based on the target language and cultural background.

9. The intelligent real-time language translation method according to claim 6, characterized in that: In multi-person dialogue scenarios, the following steps are also included: S01. Real-time identification of different speakers, calling the corresponding voiceprint model and custom rule set; S02. Maintain an independent emotional expression pattern library and context sub-model for each speaker; S03. Dynamically adjust translation, voiceprint cloning, and emotion transfer strategies based on speaker identity and dialogue context; S04. Maintain the distinctiveness of the speaker's voiceprint features and the personalization of emotional expression in the translation results; S05. Record the history of emotional interactions and rule applications in multi-person dialogues to support subsequent communication.

10. An intelligent real-time language translation terminal, characterized in that: include: processor; The memory contains computer programs, a user-defined language rule base, and a personalized voiceprint model library; A multimodal acquisition device, including a microphone array and a camera, is used to simultaneously acquire speech and facial expressions; Voiceprint-emotion processing chip, used to accelerate voiceprint feature extraction and emotion transfer calculation; An input interaction device for users to define and manage custom language rules; Audio output device for playing translated speech that incorporates voiceprints, emotions, and custom rules; The communication module is used for data interaction with the cloud; When the computer program is executed by a processor, it implements the intelligent real-time language translation method according to any one of claims 6-9.