Real-time speech translation method and device, equipment and storage medium

By constructing a dynamic user memory model, the source terms in the identified text are matched and replaced in real time, solving the problem that existing machine translation systems cannot remember translation preferences. This achieves consistency and personalized adaptation in real-time speech translation, improving the accuracy and applicability of the translation.

CN121981129APending Publication Date: 2026-05-05GOERTEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GOERTEK INC
Filing Date
2025-12-25
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing mainstream machine translation systems cannot remember previously confirmed or corrected translation preferences when processing long dialogues or streaming text, which may lead to errors in subsequent translations of the same term, resulting in a lack of consistency.

Method used

By constructing a dynamic user memory model, the system can match and identify source terms in text and memory units in real time, determine target terms based on contextual tags and semantic similarity, and replace translation terms in general translation results, supporting user-defined translation preferences.

Benefits of technology

While ensuring real-time translation, it significantly improves the consistency and personalization of terminology translation, ensuring that user-defined terms are accurately replaced in specific scenarios, thereby improving the accuracy and applicability of translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121981129A_ABST
    Figure CN121981129A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of real-time translation, and discloses a real-time speech translation method and device, equipment and a storage medium, and the method comprises the steps: carrying out the streaming recognition of a received real-time speech, and obtaining a recognition text; matching the recognition text with the source terms in each memory unit, and determining a target term corresponding to the recognition text according to a matching result; obtaining a translation result of the recognition text; and replacing the translation terms corresponding to the source terms in the translation result with the target terms to obtain a replaced translated text, and outputting the translated text. Streaming recognition is carried out on real-time voice, recognition texts are matched with source terms in a memory unit in real time, and then target terms defined by a user are adopted to replace corresponding parts in a general translation result, so that the translation preference of the user is guaranteed on the basis of guaranteeing the real-time performance of translation, and the user experience is improved. And the term translation consistency and the personalized adaptation degree are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of real-time translation technology, and in particular to a real-time speech translation method, apparatus, device, and storage medium. Background Technology

[0002] Existing mainstream machine translation systems cannot remember previously confirmed or corrected translation preferences when processing long dialogues or streaming text, which may lead to errors in subsequent translations of the same term, resulting in a lack of consistency. Summary of the Invention

[0003] The main purpose of this application is to provide a real-time speech translation method, apparatus, device, and storage medium, which aims to solve the technical problem that existing mainstream machine translation systems cannot remember previously confirmed or corrected translation preferences.

[0004] To achieve the above objectives, this application proposes a real-time speech translation method, which includes: The received real-time speech is streamed for recognition to obtain the recognized text. The identified text is matched with the source terms in each memory unit, and the target terms corresponding to the identified text are determined based on the matching results. Obtain the translation result of the identified text; The translation term corresponding to the source term in the translation result is replaced with the target term to obtain the replaced translation text, and the translated text is output.

[0005] Optionally, the step of determining the target term corresponding to the identified text based on the matching result includes: In response to the matching result being a successful match, the context label corresponding to the successfully matched source term is obtained from each memory unit stored in the memory model. The source term corresponds to different candidate terms under different context labels. The domain label of the identified text is determined based on each of the aforementioned context labels; Based on the domain label, the target term corresponding to the identified text is determined from the candidate terms.

[0006] Optionally, the step of determining the domain label of the identified text based on each of the context labels includes: Obtain the confidence weight for each context label; Calculate the overall score for each context label based on the number of context labels and their corresponding confidence weights; The context label with the highest overall score is determined as the domain label of the identified text.

[0007] Optionally, the step of determining the target term corresponding to the identified text based on the matching result includes: In response to the matching result being a failure, the identified text is matched with the context labels in each of the memory units, and the pending terms to be replaced are selected from the identified text based on the matching result; Calculate the semantic similarity between the undetermined term and each of the source terms; If the semantic similarity is greater than a preset similarity threshold, the candidate term corresponding to the source term is taken as the target term of the identified text.

[0008] Optionally, the step of obtaining the translation result of the recognized text includes: Obtain the domain adaptation attributes of each candidate translation model; Based on the domain adaptation attribute, a translation model that matches the domain label of the identified text is determined from the candidate translation models; The identified text is translated according to the translation model to obtain the translation result.

[0009] Optionally, after the step of replacing the translation term corresponding to the source term in the translation result with the target term and outputting the replaced translation text, the method further includes: Receive correction instructions from users regarding the translated text, the correction instructions being used to instruct the correction of specific terms in the translated text; In response to the existence of a source term corresponding to the specific term in the memory unit, the memory unit corresponding to the source term is updated; In response to the absence of a source term corresponding to the specific term in the memory unit, a memory unit is set under the domain label of the recognized text according to the correction instruction; The translation results containing the specific term in the historical translation text are corrected.

[0010] Optionally, after the step of correcting the translation results containing the specific term in the historical translated text, the method further includes: Detect whether there is a conflict in the updated memory unit, where the conflict is that the same source term corresponds to different candidate terms when the context labels are consistent; If a conflict exists, the usage frequency and timestamp of each memory unit are obtained; The confidence priority of each memory cell is determined based on the frequency of use; The memory cells to be deleted are determined based on the confidence priority and / or the timestamp.

[0011] Furthermore, to achieve the above objectives, this application also proposes a real-time speech translation device, which includes: The text recognition module is used to perform streaming recognition on the received real-time speech to obtain the recognized text; The terminology determination module is used to match the identified text with the source terms in each memory unit, and determine the target terms corresponding to the identified text based on the matching results; A text translation module is used to obtain the translation result of the identified text; The translation output module is used to replace the translation terms corresponding to the source terms in the translation results with the target terms, and output the replaced translation text.

[0012] In addition, to achieve the above objectives, this application also proposes a real-time speech translation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the real-time speech translation method as described above.

[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the real-time speech translation method described above.

[0014] This application discloses a method for performing streaming recognition on received real-time speech to obtain recognized text; matching the recognized text with source terms in each memory unit and determining the target terms corresponding to the recognized text based on the matching results; obtaining the translation result of the recognized text; replacing the translation terms corresponding to the source terms in the translation result with the target terms to obtain the replaced translated text, and outputting the translated text. By performing streaming recognition on real-time speech and matching the recognized text with source terms in memory units in real time, and then replacing the corresponding parts in the general translation result with user-defined target terms, the method ensures both real-time translation and user translation preferences, significantly improving the consistency and personalized adaptability of terminology translation. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the first embodiment of the real-time speech translation method of this application; Figure 2 This is a flowchart illustrating the second embodiment of the real-time speech translation method of this application; Figure 3 This is a flowchart illustrating the workflow of the real-time speech translation system described in this application. Figure 4 This is a flowchart illustrating the third embodiment of the real-time speech translation method of this application; Figure 5 This is a schematic diagram of the module structure of the real-time speech translation device according to an embodiment of this application; Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the real-time speech translation method in the embodiments of this application.

[0018] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0020] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0021] Existing mainstream machine translation systems use pre-trained terminology databases and translation models, which have long update cycles and cannot adapt to temporary or personalized terminology generated by users in specific dialogues, meetings, or projects. Furthermore, when processing long dialogues or streaming text, the systems cannot remember previously confirmed or corrected translation preferences, leading to the possibility of the same term being misused in subsequent translations, resulting in a lack of consistency. Moreover, once users discover errors, they can only passively accept them, unable to correct them in real time and apply the corrections to subsequent real-time translations.

[0022] Therefore, this application provides a real-time speech translation method that can construct a user memory model that can be edited by the user in real time (in speech or text mode), so that the translation process changes from one-way output to two-way interaction, thereby achieving higher accuracy, consistency and personalization in the process of streaming recognition and translation.

[0023] It should be noted that the executing entity in this embodiment can be a computing service device with speech conversion, text translation, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions. The following description uses a real-time speech translation system as an example to illustrate this embodiment and the subsequent embodiments.

[0024] Based on this, the embodiments of this application provide a real-time speech translation method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the real-time speech translation method of this application.

[0025] In this embodiment, the real-time speech translation method includes: Step S10: Perform streaming recognition on the received real-time speech to obtain the recognized text.

[0026] It should be noted that real-time speech refers to continuous and sequential speech signals input by users in real time through audio acquisition devices (such as microphones, voice input modules, etc.), and its content can cover various forms of spoken expression such as dialogues, speeches, meeting discussions, and operation instructions. Recognized text refers to the textual representation of the input speech content output after processing by an automatic speech recognition system.

[0027] Understandably, in real-time speech translation scenarios, to achieve a "speak and translate simultaneously" interactive experience, streaming recognition technology must be used to process the input speech in real time. Compared with traditional sentence or paragraph recognition, streaming recognition can gradually generate text as the user continues to speak, providing continuous and timely text input for subsequent terminology matching, context understanding, and real-time translation, thereby ensuring the system's response speed and consistency in streaming scenarios such as dialogues and meetings.

[0028] In practical implementation, an end-to-end speech recognition model based on deep learning can be used, combined with a streaming decoding mechanism, to achieve low-latency, high-accuracy real-time speech-to-text conversion. Furthermore, techniques such as speaker adaptation, noise suppression, and speech enhancement can be introduced to improve the robustness of recognition under different accents, environmental noise, and speaking styles.

[0029] Optionally, in multi-person dialogue or conference scenarios, the system can integrate a speaker separation and identity recognition module to distinguish the voices of different speakers and perform streaming recognition separately, thereby supporting real-time text flow and translation in multi-speaker scenarios.

[0030] Step S20: Match the identified text with the source terms in each memory unit, and determine the target term corresponding to the identified text based on the matching result.

[0031] It should be noted that a memory unit is the basic data unit stored in the dynamic user memory model. Each unit contains at least one source term and its corresponding target term, and can be associated with metadata such as context labels and confidence levels. The dynamic user memory model is a lightweight, high-speed database (such as an in-memory database like Redis or a vector database) used to store structured user memory data. Source terms refer to user-defined original words, phrases, or specific expressions that require special translation. Target terms refer to the translation results specified by the user for the corresponding source terms.

[0032] Understandably, in order to overcome the shortcomings of general translation models in terms of personalized and specialized terminology during real-time translation, it is necessary to associate the identified text with the user's memory model through a real-time matching mechanism to achieve personalized translation results. This ensures that user-defined terms (such as project codes, professional abbreviations, and specific names) are accurately and consistently replaced or prioritized in the translation, thereby improving the accuracy and applicability of the translation results in specific scenarios.

[0033] In practical implementation, efficient string matching algorithms (such as prefix trees) can be used to quickly scan the identified text to achieve accurate matching of source terms. If an accurate match fails, the semantic features of the unmatched segments in the identified text can be further extracted, and fuzzy semantic matching can be performed with the source terms in the memory unit through vector similarity calculation to improve recall. After a successful match, the system will obtain the corresponding target term and associated contextual information for subsequent translation replacement and optimization.

[0034] It should be understood that contextual tags can be used for filtering during the matching process, matching only memory units that fit the current dialogue domain or scenario to improve matching accuracy and efficiency. If the same source term corresponds to multiple target terms in the matching (such as under different contextual tags), the final target term can be determined based on tag matching degree, confidence level, or the principle of recent use.

[0035] Understandably, if a match fails, the system can mark the relevant fragments as pending terms and decide whether to add them to the memory model based on subsequent semantic analysis or user interaction, thereby enabling the system to continuously learn and expand.

[0036] Step S30: Obtain the translation result of the recognized text.

[0037] It should be noted that the translation result refers to the text output obtained after converting the identified text from the source language to the target language, which can be the initial translation of the identified text generated by directly calling a general translation model.

[0038] Furthermore, to ensure that the most suitable translation engine or strategy is automatically selected for different professional fields or language styles, thereby further improving the professional accuracy and naturalness of the output results on the basis of general translation, step S30 may include: Obtain the domain adaptation attributes of each candidate translation model; determine the translation model that matches the domain label of the identified text from the candidate translation models based on the domain adaptation attributes; translate the identified text according to the translation model to obtain the translation result.

[0039] It should be noted that candidate translation models refer to one or more translation engines or strategies that the system can invoke, including general-purpose large language models, domain-specific translation models (such as those for medicine, law, finance, and technology), and personalized models fine-tuned using historical user data. Domain adaptation attributes refer to feature information describing the text domains or language styles that each translation model excels in, typically in the form of domain labels (such as medical, contract, software development), confidence scores, or performance metrics. Domain labels are used to characterize the professional or topic category to which the identified text belongs.

[0040] Understandably, a single, general-purpose translation model cannot maintain optimal professionalism and stylistic adaptability in all scenarios. By configuring multiple candidate translation models for the system and clarifying their domain-adaptive attributes, the most suitable model can be intelligently selected for translation based on the specific context (i.e., domain label) of the currently recognized text.

[0041] In practice, the system can maintain a model registry or configuration list, which records the identifier, server endpoint, and associated domain adaptation attributes of each candidate translation model. After obtaining the identified text and its domain label, the system matches the domain label with the domain adaptation attributes of each candidate model. Upon successful matching, the system calls the corresponding translation model interface, passing in the identified text and necessary contextual information (such as matched memory units), and retrieves the returned translation result.

[0042] It should be understood that if no perfectly matching model is found, a general model can be selected or the domain-specific model that is closest in semantic similarity can be selected.

[0043] Step S40: Replace the translation term corresponding to the source term in the translation result with the target term to obtain the replaced translation text, and output the translation text.

[0044] It should be noted that translation terms refer to the translated fragments in the initial translation results generated by a general or domain-specific translation model that correspond to the source terms in the memory unit. The translated text is the output text after modifying the translation terms in the translation results to the target terms. Through precise replacement, it is possible to ensure that user-defined terms (such as brand names, project codes, professional concepts, etc.) are correctly reflected in the final translation, thereby overcoming the limitations of fixed translation methods in general translation models in specific contexts.

[0045] In practice, the segment to be replaced can be located in the translation result based on the position of the source terms in the identified text. During replacement, the grammar and sentence structure of the translated text must remain consistent. If necessary, word forms or local adjustments can be made to the target terms to integrate them into the context. In streaming output scenarios, the replacement operation must be synchronized with the translation generation to achieve seamless output sentence by sentence or segment by segment, avoiding overall delay.

[0046] It should be understood that for speech output, the final text can be fed into the speech synthesis module, ensuring the naturalness of the pronunciation after term replacement. If mapping anomalies are found during the replacement process (such as the source term not finding a corresponding segment in the translation result or being ambiguous), the system can retain the original translated terms and provide the user with highlighted prompts or logs for subsequent optimization. During output, user-defined terms can also be selectively visually highlighted (e.g., bolded, colored) to enhance the transparency and interpretability of the results.

[0047] In this embodiment, by performing streaming recognition on real-time speech and matching the recognized text with the source terms in the memory unit in real time, the corresponding part in the general translation result is replaced with user-defined target terms. This ensures the real-time nature of the translation while also guaranteeing the user's translation preferences, significantly improving the consistency and personalization of terminology translation.

[0048] Reference Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the real-time speech translation method of this application. Based on the first embodiment described above, a second embodiment of the real-time speech translation method of this application is proposed. In the second embodiment, step S20 includes: Step S201: In response to the matching result being a successful match, obtain the context label corresponding to the successfully matched source term from each memory unit stored in the memory model. The source term corresponds to different candidate terms under different context labels.

[0049] It should be noted that context tags are domain, scenario, or project classification identifiers attached by users to data pairs of source terms and target terms when creating or editing memory units, such as medical - cardiovascular, Project A - code, etc. Candidate terms refer to a target translation bound to the same source term under a specific context tag. After the system recognizes a user - defined term, it further clarifies the context to which it belongs, because the same term may correspond to different translations in different contexts.

[0050] It can be understood that merely matching the source term is not sufficient to determine the uniquely correct translation. By obtaining and utilizing context tags, the system can bind the translation selection to a specific conversation scenario, professional field, or user - set context, solving the scenario of multiple translations for one word. For example, "Cell" is translated as "细胞" in biology and "电池" in electrical engineering.

[0051] It can be understood that after a successful match, the system will use the matched source term as a query key to retrieve all memory units containing this source term from the memory model (such as an in - memory database or a vector database). Subsequently, extract the context tags stored in these memory units and their associated candidate terms. At the same time, the system may retrieve multiple memory units (i.e., the same source term corresponds to multiple context tags and candidate terms). At this time, it is necessary to combine the context of the current recognized text (such as the domain tag of the recognized text) or the implicit tag of the current user session to filter out the most relevant one or more context tags.

[0052] It should be understood that the memory model (i.e., the dynamic user memory model) is a lightweight and high - speed - access database (such as the in - memory database Redis or a vector database) used to store structured user memory data. Its core data unit is the memory unit, and each unit contains: 1. Source term: the original vocabulary, phrase, or sentence provided by the user; 2. Target term: the corresponding translation specified by the user; 3. Context tag: the domain tag added by the user to this memory pair; 4. Metadata: creation / modification timestamp, confidence / usage frequency, creator, etc.

[0053] Step S202, determine the domain tag of the recognized text according to each of the context tags.

[0054] Furthermore, in order to more accurately identify the true context of the current conversation or text when a match is successful, step S202 may include: Obtain the confidence weight of each context tag; calculate the comprehensive score of each context tag according to the quantity and corresponding confidence weight of each context tag; determine the context tag with the highest comprehensive score as the domain tag of the recognized text.

[0055] It should be noted that the confidence weight is a quantitative value of credibility or importance assigned by the system to each context label, which can be dynamically calculated based on the usage frequency, creation time, user-defined priority, or historical translation adoption of the memory unit corresponding to the label. The comprehensive score is a weighted value obtained by combining the frequency of occurrence of the context label in the current recognized text with its confidence weight, and is used to objectively evaluate the degree of contextual matching between each label and the current text.

[0056] Understandably, when multiple contextual labels are matched, relying solely on the number of labels or a single metric is insufficient to accurately determine the true domain affiliation of the current text. By introducing confidence weights and calculating a comprehensive score, the most representative domain labels can be identified more accurately, ensuring that subsequent decisions such as terminology selection and model scheduling are based on reliable contextual judgment.

[0057] In practice, the system first extracts a set of context labels from all successfully matched memory units and counts the frequency of each label. Next, it obtains the confidence weight of each label from the memory model or system configuration. The overall score can be calculated using a weighted summation formula. After calculation, the system sorts all labels by their overall scores and determines the label with the highest score as the domain label for the currently recognized text.

[0058] Optionally, if the overall scores of multiple tags are very close or all below a preset threshold, the system can simultaneously use multiple high-scoring tags as domain references, or apply processing strategies for multiple related domains in parallel during translation. If the overall scores of all tags are too low, the system can determine that the current text belongs to a general domain and adopt the default translation strategy.

[0059] Step S203: Determine the target term corresponding to the identified text from the candidate terms based on the domain label.

[0060] Understandably, after obtaining multiple candidate terms, it is essential to conduct precise screening based on the actual domain of the current text to ensure that the selected target terms are consistent with the overall translation content in terms of semantics, style, and professionalism. For example, the same source term "table" may correspond to "table" in the database domain, but to "table" in the everyday domain.

[0061] In practice, the domain label is compared with the context label bound to each candidate term. If a completely matching context label exists, the corresponding candidate term is directly selected as the target term. If multiple completely matching candidate terms exist (e.g., different users provide different translations for the same context), they can be sorted according to the metadata of the associated memory unit (e.g., confidence level, usage frequency, last update time, or creator permissions), and the one with the highest priority is selected.

[0062] It should be understood that if no perfectly matching context label is found, the system can calculate the semantic similarity between the domain label and each context label, and select candidate terms corresponding to context labels with similarity exceeding a preset threshold. If it still cannot be determined, it can combine terminology preferences used in the current session history, user settings, or system default strategies for selection, or simultaneously use multiple candidate terms to generate translation variants for the user to choose from later.

[0063] Understandably, if none of the candidate terms meet the requirements, the system can temporarily use the translation of the term from a general translation model and record this matching decision for subsequent optimization of the memory model's update strategy. After determining the target term, the system can update the usage statistics of the relevant memory units in real time and visually annotate the custom terms during the output translation to enhance the interpretability of the results.

[0064] In the second embodiment, step S20 may include: Step S211: In response to the matching result being a failure, the identified text is matched with the context labels in each of the memory units, and the pending terms to be replaced are selected from the identified text according to the matching result.

[0065] It should be noted that undefined terms are words or phrases in the identified text that are not explicitly defined in the memory model, but may have special meanings in this field or require personalized translation.

[0066] Understandably, users cannot predefine all specialized or contextual terminology. Directly using generic translations when matching fails may result in poor domain adaptability. By enabling contextual tag matching, the system can indirectly perceive the potential domain of the current conversation. For example, by identifying repeated occurrences of words related to tags such as "cardiovascular" and "surgery" in the text, the system infers that the current domain is "medical-cardiac surgery." Based on this inference, the system can more specifically scan and identify text, filtering out generic words (such as "bypass") or newly emerging phrases that may have specific meanings within that domain, marking them as pending terms, and providing clear targets for subsequent semantic similarity retrieval or user interaction.

[0067] Step S212: Calculate the semantic similarity between the undetermined term and each of the source terms.

[0068] It should be noted that semantic similarity is a numerical indicator used to quantify the degree of semantic closeness between two words, phrases, or text fragments. It can be based on natural language processing technology, which is accomplished by converting the undetermined term and each source term into semantic feature vectors in a high-dimensional space (such as word vectors and sentence vectors), and then measuring the distance or angle between these vectors, such as cosine similarity.

[0069] Understandably, abandoning the processing of the term after an exact match fails would lead to a loss of translation personalization and accuracy. By calculating semantic similarity, the system can initiate a fuzzy matching mechanism to find the term that is semantically closest to the term to be determined from the user's existing terminology database (i.e., various source terms).

[0070] In its implementation, the system first uses a pre-trained semantic embedding model to convert the undetermined term and each source term into fixed-dimensional vector representations. Then, for each undetermined term vector, the similarity score between it and all source term vectors in the memory model is calculated sequentially.

[0071] Step S213: If the semantic similarity is greater than a preset similarity threshold, the candidate term corresponding to the source term is taken as the target term of the identified text.

[0072] It should be noted that the preset similarity threshold is a pre-set critical threshold used to determine whether semantic matching is successful. If the semantic similarity between the pending term and a source term exceeds this threshold, the system considers them to be semantically close enough and can be regarded as a successful fuzzy match. At this time, the system will use the candidate term associated with the source term (i.e., the target translation defined by the user for the source term) as the target term for the current pending term.

[0073] In one example, reference Figure 3 , Figure 3 This is a flowchart of the real-time speech translation system of this application. User A initiates a dialogue, and the corresponding speech signal enters the speech recognition stage, where it is converted into a Chinese text stream. This Chinese text stream then enters the real-time memory matching stage, which scans the recognized text stream in real time to quickly find whether it contains source terms from the memory model. Simultaneously, fuzzy matching is performed using semantic vector similarity (through a small semantic model) to capture variations in the user's expression. If a match is successful, the target terms from the memory model are directly used to replace the original output of the mainstream translation model. This is then passed to the AI ​​translation engine, which converts it into an English text stream. After TTS (Text-to-Speech) processing, the speech content is delivered to user B. If real-time memory matching fails, resulting in a translation error, user B will notice the error and correct it via voice command. This correction command triggers a memory model update operation, and the updated memory model is fed back to the real-time memory matching stage to update the target terms corresponding to the source terms in the memory model.

[0074] In this embodiment, when a match is successful, the domain label is further determined based on the context label, and the corresponding target term is selected, achieving accurate scenario adaptation for translation. When an exact match fails, fuzzy matching and term recommendation are achieved through semantic similarity calculation, significantly enhancing the system's fault tolerance and applicability. Even when faced with terms that are not explicitly defined by the user but are semantically similar, the system can still provide reasonable translation suggestions.

[0075] Reference Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the real-time speech translation method of this application. Based on the above embodiments, a third embodiment of the real-time speech translation method of this application is proposed. In the third embodiment, after step S40, the method further includes: Step S401: Receive correction instructions from the user regarding the translated text, the correction instructions being used to instruct the correction of specific terms in the translated text.

[0076] It should be noted that correction instructions are user feedback triggered by the user after viewing or listening to the output of the translated text, via voice, text, or an interactive interface. They aim to point out that a specific term in the translation is incorrect or inappropriate and provide the correct translation. The core information of a correction instruction includes at least the identifier of the "specific term" to be corrected in the original or translated text, and the "correct translation" provided by the user. Triggering methods can be dedicated voice commands (such as "Correction, X should be translated as Y"), graphical interface operations (such as clicking on the term and entering a new translation), or natural language descriptions.

[0077] It should be understood that the system receives user correction commands in real time through integrated multimodal interaction interfaces (such as microphone monitoring modules or graphical user interface event listeners). The voice interface integrates automatic speech recognition and speech synthesis modules, allowing users to interact with the system using natural language; the text interface provides a graphical interface, allowing users to easily view, edit, search, and manage the entire memory model. The interaction protocol defines a set of concise voice commands, such as "Remember…", "Correct…", "Ignore this time", etc., used to seamlessly trigger updates to the memory model during streaming translation.

[0078] For voice commands, the system first performs speech recognition, then uses natural language understanding technology to parse out the terms to be corrected and the target translation provided by the user. For text or interface operations, structured command data is directly obtained. After successful reception, the system enters the command processing flow, verifies the validity of the command, and extracts key parameters, such as the original terminology (or its position in the recognized text), the target term provided by the user, and optional contextual information.

[0079] Step S402: In response to the existence of a source term corresponding to the specific term in the memory unit, update the memory unit corresponding to the source term.

[0080] It should be understood that updating the memory unit corresponding to the source term refers to modifying the content of the memory unit already stored in the memory model that corresponds to the specific term indicated by the user. The update operation includes modifying the target term field to correct it to the new translation provided by the user in the correction instruction. Furthermore, the update may also involve adjusting the metadata of the memory unit, such as increasing its confidence level, updating the last modification timestamp, recording the number of corrections, or adding correction sources to reflect the latest status and learning history of the term.

[0081] Understandably, when a user requests a correction for a term that already exists in the memory model, it indicates that the translation currently stored in the system is inadequate or incorrect.

[0082] In practice, the system first uses the user-specified "specific term" (usually the source term) as the query key to search for the corresponding memory cell in the memory model. Upon successful query, the system replaces the original target term in the memory cell with the "correct translation" provided in the correction instruction. Simultaneously, the system updates relevant metadata, such as resetting the confidence level of the corresponding memory cell or weighting it according to the correction type, and updating the timestamp. If the user instruction contains new contextual information, the context label of the memory cell can also be updated synchronously.

[0083] Understandably, when a source term corresponding to a specific term exists in a memory unit, the user can directly say "delete this memory" to the translation result to delete the corresponding memory unit.

[0084] Step S403: In response to the absence of a source term corresponding to the specific term in the memory unit, a memory unit is set under the domain label of the recognized text according to the correction instruction.

[0085] It should be noted that setting up a memory unit refers to creating a completely new, structured data record within the dynamic user memory model. The core fields of the memory unit will include the "specific term" specified in the user's correction instructions as the source term, the correct translation provided by the user as the target term, and the domain label associated with the currently identified text as its initial context label. Furthermore, the memory unit will generate metadata such as creation time, creator, and initial confidence level, thus completing the full modeling and storage of new knowledge.

[0086] Understandably, when a user-corrected term does not exist in the memory model, it indicates that the user is introducing a completely new, personalized translation into the system.

[0087] In its implementation, the system first obtains the domain labels of the identified text from the context inference of the current session. Then, using specific terms and correct translations extracted from the user's correction instructions as the core, and combining them with the determined domain labels, a standardized memory unit data structure is constructed. Subsequently, the memory units are persistently stored in the memory model (such as an in-memory database or a vector database).

[0088] Step S404: Correct the translation results of the specific terms contained in the historical translation text.

[0089] It should be noted that historical translation text refers to all translation texts generated and output in the same session or related sessions before receiving this correction instruction.

[0090] In practice, the system initiates a backtracking process immediately after updating or creating the memory unit. First, based on the specific term (source term) and its associated contextual tags, the system retrieves all output translated text records from the current session's cache or persistent logs. Next, it locates all positions within these records containing the old translation of the term. Then, the system replaces each translation one by one with the new target term. The replacement process must consider the text's contextual structure to ensure the grammatical and semantic coherence of the replaced sentences.

[0091] Furthermore, to avoid translation conflicts arising from multiple different candidate terms for the same source term under the same context label, and to ensure the consistency and validity of the data in the memory model, step S404 is followed by: The updated memory units are checked for conflicts, where the conflict is that the same source term corresponds to different candidate terms when the context labels are consistent. If a conflict exists, the usage frequency and timestamp of each memory unit are obtained. The confidence priority of each memory unit is determined based on the usage frequency. The memory units to be deleted are determined based on the confidence priority and / or the timestamp.

[0092] It should be noted that in this embodiment, a conflict refers to two or more memory units under the same context label whose source terms (after semantic normalization) essentially point to the same concept or entity, but whose corresponding target terms (i.e., translations) are different. Usage frequency refers to the statistical count of the number of times each conflicting memory unit has been successfully matched and adopted in historical translations. Timestamp refers to the last modification time of the memory unit. Confidence priority is a ranking index calculated based on indicators such as usage frequency, used to rank the confidence of each memory unit when adjudicating conflicts.

[0093] Understandably, in scenarios supporting multi-user collaborative editing or allowing users to make multiple revisions, the memory model can easily generate multiple translation records for the same term in the same context, leading to conflicts. If these conflicts are not detected and resolved, the system's translation output in that context will be unstable or inconsistent, severely impacting user experience and the reliability of the translation results. Therefore, it is necessary to detect the modified memory units to eliminate conflicts.

[0094] In one feasible embodiment, the system can determine the confidence level priority based on the usage frequency of memory units, and identify the memory units with the lowest confidence level priority as the objects to be deleted. In this way, memory units that are used more frequently by users and that better match their terminology usage habits can be retained first, avoiding the accidental deletion of frequently used and effective memory units.

[0095] In another feasible embodiment, the system can use the creation or update timestamp of the memory unit as the basis for judgment, identifying the memory unit with the earliest timestamp as the object to be deleted. Newer memory units can be preferentially retained, allowing the memory model to align with the user's recent terminology usage habits and avoiding the use of outdated terminology matching rules.

[0096] In another feasible embodiment, the system can first perform preliminary screening based on confidence priority, listing memory units with lower confidence priority as candidates for deletion; if multiple memory units have the same confidence priority, then the timestamps are further compared, and the memory unit with the earlier timestamp is identified as the object to be deleted. This approach can both retain memory units frequently used by the user and prioritize recently updated content among memory units of the same priority, taking into account both the practicality and timeliness of memory units, so that the terminology matching of the memory model not only conforms to the user's common habits, but also meets their recent terminology adjustment needs.

[0097] In this embodiment, by introducing user correction instructions and supporting real-time updates and historical backtracking corrections of the memory unit, users are transformed from passive recipients of results into collaborative trainers of the system. They can provide immediate feedback to the system when errors are discovered and apply the correction results to subsequent real-time translations. At the same time, the consistency between historical translations and current translations is ensured, enabling the translation system to have the ability to continuously optimize and meet the personalized needs of users.

[0098] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the real-time speech translation method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0099] This application also provides a real-time voice translation device; please refer to... Figure 5 The real-time voice translation device includes: The text recognition module 10 is used to perform streaming recognition on the received real-time speech to obtain the recognized text; The terminology determination module 20 is used to match the identified text with the source terms in each memory unit, and determine the target term corresponding to the identified text based on the matching result; The text translation module 30 is used to obtain the translation result of the recognized text; The translation output module 40 is used to replace the translation term corresponding to the source term in the translation result with the target term, and output the replaced translation text.

[0100] The real-time speech translation device provided in this application, employing the real-time speech translation method in the above embodiments, can solve the technical problem that existing mainstream machine translation systems cannot remember previously confirmed or corrected translation preferences. Compared with the prior art, the beneficial effects of the real-time speech translation device provided in this application are the same as those of the real-time speech translation method provided in the above embodiments, and other technical features in the real-time speech translation device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0101] This application provides a real-time speech translation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the real-time speech translation method in Embodiment 1 above.

[0102] The following is for reference. Figure 6 The diagram illustrates a structural schematic suitable for implementing the real-time speech translation device of the embodiments of this application. The real-time speech translation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The real-time speech translation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0103] like Figure 6As shown, the real-time speech translation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the real-time speech translation device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the real-time speech translation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show real-time speech translation devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0104] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0105] The real-time speech translation device provided in this application, employing the real-time speech translation method described in the above embodiments, can solve the technical problem that existing mainstream machine translation systems cannot remember previously confirmed or corrected translation preferences. Compared with the prior art, the beneficial effects of the real-time speech translation device provided in this application are the same as those of the real-time speech translation method provided in the above embodiments, and other technical features of this real-time speech translation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0106] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0107] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0108] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the real-time speech translation method in the above embodiments.

[0109] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0110] The aforementioned computer-readable storage medium may be included in the real-time speech translation device; or it may exist independently and not be assembled into the real-time speech translation device.

[0111] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the real-time speech translation device, cause the real-time speech translation device to perform the real-time speech translation method described above.

[0112] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0114] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0115] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described real-time speech translation method. This addresses the technical problem that existing mainstream machine translation systems cannot remember previously confirmed or corrected translation preferences. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the real-time speech translation method provided in the above embodiments, and will not be elaborated upon here.

[0116] The above description is only a part of the embodiments of this application and does not limit the scope of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of this application.

Claims

1. A real-time speech translation method, characterized in that, The real-time speech translation method includes: The received real-time speech is streamed for recognition to obtain the recognized text. The identified text is matched with the source terms in each memory unit, and the target terms corresponding to the identified text are determined based on the matching results. Obtain the translation result of the identified text; The translation term corresponding to the source term in the translation result is replaced with the target term to obtain the replaced translation text, and the translated text is output.

2. The real-time speech translation method as described in claim 1, characterized in that, The step of determining the target term corresponding to the identified text based on the matching result includes: In response to the matching result being a successful match, the context label corresponding to the successfully matched source term is obtained from each memory unit stored in the memory model. The source term corresponds to different candidate terms under different context labels. The domain label of the identified text is determined based on each of the aforementioned context labels; Based on the domain label, the target term corresponding to the identified text is determined from the candidate terms.

3. The real-time speech translation method as described in claim 2, characterized in that, The step of determining the domain label of the identified text based on each of the context labels includes: Obtain the confidence weight for each context label; Calculate the overall score for each context label based on the number of context labels and their corresponding confidence weights; The context label with the highest overall score is determined as the domain label of the identified text.

4. The real-time speech translation method as described in claim 1, characterized in that, The step of determining the target term corresponding to the identified text based on the matching result includes: In response to the matching result being a failure, the identified text is matched with the context labels in each of the memory units, and the pending terms to be replaced are selected from the identified text based on the matching result; Calculate the semantic similarity between the undetermined term and each of the source terms; If the semantic similarity is greater than a preset similarity threshold, the candidate term corresponding to the source term is taken as the target term of the identified text.

5. The real-time speech translation method as described in claim 1, characterized in that, The step of obtaining the translation result of the recognized text includes: Obtain the domain adaptation attributes of each candidate translation model; Based on the domain adaptation attribute, a translation model that matches the domain label of the identified text is determined from the candidate translation models; The identified text is translated according to the translation model to obtain the translation result.

6. The real-time speech translation method as described in any one of claims 1 to 5, characterized in that, After the step of replacing the translation term corresponding to the source term in the translation result with the target term and outputting the replaced translation text, the method further includes: Receive correction instructions from users regarding the translated text, the correction instructions being used to instruct the correction of specific terms in the translated text; In response to the existence of a source term corresponding to the specific term in the memory unit, the memory unit corresponding to the source term is updated; In response to the absence of a source term corresponding to the specific term in the memory unit, a memory unit is set under the domain label of the recognized text according to the correction instruction; The translation results containing the specific term in the historical translation text are corrected.

7. The real-time speech translation method as described in claim 6, characterized in that, Following the step of correcting the translation results containing the specific term in the historical translated text, the method further includes: Detect whether there is a conflict in the updated memory unit, where the conflict is that the same source term corresponds to different candidate terms when the context labels are consistent; If a conflict exists, the usage frequency and timestamp of each memory unit are obtained; The confidence priority of each memory cell is determined based on the frequency of use; The memory cells to be deleted are determined based on the confidence priority and / or the timestamp.

8. A real-time voice translation device, characterized in that, The device includes: The text recognition module is used to perform streaming recognition on the received real-time speech to obtain the recognized text; The terminology determination module is used to match the identified text with the source terms in each memory unit, and determine the target terms corresponding to the identified text based on the matching results; A text translation module is used to obtain the translation result of the identified text; The translation output module is used to replace the translation terms corresponding to the source terms in the translation results with the target terms, and output the replaced translation text.

9. A real-time voice translation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the real-time speech translation method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the real-time speech translation method as described in any one of claims 1 to 7.