Streaming speech translation method and device, equipment and medium
By introducing thinking prompts and fine-grained text component processing into streaming speech translation technology, the context understanding problem in streaming speech translation is solved, and more accurate and smooth translation results are achieved, meeting the demand for efficient processing of streaming translation.
Patent Information
- Application Number
- CN202411855127.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-09
AI Technical Summary
Streaming speech translation technology faces contextual understanding problems caused by the difference in input characteristics from traditional machine translation, which may lead to ambiguity, misunderstanding or omission of important information in the translation results.
By introducing thinking prompts and fine-grained text component processing methods, the voice text to be translated is obtained and divided into multiple text components, and combined with the thinking prompt template for translation, ensuring that the model starts translation when receiving part of the information, and continuously corrects and improves the translation results as the input increases.
It significantly improves the accuracy and fluency of streaming speech translation, ensures real-time and coherence of translation results, avoids ambiguity, misunderstanding or omission of important information, and broadens the application scope of streaming speech translation technology.
Smart Images

Figure CN119962545A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech translation, and in particular to a streaming speech translation method, device, equipment and medium. Background Art
[0002] As a key technology for connecting speakers of different languages around the world, streaming speech translation has shown great application potential in many fields in recent years, such as global communication, international conferences, and tourism services. However, despite the significant progress made in this technology, it still faces a series of complex and urgent technical challenges in practical applications.
[0003] On the one hand, the input characteristics of streaming speech translation are fundamentally different from those of traditional machine translation. In traditional machine translation, the model usually receives complete and structured sentences as input, which makes the translation process relatively simple and accurate. However, in streaming speech translation, the input is often a real-time, continuous and incomplete speech text stream. This input mode requires the translation model to be highly flexible and adaptable, and to be able to start translating when receiving partial information, and to continuously correct and improve the translation results as the input increases.
[0004] On the other hand, the problem of context understanding in streaming speech translation is also particularly prominent. Since the input is fragmented, the model often lacks sufficient context information to accurately understand the semantics during the translation process. This may cause the model to produce ambiguity, misunderstanding, or omission of important information during translation, thus affecting the accuracy and fluency of the translation.
[0005] Therefore, in summary, the field of streaming speech translation technology urgently needs a translation method that can effectively solve the context understanding problem caused by differences in input characteristics. Summary of the invention
[0006] In view of the above problems, embodiments of the present invention are proposed to provide a streaming speech translation method, apparatus, device and medium that overcome the above problems or at least partially solve the above problems.
[0007] In order to solve the above problems, an embodiment of the present invention discloses a streaming speech translation method, which includes:
[0008] Obtaining a voice text to be translated and a thought prompt; the thought prompt is used to guide how to translate the voice text;
[0009] Determining at least one language text component to be translated according to the speech text to be translated;
[0010] Determining at least one target translation text component according to the at least one language text component to be translated and the thought prompt template;
[0011] A target translation text is determined according to the at least one target translation text component.
[0012] Optionally, determining at least one language text component to be translated according to the speech text to be translated includes:
[0013] Determine, according to the speech text to be translated, a minimum text unit corresponding to the language type of the speech text to be translated; the minimum text unit is used to represent the smallest understandable text unit in the language type of the speech text to be translated;
[0014] At least one language text component to be translated is determined according to the speech text to be translated and the minimum text unit.
[0015] Optionally, the thought prompt includes a text combination unit prompt, and determining at least one target translation text component according to the at least one language text component to be translated and the thought prompt template includes:
[0016] Determine at least one language combination text to be translated according to the at least one language text component to be translated and the text combination unit prompt;
[0017] At least one target translation text component is determined according to the at least one language combination text to be translated.
[0018] Optionally, the thought prompt includes a context prompt, and the determining of at least one target translation text component according to the at least one language combination text to be translated includes:
[0019] According to the context prompt, a first text memory is acquired; the first text memory is used to represent the historical translation results in the current text translation process;
[0020] At least one target translation text component is determined according to the at least one language combination text to be translated and the first text memory.
[0021] Optionally, the method further comprises:
[0022] In the process of determining at least one target translation text component according to the at least one language combination text to be translated and the first text memory, each time a target translation text component is determined, the first text memory is updated according to the determined target translation text component.
[0023] Optionally, determining at least one target translation text component according to the at least one language combination text to be translated and the first text memory includes:
[0024] According to the context prompt, a second text memory is acquired; the second text memory has all historical translation results in the text translation process;
[0025] At least one target translation text component is determined according to the at least one language combination text to be translated, the first text memory and the second text memory.
[0026] Optionally, the method further comprises:
[0027] After determining the target translation text according to the at least one target translation text component, the second text memory is updated according to the current translation result.
[0028] On the other hand, an embodiment of the present invention further discloses a streaming speech translation device, the device comprising:
[0029] A basic data acquisition module is used to acquire the voice text to be translated and the thought prompt; the thought prompt is used to guide how to translate the voice text;
[0030] A data component determination module, used to determine at least one language text component to be translated according to the speech text to be translated;
[0031] A translation component determination module, used to determine at least one target translation text component according to the at least one language text component to be translated and the thought prompt template;
[0032] The translation integration module is used to determine a target translation text according to the at least one target translation text component.
[0033] Optionally, the data component determination module includes:
[0034] A minimum text unit acquisition submodule is used to determine the minimum text unit corresponding to the language type of the speech text to be translated according to the speech text to be translated; the minimum text unit is used to represent the smallest understandable text unit in the language type of the speech text to be translated;
[0035] The first text component acquisition submodule is used to determine at least one language text component to be translated according to the speech text to be translated and the minimum text unit.
[0036] Optionally, the thought prompt includes a text combination unit prompt, and the translation component determination module includes:
[0037] A combined text acquisition submodule, used to determine at least one language combined text to be translated according to the at least one language text component to be translated and the text combination unit prompt;
[0038] The first target component determination submodule is used to determine at least one target translation text component according to the at least one language combination text to be translated.
[0039] Optionally, the thought prompt includes a context prompt, and the first target component determination submodule includes:
[0040] A first memory acquisition unit is used to acquire a first text memory according to the context prompt; the first text memory is used to represent the historical translation results in the current text translation process;
[0041] The second translation unit is configured to determine at least one target translation text component according to the at least one language combination text to be translated and the first text memory.
[0042] Optionally, the device further comprises:
[0043] The first memory updating submodule is used for, in the process of determining at least one target translation text component according to the at least one language combination text to be translated and the first text memory, updating the first text memory according to each target translation text component determined.
[0044] Optionally, the second translation unit includes:
[0045] A second memory acquisition subunit is used to acquire a second text memory according to the context prompt; the second text memory contains historical translation results of all text translation processes;
[0046] The third translation unit is used to determine at least one target translation text component according to the at least one language combination text to be translated, the first text memory and the second text memory.
[0047] Optionally, the device further comprises:
[0048] The second memory updating submodule is used to update the second text memory according to the current translation result after determining the target translation text according to the at least one target translation text component.
[0049] Accordingly, an embodiment of the present invention discloses an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the various steps of the above-mentioned streaming speech translation method embodiment when executed by the processor.
[0050] Accordingly, an embodiment of the present invention discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each step of the above-mentioned streaming speech translation method embodiment is implemented.
[0051] The embodiment of the present invention includes the following advantages: the embodiment of the present invention significantly improves the accuracy and fluency of streaming speech translation by introducing thought prompts and combining fine-grained text component processing methods; thought prompts provide additional guidance information for the translation model, which can help the model understand the semantics more accurately in fragmented input. The speech text to be translated is divided into multiple text components for processing, and this fine-grained processing method makes the translation process more flexible. The model can start translation when receiving partial information, and continuously correct and improve the translation results as the input increases. This flexibility ensures the real-time and coherence of the translation results and meets the requirements of streaming translation for efficient processing. By combining thought prompts with the language text components to be translated, the model can be more accurately guided to translate. The translation strategy, contextual information or language rules in the thought prompts can help the model make more accurate decisions during the translation process, avoid ambiguity, misunderstanding or omission of important information, thereby significantly improving the accuracy of translation. The thought prompts and text component processing methods enable the model to learn more about context, grammar and vocabulary selection during the training process. This learning method enhances the generalization ability of the model, enabling it to translate and adapt more flexibly when faced with voice input from different fields, styles or language habits, thereby broadening the application scope of streaming speech translation technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a flow chart of steps of an embodiment of a streaming speech translation method of the present invention;
[0053] Figure 2 is a flow chart of steps of another embodiment of a streaming speech translation method of the present invention;
[0054] Figure 3 It is a schematic diagram of a thought prompt chain of an embodiment of a streaming speech translation method of the present invention;
[0055] Figure 4 It is a flow chart of an embodiment of a streaming speech translation method of the present invention;
[0056] Figure 5 It is a structural block diagram of an embodiment of a streaming speech translation device of the present invention. DETAILED DESCRIPTION
[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0058] Streaming Speech refers to a technology that can process and generate voice data in real time. In streaming speech processing, voice signals are continuously input into the processing system, and the system can output the processing results in real time, such as speech-to-text (speech recognition), speech synthesis, etc. This technology is widely used in scenarios such as intelligent voice assistants, voice chat robots, voice translation, and voice broadcasting, and can provide real-time and smooth voice interaction experience. The core of streaming speech technology lies in its real-time and continuity. Compared with traditional speech processing technology, streaming speech technology does not need to wait for the entire voice input to be completed before processing, but can start processing and output results during the voice input process. This gives streaming speech technology a significant advantage in scenarios that require real-time response.
[0059] A thinking prompt template is a tool used to guide or inspire thinking and assist decision-making. It is usually presented in a structured form and contains a series of questions, suggestions or guiding principles. It aims to help the model analyze and judge more comprehensively and systematically when facing complex problems.
[0060] One of the core concepts of the embodiments of the present invention is to enhance the accuracy and fluency of streaming speech translation by constructing and introducing thought prompt templates and combining them with fine-grained text component processing.
[0061] Reference Figure 1 , shows a flow chart of a method for streaming speech translation according to an embodiment of the present invention, which may specifically include the following steps:
[0062] Step 101, obtaining a voice text to be translated and a thought prompt; the thought prompt is used to guide how to translate the voice text;
[0063] The speech text to be translated is generally the speech content spoken by the user in real time and needs to be translated into another language, or a training text prepared in advance, etc. In streaming speech translation, this text is generated continuously and in real time, rather than a complete sentence given in advance.
[0064] In one example, the voice text to be translated may be obtained by collecting language recognition samples corresponding to real-time audio as the voice text to be translated.
[0065] For example, real-time audio is obtained, and the continuous speech is converted into streaming speech text using ASR technology (Automatic Speech Recognition). The converted streaming speech text is used as a language recognition sample. It should be noted that the technology for obtaining streaming speech can also be other technologies, and the present invention does not limit this.
[0066] A mind cue is additional information that guides the translation process. It can contain translation strategies, contextual information, a domain-specific glossary, or any other information that helps to translate accurately. Mind cue can be static, such as predefined rules or glossaries, or dynamic, such as generated based on the current context.
[0067] In one example, thought prompts can be used to guide a large language model to translate streaming speech. At this time, thought prompts may include five parts: system prompts, user queries, memory content, output format, and thought prompt chains.
[0068] System prompts refer to instructions used to set roles, provide context and specific tasks when interacting with a large language model, guiding the large model to generate responses that better meet task requirements in specific situations.
[0069] For example, the system prompt is: "Suppose you are a machine simultaneous interpretation expert, responsible for translating streaming input Chinese sentences into English in real time."
[0070] User queries are explicit translation instructions given to the big model.
[0071] The memory content includes long-term memory and short-term memory, which is realized through an external dynamically updated memory module.
[0072] Output format: To ensure the readability of the output content of the large model, the output format can be required to be a structured JSON format in the prompt template. For example, the JSON format consisting of three fields, "reading and writing judgment", "translation result", and "previous sentence translation correction", respectively saves the output results of the three reasoning units in the thinking prompt chain, namely "semantic unit recognition", "translation", and "self-correction".
[0073] Thinking prompt chain. The complex annotation process of streaming translation text is decomposed into three serial reasoning units: "semantic unit recognition", "translation", and "self-correction". In each reasoning unit, natural language prompts are used to describe sub-problems, guiding the model to think and decide on the inferences that need to be made at this stage based on the results of the previous step and the requirements of the current problem;
[0074] It should be emphasized that the output format of the thought prompt can be other types of structured formats, and the present invention is not limited to this.
[0075] Step 102, determining at least one language text component to be translated according to the speech text to be translated;
[0076] Analyze the speech text to be translated and identify segments or phrases that can be translated independently. These segments or phrases are the language text components to be translated. The segmentation method can be based on factors such as natural pauses in speech, grammatical structure, word boundaries, or based on the construction of expert dictionaries.
[0077] Step 103, determining at least one target translation text component according to the at least one language text component to be translated and the thought prompt template;
[0078] For each language text component to be translated, the translation is performed in combination with the information in the thought prompts, such as translation strategy, context, etc. A target language text component corresponding to the source language text component is generated. In this process, a machine translation model, a rule-based translation system or other translation technology may be applied. It should be noted that the translation technology applied in the translation process may also be other translation technologies, which is not limited in the embodiments of the present invention.
[0079] Step 104 is used to determine a target translation text according to the at least one target translation text component.
[0080] All target language text components generated in step 103 are combined according to their order or logical structure in the original speech text. The combined text is adjusted and optimized as necessary to ensure the accuracy and fluency of the translation. This may include correcting grammatical errors, adjusting vocabulary selection, adding necessary conjunctions, etc. It should be noted that adjustment and optimization include other methods, which are not limited in the embodiment of the present invention.
[0081] The embodiment of the present invention significantly improves the accuracy and fluency of streaming speech translation by introducing thought prompts and combining fine-grained text component processing methods; thought prompts provide additional guidance information for the translation model, which can help the model understand semantics more accurately in fragmented input. The speech text to be translated is divided into multiple text components for processing, and this fine-grained processing method makes the translation process more flexible. The model can start translating when it receives partial information, and continuously correct and improve the translation results as the input increases. This flexibility ensures the real-time and coherence of the translation results and meets the requirements of streaming translation for efficient processing. By combining thought prompts with the language text components to be translated, the model can be more accurately guided to translate. The translation strategies, contextual information or language rules in the thought prompts can help the model make more accurate decisions during the translation process, avoid ambiguity, misunderstanding or omission of important information, thereby significantly improving the accuracy of translation. The thought prompts and text component processing methods enable the model to learn more about context, grammar and vocabulary selection during the training process. This learning method enhances the generalization ability of the model, enabling it to translate and adapt more flexibly when faced with voice input from different fields, styles or language habits, thereby broadening the application scope of streaming speech translation technology.
[0082] Reference Figure 2, shows a flow chart of another embodiment of a streaming speech translation method of the present invention, which may specifically include the following steps:
[0083] Step 201, obtaining a voice text to be translated and a thought prompt; the thought prompt is used to guide how to translate the voice text;
[0084] The speech text to be translated is generally the speech content spoken by the user in real time and needs to be translated into another language, or a training text prepared in advance, etc. In streaming speech translation, this text is generated continuously and in real time, rather than a complete sentence given in advance.
[0085] In one example, the voice text to be translated may be obtained by collecting language recognition samples corresponding to real-time audio as the voice text to be translated.
[0086] A mind cue is additional information that guides the translation process. It can contain translation strategies, contextual information, a domain-specific glossary, or any other information that helps to translate accurately. Mind cue can be static, such as predefined rules or glossaries, or dynamic, such as generated based on the current context.
[0087] In one example, thought prompts can be used to guide a large language model to translate streaming speech. At this time, thought prompts may include five parts: system prompts, user queries, memory content, output format, and thought prompt chains.
[0088] System prompts refer to instructions used to set roles, provide context and specific tasks when interacting with a large language model, guiding the large model to generate responses that better meet task requirements in specific situations.
[0089] For example, the system prompt is: "Suppose you are a machine simultaneous interpretation expert, responsible for translating streaming input Chinese sentences into English in real time."
[0090] User queries are explicit translation instructions given to the big model.
[0091] The memory content includes long-term memory and short-term memory, which is realized through an external dynamically updated memory module.
[0092] Output format: To ensure the readability of the output content of the large model, the output format can be required to be a structured JSON format in the prompt template. For example, the JSON format consisting of three fields, "reading and writing judgment", "translation result", and "previous sentence translation correction", respectively saves the output results of the three reasoning units in the thinking prompt chain, namely "semantic unit recognition", "translation", and "self-correction".
[0093] Thinking prompt chain. The complex annotation process of streaming translation text is decomposed into three serial reasoning units: "semantic unit recognition", "translation", and "self-correction". In each reasoning unit, natural language prompts are used to describe sub-problems, guiding the model to think and decide on the inferences that need to be made at this stage based on the results of the previous step and the requirements of the current problem;
[0094] It should be emphasized that the output format of the thought prompt can be other types of structured formats, and the present invention is not limited to this.
[0095] Step 202: determining the minimum text unit corresponding to the language type of the speech text to be translated according to the speech text to be translated; the minimum text unit is used to represent the minimum understandable text unit in the language type of the speech text to be translated;
[0096] Depending on the language of the speech text to be translated, the system needs to determine the smallest text unit in the language. These units are usually words, phrases or characters. They are the basic elements of the language and the smallest understandable unit in the translation process.
[0097] In one example, this process can be represented as: simulating a streaming input and output mode in a word-by-word increment manner using the smallest language unit, such as characters in Chinese and words in English. A streaming text is constructed for each complete sentence of the source language seed text.
[0098] Step 203: Determine at least one language text component to be translated according to the speech text to be translated and the minimum text unit.
[0099] Based on the speech text to be translated and the smallest text unit, the system splits the text into multiple language text components to be translated. These components may be sentences, phrases or words, which will be translated separately and combined later.
[0100] In one example, this process can be expressed as: the source language corpus is recorded as a dataset D = {X}, where x i =[w1w2…w j ,…,w n ] represents the source language text of the i-th data, w j Represents x i The jth smallest language unit in .
[0101] Step 204, determining at least one target translation text component according to the at least one language text component to be translated and the thought prompt template;
[0102] For each language text component to be translated, the translation is performed in combination with the information in the thought prompts, such as translation strategy, context, etc. A target language text component corresponding to the source language text component is generated. In this process, a machine translation model, a rule-based translation system or other translation technology may be applied. It should be noted that the translation technology applied in the translation process may also be other translation technologies, which is not limited in the embodiments of the present invention.
[0103] In one embodiment, the thought prompt includes a text combination unit prompt, and step 204 includes the following sub-steps:
[0104] Sub-step S11, determining at least one language combination text to be translated according to the at least one language text component to be translated and the text combination unit prompt;
[0105] According to the language text components to be translated and the text combination unit prompts (such as sentence structure, phrase collocation, etc.), the system generates the language combination text to be translated, that is, recombines the text components to be translated according to the habits of the target language.
[0106] Sub-step S12: Based on these combined texts, the system generates a target translation text component.
[0107] In one example, the text combination unit prompt includes a combination of professional terms in a corresponding field in a corresponding language context;
[0108] Sub-step S13, determining at least one target translation text component according to the at least one language combination text to be translated.
[0109] In one embodiment, the thought prompt includes a context prompt, and step 204 includes the following sub-steps:
[0110] Sub-step S21, obtaining a first text memory according to the context prompt; the first text memory is used to represent the historical translation results in the current text translation process;
[0111] According to the context prompt, the system obtains the first text memory, which includes the historical translation results in the current translation process to maintain the coherence of the translation.
[0112] Sub-step S22, determining at least one target translation text component according to the at least one language combination text to be translated and the first text memory.
[0113] The system generates a target translation text component by combining the language combination text to be translated and the first text memory. In this process, each time a target translation text component is determined, the first text memory is updated to reflect the latest translation status.
[0114] In one embodiment, the method further comprises the following steps:
[0115] In the process of determining at least one target translation text component according to the at least one language combination text to be translated and the first text memory, each time a target translation text component is determined, the first text memory is updated according to the determined target translation text component.
[0116] By accurately identifying and combining specialized terms, the system ensures professionalism and accuracy in translation. This is particularly important in fields such as medicine, law, and technology, where accurate translation of terms is essential to conveying the correct information. Text combination unit prompts help the system reassemble text components according to the sentence structure and phrase collocation rules of the target language, thereby generating more coherent and natural translation results.
[0117] In one embodiment, sub-step S22 includes the following sub-steps:
[0118] Sub-step S221, acquiring a second text memory according to the context prompt; the second text memory stores historical translation results of all text translation processes;
[0119] The system can also access a second text memory, which contains historical translation results of all text translation processes for a broader contextual understanding.
[0120] Sub-step S222: determining at least one target translation text component according to the at least one language combination text to be translated, the first text memory and the second text memory.
[0121] By combining the language combination text to be translated, the first text memory and the second text memory, the system generates a more accurate target translation text component.
[0122] In one example, the process of combining the first text memory and the second text and generating the target translation text component can be represented as:
[0123] Based on each source language text x i Constructing a streaming text can be expressed as: S i ={[w1],[w1w2],…,[w1w2…w j ],…,[w1w2…w j …w n ]}, a source language text of length n can construct n stream texts, where [w1w2…w j …w n ] is a complete sentence, and the rest of the streaming text is called "sentence prefix text".
[0124] During translation, the memory module is initialized and will store all previously translated streaming translation pairs <source language sentence prefix or sentence, target language sentence prefix or sentence> as long-term memory;
[0125] The memory module will save the streaming translation pair <source language sentence prefix, target language sentence prefix> of the preceding sentence prefix that is homologous to the current input as short-term memory;
[0126] Initially, both the long-term memory and the short-term memory of the memory module are empty. In particular, if the source language corpus contains proper nouns, the specific translation of the proper nouns can be stored in the long-term memory in advance to ensure the accurate translation of the proper nouns.
[0127] By accurately identifying and combining specialized terms, the system ensures professionalism and accuracy in translation. This is particularly important in fields such as medicine, law, and technology, where accurate translation of terms is essential to conveying the correct information. Text combination unit prompts help the system reassemble text components according to the sentence structure and phrase collocation rules of the target language, thereby generating more coherent and natural translation results.
[0128] In one embodiment, the method further comprises the steps of:
[0129] After determining the target translation text according to the at least one target translation text component, the second text memory is updated according to the current translation result.
[0130] Step 205: Determine a target translation text according to the at least one target translation text component.
[0131] The embodiment of the present invention significantly improves the accuracy and fluency of streaming speech translation by introducing thought prompts and combining fine-grained text component processing methods; thought prompts provide additional guidance information for the translation model, which can help the model understand semantics more accurately in fragmented input. The speech text to be translated is divided into multiple text components for processing, and this fine-grained processing method makes the translation process more flexible. The model can start translating when it receives partial information, and continuously correct and improve the translation results as the input increases. This flexibility ensures the real-time and coherence of the translation results and meets the requirements of streaming translation for efficient processing. By combining thought prompts with the language text components to be translated, the model can be more accurately guided to translate. The translation strategies, contextual information or language rules in the thought prompts can help the model make more accurate decisions during the translation process, avoid ambiguity, misunderstanding or omission of important information, thereby significantly improving the accuracy of translation. The thought prompts and text component processing methods enable the model to learn more about context, grammar and vocabulary selection during the training process. This learning method enhances the generalization ability of the model, enabling it to translate and adapt more flexibly when faced with voice input from different fields, styles or language habits, thereby broadening the application scope of streaming speech translation technology.
[0132] Reference Figure 3 , showing a schematic diagram of a thought prompt chain of an embodiment of a streaming speech translation method of the present invention;
[0133] In one embodiment, the thought prompt chain can be divided into three serially connected reasoning units, namely speech unit recognition, translation and self-correction; natural language prompts are used in each reasoning unit to describe sub-problems, guiding the model to think and decide on the inferences that need to be made at this stage based on the results of the previous step and the requirements of the current problem. Among them, the "semantic unit recognition" unit uses high-quality prompt words to guide the model to judge whether the current input text stream can constitute a combination of the smallest semantic units. The reasoning result of this unit is a clear instruction of "[translate]" or "[temporarily save]".
[0134] In the prompt words, when the text to be translated can be well understood and translated without waiting for more words to enter, the reasoning unit outputs the instruction of [translate] and enters the next reasoning unit; otherwise, the reasoning unit outputs the instruction of [temporarily store] and no longer enters the subsequent reasoning unit, and takes the translation result of the prefix of the previous sentence as the translation result of this round.
[0135] The "Translation" unit prompts the large model to refer to the translations of all sentence prefixes and similar sentences, complete the translation of the current sentence, and record the translation result in the "Translation Result" field in the output JSON.
[0136] The "self-correction" unit combines the translation results in the memory content and the current text flow to determine whether the translation result of the previous sentence prefix needs to be corrected. If no correction is required, the "Previous sentence translation correction" field in JSON outputs the [keep previous sentence translation] instruction; if correction is required, the corrected translation result is generated and written into the "Previous sentence translation correction" field in the output JSON.
[0137] Reference Figure 4 , showing a flow chart of a streaming speech translation method embodiment of the present invention:
[0138] First, we collect language recognition samples corresponding to the audio as the source language seed text. The ways to obtain the source language seed text include: downloading the parallel corpus of machine translation; collecting public audio resources such as daily conversations, keynote speeches, conference recordings, film and television videos, and transcribing them into source language text through automatic language recognition systems or manual transcription. The specific method can be dynamically set according to business needs.
[0139] After collecting and organizing the source language seed text, a word-by-word streaming text is constructed based on the source language text of the complete sentence.
[0140] An optional process is: for each source language seed text, construct streaming text data through string slicing in a word-by-word increment manner using the smallest language unit, such as characters for Chinese and words for English.
[0141] Taking Chinese as an example, if the complete source language text is "Today the weather is really nice", then the corresponding streaming text sequence includes: "today", "today", "today's weather", "today's weather", "today's weather is really nice", and "today's weather is really nice".
[0142] A thinking prompt template is constructed based on five parts: system prompts, user queries or specific tasks, memory content or translation references, output format, and thinking prompt chain.
[0143] For example, an alternative example of a thought prompt template is:
[0144]
[0145]
[0146]
[0147] Then, the streaming text is translated and synthesized using the large language model and the thought chain prompts in the thought prompt template.
[0148] At the beginning of the translation, the short-term memory and long-term memory of the memory module are empty. The memory module will be continuously updated according to the translation process;
[0149] Each stream of speech from the same source is sequentially embedded into a thought prompt template and input into the large language model to obtain the inference result;
[0150] For example, an optional translation process is: read short-term memory from the memory, and the short-term memory is a streaming translation pair of the prefix of the previous sentence that is homologous to the current input text stream. If the short-term memory is empty at this time, write "None" to the corresponding position in the template, remove step 3 from the "reasoning hint" part of the template, and fill the {target language sentence prefix i-1} position in step 2 with an empty string.
[0151] Retrieve long-term memory from the memory, the long-term memory is the streaming translation pairs of all sentences or sentence prefixes similar to the current input text stream. If the long-term memory is empty at this time, "None" is written in the corresponding position in the template. One optional implementation method of long-term memory is to use the Jaccard similarity algorithm. The implementation method can also be other algorithms besides the Jaccard similarity algorithm, which is not limited in the embodiment of the present invention;
[0152] The streaming text to be translated and the long and short-term memory are embedded in the constructed thinking chain prompt template to obtain the translation prompt words of this streaming text.
[0153] The translation prompt of each streaming text is used as a big model query and input into the big model for reasoning. After the reasoning of each streaming text is completed, the translated target language text is obtained from the structured reasoning results.
[0154] The <source language sentence prefix or sentence, target language sentence prefix or sentence> of this round of translation is added to the short-term memory, and the short-term memory of the memory module is updated.
[0155] If the inference result adjusts the target language sentence prefix of the previous source language sentence prefix, the short-term memory in the memory module is updated synchronously.
[0156] If the same content as the current translation does not exist in the long-term memory, the source language and target language sentence prefixes or sentences of the current translation are added to the long-term memory of the memory.
[0157] After all streaming text translations of the source language seed text are completed, the short-term memory in the memory is read as the streaming translation data of this source language text and appended to the constructed streaming translation dataset.
[0158] Read the next complete source language text, initialize the short-term memory in the memory module to be empty, and repeat the above translation steps until all complete source language texts are marked.
[0159] After the initial translation is completed, the source language text and the streaming translation text are integrated according to the pre-defined dataset format to obtain a preliminary synthesized streaming translation dataset.
[0160] For example, an optional dataset format is: including a "collection_method" field for storing the data source, a "source" field for storing complete source language sentences, a "translation" field for storing complete target language sentences, a "streaming_translation" field for storing each corresponding streaming text and its translation, and an "annotation" field for storing the large language model model used. The "manual_modify" field is used to indicate whether the sentence has been manually modified. 0 represents no manual modification, 1 represents manual modification, and other fields;
[0161] A specific example is as follows:
[0162]
[0163]
[0164] After obtaining the preliminary synthesized streaming translation dataset, the streaming translation text is scored for translation quality based on the large language referee model. The average quality score of all streaming translation texts under the source language text is used as the quality score of the streaming translation of this source language text.
[0165] For example, when the Big Language Judge model evaluates a translation dataset, it will comprehensively score the translation dataset based on the following dimensions:
[0166] The accuracy dimension, for example, when scoring, focuses on whether the translation dataset accurately conveys the meaning and information of the original text, including key concepts, details, and context in the original text.
[0167] For example, when scoring, the fluency dimension focuses on whether the language of the translation dataset is natural and fluent, and conforms to the expression habits of the target language. The translation dataset is scored based on whether there are grammatical errors, such as tense, voice, subject-verb agreement, etc.
[0168] Sentence coherence: Ensure that the sentences in the translation dataset are clearly structured, logically coherent, and easy for readers to understand.
[0169] For example, when scoring, attention will be paid to whether the translation dataset maintains the style and tone of the original text in order to faithfully convey the characteristics of the original text, and whether the translation dataset arbitrarily changes the viewpoint or position of the original text, and scoring will be based on this.
[0170] The cultural adaptability dimension, for example, takes into account the cultural background and language habits of the target language when scoring, and whether authentic expressions of the target language are used in the translation dataset to make the translation dataset easier for the target audience to accept and understand, and scores accordingly.
[0171] The terminology consistency dimension, for example, when scoring texts involving specific fields, whether the translation dataset ensures that the translation of professional terms is accurate and consistent, and whether synonyms or ambiguous words are used in the translation dataset, resulting in unclear or confusing meanings.
[0172] Layout and format dimensions, for example, when scoring, we will pay attention to whether the translation dataset maintains the specific layout of the original text
[0173] Speed and efficiency dimensions, for example, when scoring, we will focus on the speed and efficiency of translation
[0174] Subjective interpretation and objective evaluation dimensions, for example, in translation datasets, one should avoid adding personal subjective interpretations or opinions, and should faithfully convey the meaning of the original text.
[0175] For the above-mentioned multiple scoring dimensions, corresponding weights can be set according to the business scenario to obtain scoring rules that are more in line with the business scenario, thereby obtaining a more accurate translation data set.
[0176] Based on the quality score, low-scoring translation data are screened out and manually verified and corrected to obtain the final synthetic dataset.
[0177] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0178] Reference Figure 5 , shows a structural block diagram of an embodiment of a streaming speech translation device of the present invention, which may specifically include the following modules:
[0179] The basic data acquisition module 301 is used to acquire the voice text to be translated and the thought prompt; the thought prompt is used to guide how to translate the voice text;
[0180] A data component determination module 302, configured to determine at least one language text component to be translated according to the speech text to be translated;
[0181] A translation component determination module 303, configured to determine at least one target translation text component according to the at least one language text component to be translated and the thought prompt template;
[0182] The translation integration module 304 is configured to determine a target translation text according to the at least one target translation text component.
[0183] In one embodiment, the data component determination module includes:
[0184] A minimum text unit acquisition submodule is used to determine the minimum text unit corresponding to the language type of the speech text to be translated according to the speech text to be translated; the minimum text unit is used to represent the smallest understandable text unit in the language type of the speech text to be translated;
[0185] The first text component acquisition submodule is used to determine at least one language text component to be translated according to the speech text to be translated and the minimum text unit.
[0186] In one embodiment, the thought prompt includes a text combination unit prompt, and the translation component determination module includes:
[0187] A combined text acquisition submodule, used to determine at least one language combined text to be translated according to the at least one language text component to be translated and the text combination unit prompt;
[0188] The first target component determination submodule is used to determine at least one target translation text component according to the at least one language combination text to be translated.
[0189] In one embodiment, the thought prompt includes a context prompt, and the first target component determination submodule includes:
[0190] A first memory acquisition unit is used to acquire a first text memory according to the context prompt; the first text memory is used to represent the historical translation results in the current text translation process;
[0191] The second translation unit is configured to determine at least one target translation text component according to the at least one language combination text to be translated and the first text memory.
[0192] In one embodiment, the device further comprises:
[0193] The first memory updating submodule is used for, in the process of determining at least one target translation text component according to the at least one language combination text to be translated and the first text memory, updating the first text memory according to each target translation text component determined.
[0194] In one embodiment, the second translation unit includes:
[0195] A second memory acquisition subunit is used to acquire a second text memory according to the context prompt; the second text memory contains historical translation results of all text translation processes;
[0196] The third translation unit is used to determine at least one target translation text component according to the at least one language combination text to be translated, the first text memory and the second text memory.
[0197] In one embodiment, the device further comprises:
[0198] The second memory updating submodule is used to update the second text memory according to the current translation result after determining the target translation text according to the at least one target translation text component.
[0199] The embodiment of the present invention significantly improves the accuracy and fluency of streaming speech translation by introducing thought prompts and combining fine-grained text component processing methods; thought prompts provide additional guidance information for the translation model, which can help the model understand semantics more accurately in fragmented input. The speech text to be translated is divided into multiple text components for processing, and this fine-grained processing method makes the translation process more flexible. The model can start translating when it receives partial information, and continuously correct and improve the translation results as the input increases. This flexibility ensures the real-time and coherence of the translation results and meets the requirements of streaming translation for efficient processing. By combining thought prompts with the language text components to be translated, the model can be more accurately guided to translate. The translation strategies, contextual information or language rules in the thought prompts can help the model make more accurate decisions during the translation process, avoid ambiguity, misunderstanding or omission of important information, thereby significantly improving the accuracy of translation. The thought prompts and text component processing methods enable the model to learn more about context, grammar and vocabulary selection during the training process. This learning method enhances the generalization ability of the model, enabling it to translate and adapt more flexibly when faced with voice input from different fields, styles or language habits, thereby broadening the application scope of streaming speech translation technology.
[0200] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0201] An embodiment of the present invention further provides an electronic device, including:
[0202] The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, each process of the above-mentioned streaming speech translation method embodiment is implemented and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0203] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned streaming speech translation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0204] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0205] It will be appreciated by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0206] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0207] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0208] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0209] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0210] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0211] The above is a detailed introduction to a streaming speech translation method, device, equipment and medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A streaming speech translation method, characterized in that: The method comprises: Obtaining a voice text to be translated and a thought prompt; the thought prompt is used to guide how to translate the voice text; Determining at least one language text component to be translated according to the speech text to be translated; Determining at least one target translation text component according to the at least one language text component to be translated and the thought prompt template; A target translation text is determined according to the at least one target translation text component.
2. The method according to claim 1, characterized in that: The step of determining at least one language text component to be translated according to the speech text to be translated comprises: Determine, according to the speech text to be translated, a minimum text unit corresponding to the language type of the speech text to be translated; the minimum text unit is used to represent the smallest understandable text unit in the language type of the speech text to be translated; At least one language text component to be translated is determined according to the speech text to be translated and the minimum text unit.
3. The method according to claim 2, characterized in that The thought prompt includes a text combination unit prompt, and determining at least one target translation text component according to the at least one language text component to be translated and the thought prompt template includes: Determine at least one language combination text to be translated according to the at least one language text component to be translated and the text combination unit prompt; At least one target translation text component is determined according to the at least one language combination text to be translated.
4. The method according to claim 3, characterized in that The thought prompt includes a context prompt, and the determining at least one target translation text component according to the at least one language combination text to be translated includes: According to the context prompt, a first text memory is acquired; the first text memory is used to represent the historical translation results in the current text translation process; At least one target translation text component is determined according to the at least one language combination text to be translated and the first text memory.
5. The method according to claim 4, characterized in that The method further comprises: In the process of determining at least one target translation text component according to the at least one language combination text to be translated and the first text memory, each time a target translation text component is determined, the first text memory is updated according to the determined target translation text component.
6. The method according to claim 5, characterized in that The step of determining at least one target translation text component according to the at least one language combination text to be translated and the first text memory comprises: According to the context prompt, a second text memory is acquired; the second text memory has all historical translation results in the text translation process; At least one target translation text component is determined according to the at least one language combination text to be translated, the first text memory and the second text memory.
7. The method according to claim 6, characterized in that The method further comprises: After determining the target translation text according to the at least one target translation text component, the second text memory is updated according to the current translation result.
8. A streaming speech translation device, characterized in that: The device comprises: A basic data acquisition module is used to acquire the voice text to be translated and the thought prompt; the thought prompt is used to guide how to translate the voice text; A data component determination module, used to determine at least one language text component to be translated according to the speech text to be translated; A translation component determination module, used to determine at least one target translation text component according to the at least one language text component to be translated and the thought prompt template; The translation integration module is used to determine a target translation text according to the at least one target translation text component.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the streaming speech translation method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the streaming speech translation method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Machine translation method and system based on large model fusion task disassembly and cognitive planning
CN120745656A