A speech translation method, system and related device

By identifying audio languages, generating prompt texts and obtaining reference vocabulary, the problem of inaccurate audio translation in translation scenarios in the prior art is solved, and a more efficient and accurate voice translation effect is achieved.

CN119692368BActive Publication Date: 2025-06-27IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510205232.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

In the translation scenarios, the existing technology is due to high degree of diversity, rich term expression, and large cultural differences, which lead to the inability to translate the audio input by users accurately and effectively, and translation errors, text nesting and other situations often occur.

Method used

By obtaining the language to be translated for input audio, the audio is recognized and the initial recognition text is generated, the prompt text is obtained based on the language and text, and reference vocabulary is obtained from the candidate vocabulary, and the translated text is generated in combination with the initial recognition text, prompt text and reference vocabulary.

Benefits of technology

Improve the accuracy and efficiency of speech translation, and reduce the situation of translation errors and text nesting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119692368B_ABST
    Figure CN119692368B_ABST
Patent Text Reader

Abstract

The present application discloses a voice translation method, system and related device. The method includes: obtaining input audio at at least one end, determining the language to be translated of the input audio, using each end of the input audio as the audio to be translated respectively, and determining the initial recognition text corresponding to the audio to be translated; wherein, at least one recognition language is matched with the initial recognition text; based on the language to be translated, the initial recognition text and its corresponding recognition language, obtaining a prompt text matched with the audio to be translated; wherein, the prompt text includes a conversion language matched with the recognition language; obtaining reference words matched with the initial recognition text from a candidate word library; and based on the initial recognition text, the prompt text and the reference words, obtaining a translation text corresponding to the audio to be translated. By the above method, the present application can improve the accuracy and efficiency of voice translation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular to a speech translation method, system and related devices. Background Art

[0002] With the rapid development of artificial intelligence, its application in the field of translation is becoming more and more extensive. At present, there are various pre-trained translation models for audio translation in the prior art. However, in practical applications, due to factors such as high diversification of translation scenarios, rich term expressions, and large cultural differences, it is impossible to accurately and effectively translate the input audio by the user only, and translation errors, text nesting and other situations often occur.

[0003] Therefore, how to accurately and effectively translate the text to be translated is an urgent problem to be solved in the prior art. Summary of the Invention

[0004] The main technical problem to be solved by the present application is to provide a speech translation method, system and related devices, which can improve the accuracy and efficiency of speech translation.

[0005] To solve the above technical problem, a technical solution adopted by the present application is: to provide a speech translation method, including: obtaining input audio of at least one end, determining the language to be translated of the input audio, taking the input audio of each end as the audio to be translated respectively, and determining the initial recognition text corresponding to the audio to be translated; wherein, at least one recognition language is matched with the initial recognition text; based on the language to be translated, the initial recognition text and its corresponding recognition language, obtaining a prompt text matched with the audio to be translated; wherein, the prompt text includes a conversion language matched with the recognition language; obtaining reference words matched with the initial recognition text from a candidate word library; and obtaining a translation text corresponding to the audio to be translated based on the initial recognition text, the prompt text and the reference words.

[0006] To solve the above technical problems, another technical solution adopted in this application is: to provide a voice translation system, including: an identification module, configured to obtain input audio at at least one end, determine the language to be translated of the input audio, use each end of the input audio as the audio to be translated respectively, and determine the initial identification text corresponding to the audio to be translated; wherein, at least one identification language is matched with the initial identification text; a first processing module, configured to obtain a prompt text matching the audio to be translated based on the language to be translated, the initial identification text and its corresponding identification language; wherein, the prompt text includes a conversion language matching the identification language; a second processing module, configured to obtain reference words matching the initial identification text from a candidate word library; a translation module, configured to obtain a translation text corresponding to the audio to be translated based on the initial identification text, the prompt text and the reference words.

[0007] To solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, including: a memory and a processor coupled to each other, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the method mentioned in the above technical solution.

[0008] To solve the above technical problems, yet another technical solution adopted in this application is: to provide a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the method mentioned in the above technical solution is implemented.

[0009] The beneficial effect of this application is: different from the prior art, the voice translation method proposed in this application obtains input audio at at least one end from a scenario and determines the language to be translated of the input audio. The audio to be translated is identified to obtain the corresponding initial identification text, and a prompt text is determined according to the audio to be translated, the initial identification text and its matching identification language. Also, reference words matching the initial identification text are obtained from a candidate word library, and a translation text is obtained by combining the initial identification text, the prompt text and the reference words, improving the accuracy of voice translation. Description of the Drawings

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:

[0011] Figure 1 is a schematic flowchart of an embodiment of the voice translation of this application;

[0012] Figure 2 isFigure 1 The flow chart corresponding to step S102 in another embodiment;

[0013] Figure 3 is Figure 1 The flow chart corresponding to step S102 in yet another embodiment;

[0014] Figure 4 is Figure 1 The flow chart corresponding to step S103 in another embodiment;

[0015] Figure 5 is Figure 4 The flow chart corresponding to step S404 in another embodiment;

[0016] Figure 6 is Figure 1 The flow chart corresponding to step S104 in another embodiment;

[0017] Figure 7 It is a schematic structural diagram of an embodiment of the voice translation system of the present application;

[0018] Figure 8 It is a schematic structural diagram of an embodiment of the electronic device of the present application;

[0019] Figure 9 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Specific embodiments

[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments, and different embodiments can be adaptively combined. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0021] Please refer to Figure 1 , Figure 1 It is a schematic flow chart of an embodiment of the voice translation of the present application. The method includes:

[0022] S101: Obtain the input audio at at least one end, determine the language to be translated of the input audio, use the input audio at each end as the audio to be translated respectively, and determine the initial recognition text corresponding to the audio to be translated; wherein, at least one recognition language is matched with the initial recognition text.

[0023] In one embodiment, input audio collected at at least one end is obtained, and the language to be translated of the input audio is determined. The input audio at each end is respectively used as the audio to be translated, and a preliminary text recognition is performed on the audio to be translated to obtain the initial recognition text corresponding to the audio to be translated.

[0024] Specifically, after using the input audio at each end as the audio to be translated, the input audio is encoded to obtain the corresponding target audio encoding. The target audio encoding is decoded to obtain the initial recognition text corresponding to the audio to be translated, and at least one recognition language corresponding to the initial recognition text is determined.

[0025] In a specific application scenario, the above voice translation method is applied to a multi-person meeting scenario, and input audio of a corresponding target object in the meeting scenario is collected by using a plurality of microphone collection ends, and a plurality of languages to be translated corresponding to the input audio are determined. For example, the above target object is a participant in the meeting. When the languages expressed by a plurality of participants include Chinese, English, and Korean, then the plurality of languages to be translated are determined to be Chinese, English, and Korean.

[0026] S102: Based on the language to be translated, the initial recognition text, and its corresponding recognition language, obtain a prompt text that matches the audio to be translated; wherein, the prompt text includes a conversion language that matches the recognition language.

[0027] In one embodiment, a prompt text is generated according to the language to be translated, the initial recognition text, and its corresponding recognition language. The prompt text includes a conversion language that is different from the recognition language determined from the language to be translated, and the prompt text is used to prompt to translate the audio to be translated into a translation text that matches the corresponding conversion language.

[0028] In a specific application scenario, according to the language expressed by the target object corresponding to any end, the corresponding conversion language is determined. For example, when the language spoken by the target object is Chinese, the corresponding conversion language is determined to be Chinese, so that the subsequent translated text can be Chinese text that the target object can understand.

[0029] In another embodiment, a conversion language that is different from the recognition language is obtained from the language to be translated, and the conversion language can be specified by a relevant target object or a technician.

[0030] S103: Obtain a reference vocabulary that matches the initial recognition text from the candidate vocabulary.

[0031] In one embodiment, in response to the initial recognition text including multiple recognition characters and each candidate word in the candidate word library including multiple reference characters, reference words related to the initial recognition text are filtered out according to the matching degree between the recognition characters and the corresponding reference characters. By selecting reference words from the candidate word library to provide a reference for the subsequent translation process, the accuracy of text translation in different fields is improved.

[0032] Specifically, for the initial recognition text and each candidate word, the number of identical characters in the corresponding recognition characters and the corresponding reference characters is obtained, and the candidate word corresponding to the maximum number is used as the reference word.

[0033] S104: Based on the initial recognition text, the prompt text, and the reference words, obtain the translation text corresponding to the audio to be translated.

[0034] In one embodiment, the audio to be translated is translated in combination with the obtained initial recognition text, prompt text, and reference words to obtain the corresponding translation text.

[0035] Specifically, the obtained initial recognition text, prompt text, and reference words are input into the trained audio translation model, and the audio translation model is used to translate the audio to be translated to obtain the corresponding translation text. The specific structure of the above audio translation model can refer to the existing neural network model structure.

[0036] The speech translation method proposed in this application obtains at least one input audio from the scene and determines the language to be translated of the input audio. The audio to be translated is recognized to obtain the corresponding initial recognition text, and the prompt text is determined according to the audio to be translated, the initial recognition text, and the recognized language that matches it. Also, reference words matching the initial recognition text are obtained from the candidate word library, and the translation text is obtained by combining the initial recognition text, the prompt text, and the reference words, improving the accuracy of speech translation.

[0037] Please refer to Figure 2 , Figure 2 is Figure 1 The flowchart of another embodiment corresponding to step S102 in. Specifically, the implementation process of step S102 includes:

[0038] S201: Based on the recognized language, divide the initial recognition text into multiple text segments; wherein, each text segment matches the same recognized language, and adjacent text segments match different recognized languages.

[0039] In one embodiment, the initial recognition text includes multiple consecutive recognition characters. According to the recognized language, the recognition characters in the initial recognition text are divided to obtain the corresponding multiple text segments. Among them, the recognition characters in the same text segment match the same recognized language, and adjacent text segments match different recognized languages.

[0040] In a specific application scenario, the initial recognition text obtained is "Ladies and gentlemen, good morning everyone!", then it is determined that the initial recognition text matches the adjacent first text segment "Ladies and gentlemen, good morning everyone!" and the second text segment "Ladies and gentlemen, good morning everyone", and the recognition language matched by the first text segment is Chinese, and the recognition language matched by the second text segment is English.

[0041] S202: Based on all text segments and their matching recognition languages, determine the conversion language matched by each text segment.

[0042] In one embodiment, based on all text segments and their matched recognized languages, a conversion language different from the corresponding recognized language is determined from the languages ​​to be translated.

[0043] Specifically, the conversion language matching each text segment is determined according to the expression language of the target object at either end and the recognition language corresponding to each text segment. For example, when the target object speaks Chinese, the corresponding conversion language is determined to be Chinese for the text segment whose recognition language is English.

[0044] In another embodiment, a predetermined translation strategy is obtained, the translation strategy includes a conversion language determined from the languages ​​to be translated that is different from the corresponding recognition language, and the translation strategy is used to represent the translation between the recognition language and the conversion language. For example, the recognition language and the conversion language are Chinese and English, respectively. When the initial recognition text includes text segments corresponding to Chinese and English, respectively, it is determined that the conversion language corresponding to the Chinese text segment is English, and the conversion language corresponding to the English text segment is Chinese.

[0045] S203: Based on the text segments and the matching conversion languages, obtaining the prompt texts matching the respective audio segments in the audio to be translated; wherein the audio segments correspond to the text segments one by one.

[0046] In one embodiment, the audio to be translated is divided according to the obtained text segments to obtain audio segments corresponding to the text segments. For each audio segment, a prompt text matching the audio segment is generated according to the corresponding conversion language, and the prompt text is used to prompt the corresponding audio segment to be translated into a translation text matching the corresponding conversion language.

[0047] In a specific application scenario, when the recognition language of the text segment is Chinese and its corresponding conversion language is English, the generated prompt text is: "translate Chinese to English".

[0048] Please refer to Figure 3 , Figure 3 which Figure 1 is a schematic flowchart of another implementation manner corresponding to step S102 in

[0049] S301: Obtain the recognition characters matched by different recognition languages in the initial recognition text.

[0050] In one implementation manner, the initial recognition text includes a plurality of consecutive recognition characters. Determine the recognition characters corresponding to each recognition language, and count the number of recognition characters corresponding to each recognition language.

[0051] S302: Determine the translation strategy based on the number of recognition characters matched by each recognition language.

[0052] In one implementation manner, determine the translation strategy according to the number of recognition characters matched by each recognition language. Among them, the above translation strategy is used to determine the conversion language corresponding to the audio to be translated.

[0053] Specifically, according to the number of recognition characters corresponding to each recognition language, determine the proportion of the corresponding recognition language in the initial recognition text. According to the above proportion, determine the conversion text corresponding to the initial recognition text.

[0054] In one implementation scenario, for the proportion corresponding to each recognition language, obtain the maximum proportion. Obtain a preset ratio threshold. When the maximum proportion is greater than the ratio threshold, determine the corresponding conversion language according to the recognition language corresponding to the maximum proportion. Otherwise, for each recognition language, determine the corresponding conversion language. For example, when the recognition language corresponding to the maximum proportion is Chinese and the maximum proportion is greater than the ratio threshold, if the language spoken by the corresponding target object is English, then use English as the conversion language corresponding to the entire audio to be translated; or, when the recognition language corresponding to the maximum proportion is Chinese and the maximum proportion is less than the ratio threshold, if the language spoken by the corresponding target object is English, then determine multiple audio segments corresponding to the audio to be translated according to the recognition language, and determine that the conversion language corresponding to each audio segment is English.

[0055] S303: Obtain the prompt text based on the recognition language and the translation strategy.

[0056] In one implementation scenario, generate the prompt text according to the recognition language and the translation strategy. Among them, the prompt text includes a conversion language determined from the languages to be translated and different from the recognition language. The prompt text is used to prompt to translate the audio to be translated into a translation text matching the corresponding conversion language.

[0057] The above solution helps to improve the translation efficiency by determining the corresponding translation strategy according to the number of recognition characters.

[0058] In yet another embodiment, after obtaining the recognition characters matching different recognition languages in the initial recognition text, corresponding translation strategies are determined according to the recognition characters matching each recognition language.

[0059] Specifically, professional nouns in the initial recognition text are recognized, and the recognition language corresponding to the professional noun is used as the corresponding conversion language.

[0060] In one implementation scenario, corresponding professional nouns are determined according to special characters in the initial recognition text. By using the recognition language corresponding to the professional noun as the corresponding conversion language, it is ensured that professional nouns are not translated during subsequent translation, thus guaranteeing the accuracy of translation. For example, quotes in the initial recognition text are recognized, and the part within the quotes is regarded as a professional noun.

[0061] Please refer to Figure 4 , Figure 4 is Figure 1 a schematic flowchart of another embodiment corresponding to step S103 in

[0062] S401: Obtain the target audio encoding corresponding to the audio to be translated.

[0063] In one embodiment, the target audio encoding corresponding to the audio to be translated is obtained.

[0064] In one implementation scenario, to improve the execution efficiency, the target audio encoding is obtained through the process of generating the initial recognition text. The specific acquisition process can refer to the corresponding above embodiments and will not be elaborated in detail here.

[0065] S402: Obtain the reference encoding corresponding to each candidate vocabulary in the candidate vocabulary library.

[0066] In one embodiment, the candidate vocabulary library includes multiple candidate vocabularies, and each candidate vocabulary is encoded separately to obtain the corresponding reference encoding.

[0067] Specifically, a semantic encoding model is obtained, and the above candidate vocabularies are input into the semantic encoding model to obtain the corresponding reference encoding. Among them, the specific structure of the reference encoding model can refer to existing neural network structures.

[0068] S403: Based on the target audio encoding and the reference encoding, obtain the similarity scores between the audio to be translated and each candidate vocabulary.

[0069] In one embodiment, according to the target audio encoding and the reference encoding corresponding to each candidate vocabulary, the similarity scores between the audio to be translated and each candidate vocabulary are obtained.

[0070] Specifically, by calculating the cosine similarity between the target audio encoding and the reference encoding, and using this cosine similarity as the similarity score between the audio to be translated and the corresponding candidate words.

[0071] S404: Determine the reference word from the candidate words based on the similarity score.

[0072] In one embodiment, at least one reference word is selected from all candidate words according to the similarity scores between the audio to be translated and each candidate word.

[0073] Specifically, compare the similarity score with a preset first score threshold, and use the candidate word corresponding to the similarity score greater than the first score threshold as the reference word. Among them, the above first score threshold can be obtained by reverse deduction through multiple experiments, or can also be obtained by estimation by relevant technical personnel.

[0074] In another embodiment, for the similarity scores between the audio to be translated and each candidate word, use the candidate word corresponding to the maximum similarity score as the reference word.

[0075] The above solution provides a reference basis for the subsequent translation process of the audio to be translated by determining the reference word, thereby improving the translation accuracy.

[0076] Please refer to Figure 5 , Figure 5 is Figure 4 the schematic flowchart of another embodiment corresponding to step S404 in

[0077] S501: Determine at least one first screening word from all candidate words based on the similarity score.

[0078] In one embodiment, use the candidate word corresponding to the similarity score greater than the score threshold as the first screening word. Or, for the similarity scores between the audio to be translated and each candidate word, use the candidate word corresponding to the maximum similarity score as the first screening word.

[0079] S502: Obtain the matching scores between the initial recognition text and each candidate word, and determine at least one second screening word from all candidate words based on the matching scores.

[0080] In one embodiment, according to the initial recognition text and each candidate word, obtain the corresponding matching scores, and determine at least one second screening word from all candidate words according to the matching scores.

[0081] Specifically, obtain the edit distance between the initial recognition text and each candidate word. According to the above edit distance, determine the corresponding matching score. The edit distance is inversely proportional to the corresponding matching score, that is, the greater the edit distance, the smaller the corresponding matching score. Obtain a preset second score threshold, and use the candidate words corresponding to the matching scores greater than the second score threshold as the second screened words. Alternatively, for the matching scores between the audio to be translated and each candidate word, use the candidate word corresponding to the maximum matching score as the second screened word. Among them, the above second score threshold can be obtained by inverse deduction through multiple experiments, or can also be estimated by relevant technical personnel.

[0082] In another embodiment, perform fuzzy matching on the obtained initial recognition text and each candidate word to obtain the corresponding matching score. According to the matching score, determine at least one second screened word from all candidate words.

[0083] In addition, it should be noted that the execution order of the above steps S501 and S502 can also be other, for example, first execute step S502, and then execute step S501; or, execute steps S501 and S502 simultaneously.

[0084] S503: Determine the reference word based on the first screened word and the second screened word.

[0085] In one embodiment, both the first screened word and the second screened word are used as reference words.

[0086] In another embodiment, in response to obtaining multiple first screened words and / or multiple second screened words, obtain the same words among the multiple first screened words and the multiple second screened words, and use the above same words as reference words.

[0087] The above solution determines the reference word by combining the similarity score and the matching score, further improving the accuracy of determining the reference word and helping to improve the accuracy of subsequent translation.

[0088] Please refer to Figure 6 , Figure 6 is Figure 1 a schematic flowchart of another embodiment corresponding to step S104 in

[0089] S601: Use the target audio encoding corresponding to the audio to be translated as the first prompt information.

[0090] In one embodiment, use the target audio encoding corresponding to the audio to be translated as the first prompt information. The specific acquisition process of the target audio encoding can refer to the corresponding above embodiment.

[0091] S602: Obtain a second prompt message based on encoding the prompt text.

[0092] In one implementation, encode the prompt text obtained in the corresponding implementation above to obtain a second prompt message.

[0093] Specifically, input the prompt text into a semantic encoding model to encode the prompt text using the semantic encoding model, and use the obtained features after encoding as the second prompt message. The specific implementation process can refer to the corresponding implementation above.

[0094] S603: In response to the candidate word library including reference translation words that match each candidate word, obtain the reference translation words corresponding to the reference words, and obtain a third prompt message based on the reference words and their matching reference translation words; wherein, the third prompt message is used to represent the conversion relationship between the reference words and the reference translation words.

[0095] In one implementation, the candidate word library includes candidate words that match multiple different languages, and each candidate word is matched with a reference translation word corresponding to another language. For each reference word, construct a reference translation phrase based on the reference word and the corresponding reference translation word. Obtain the encoded features corresponding to the reference translation phrase, and use the encoded features as the third prompt message.

[0096] Specifically, for each obtained reference word, generate a reference translation phrase according to the reference word, the reference translation word corresponding to the reference word, and the conversion relationship between the two. Use the semantic encoding model to encode the reference translation phrase to obtain the third prompt message.

[0097] In a specific application scenario, when the determined reference word is "Good morning", and the obtained corresponding reference translation word is "good morning", the corresponding reference translation phrase is generated as: "Good morning" translate to "good morning". Encode it using the semantic encoding model to obtain the third prompt message.

[0098] In another implementation, in response to each candidate word in the candidate word library being matched with reference translation words corresponding to multiple other languages. For the obtained reference word, generate a reference translation phrase according to the corresponding conversion language, the reference word, the reference translation word corresponding to the reference word, and the conversion relationship between the two. Use the semantic encoding model to encode the reference translation phrase to obtain the third prompt message.

[0099] In a specific application scenario, when the determined reference word is "good morning" and the converted language that the audio to be translated matches is English, obtain the corresponding reference translation word "good morning", and generate the corresponding third prompt message: "good morning" translates to "good morning".

[0100] S604: Input the first prompt message, the second prompt message, and the third prompt message into the intelligent analysis model to obtain the translation text output by the intelligent analysis model.

[0101] In a real-time manner, use the intelligent analysis model to analyze according to the first prompt message, the second prompt message, and the third prompt message, so as to obtain the translation text output by the intelligent analysis model.

[0102] Specifically, concatenate the first prompt message, the second prompt message, and the third prompt message in sequence as the translation task text, which is used to prompt the intelligent analysis model to translate the audio to be translated into a translation text that matches the converted language. Use the intelligent analysis model to analyze the above translation task text and output the corresponding translation text.

[0103] In an implementation scenario, the intelligent analysis model is a large language model with relatively excellent data analysis capabilities. By inputting the generated translation task text into the intelligent analysis model, after the intelligent analysis model interprets it in detail, it outputs a translation text that matches the audio to be translated.

[0104] In a specific application scenario, the above large language model can include but is not limited to Deep Neural Networks (DNNs), Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM), and Generative Pretrained Transformer models, etc. There is no specific limitation on the specific structure and specific deployment of the large language model here. In addition, it should be noted that the specific structure and specific deployment of the intelligent analysis model mentioned in other implementation manners of this application can all refer to this implementation manner.

[0105] In another implementation manner, obtain a pre-set translation task template, fill in the obtained first prompt message, second prompt message, and third prompt message into the corresponding positions in the translation task template in sequence to obtain the translation task text. Input the translation task text into the intelligent analysis model to obtain the output translation text.

[0106] In a specific application scenario, the pre-set translation task template is: "You are a professional translator. Please translate the audio to be translated according to the first prompt information, the second prompt information, and the third prompt information, and output the corresponding translation text."

[0107] In the above solution, by using the intelligent analysis model in combination with various prompt information, the generated translation text corresponding to the audio to be translated has a relatively high accuracy, improving the effect of audio translation.

[0108] Please refer to Figure 7 , Figure 7 , which is a schematic structural diagram of an embodiment of the voice translation system of the present application. Specifically, the voice translation system includes an identification module 10, a first processing module 20, a second processing module 30, and a translation module 40 that are mutually coupled.

[0109] Specifically, the identification module 10 is used to obtain the input audio at at least one end, determine the language to be translated of the input audio, use each end of the input audio as the audio to be translated respectively, and determine the initial recognition text corresponding to the audio to be translated; wherein, at least one recognition language is matched with the initial recognition text.

[0110] The first processing module 20 is used to obtain the prompt text matching the audio to be translated based on the language to be translated, the initial recognition text, and its corresponding recognition language; wherein, the prompt text includes the conversion language matching the recognition language.

[0111] The second processing module 30 is used to obtain the reference vocabulary matching the initial recognition text from the candidate vocabulary library.

[0112] The translation module 40 is used to obtain the translation text corresponding to the audio to be translated based on the initial recognition text, the prompt text, and the reference vocabulary.

[0113] In an embodiment, multiple recognition languages are matched with the initial recognition text. The first processing module 20 obtains the prompt text matching the audio to be translated based on the language to be translated, the initial recognition text, and its corresponding recognition language, including: dividing the initial recognition text into multiple text segments based on the recognition language; wherein, each text segment is matched with the same recognition language, and adjacent text segments are matched with different recognition languages; determining the conversion language matched by each text segment based on all text segments and their matched recognition languages; obtaining the prompt text matched by each audio segment in the audio to be translated based on the text segment and its matched conversion language; wherein, the audio segment corresponds to the text segment one by one.

[0114] In one embodiment, the initial recognition text is matched with multiple recognized languages. The first processing module 20 obtains a prompt text that matches the audio to be translated based on the language to be translated, the initial recognition text, and its corresponding recognized language, including: obtaining the recognized characters matched by different recognized languages in the initial recognition text; determining a translation strategy based on the number of recognized characters matched by each recognized language; and obtaining the prompt text based on the recognized language and the translation strategy.

[0115] In one embodiment, the candidate word library includes multiple candidate words. The second processing module 30 obtains a reference word that matches the initial recognition text from the candidate word library, including: obtaining the target audio encoding corresponding to the audio to be translated; and obtaining the reference encoding corresponding to each candidate word in the candidate word library; obtaining the similarity score between the audio to be translated and each candidate word based on the target audio encoding and the reference encoding; and determining the reference word from all the candidate words based on the similarity score.

[0116] In one embodiment, the second processing module 30 determines a reference word from all the candidate words based on the similarity score, including: determining at least one first screened word from all the candidate words based on the similarity score; and obtaining the matching score between the initial recognition text and each candidate word, and determining at least one second screened word from all the candidate words based on the matching score; and determining the reference word based on the first screened word and the second screened word.

[0117] In one embodiment, the translation module 40 obtains a translation text corresponding to the audio to be translated based on the initial recognition text, the prompt text, and the reference word, including: using the target audio encoding corresponding to the audio to be translated as the first prompt information; and obtaining the second prompt information based on encoding the prompt text; and in response to the candidate word library including reference translation words that match each candidate word, obtaining the reference translation words corresponding to the reference words, and obtaining the third prompt information based on the reference words and their matching reference translation words, where the third prompt information is used to represent the conversion relationship between the reference words and the reference translation words; and inputting the first prompt information, the second prompt information, and the third prompt information into an intelligent analysis model to obtain the translation text output by the intelligent analysis model.

[0118] In one embodiment, the translation module 40 obtains the third prompt information based on the reference words and their matching reference translation words, including: for each reference word, constructing a reference translation phrase based on the reference word and the corresponding reference translation word; and obtaining the encoding feature corresponding to the reference translation phrase and using the encoding feature as the third prompt information.

[0119] Please refer to Figure 8 , Figure 8It is a schematic structural diagram of an embodiment of the electronic device of the present application. The electronic device includes: a memory 50 and a processor 60 that are coupled to each other. Program instructions are stored in the memory 50, and the processor 60 is configured to execute the program instructions to implement the methods described in any of the above embodiments. Specifically, the electronic device includes, but is not limited to: desktop computers, laptop computers, tablet computers, servers, etc., which are not limited herein. In addition, the processor 60 may also be referred to as a CPU (Center Processing Unit, central processing unit). The processor 60 may be an integrated circuit chip with signal processing capabilities. The processor 60 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 60 may be implemented jointly by integrated circuit chips.

[0120] Please refer to Figure 9 , Figure 9 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Program instructions 80 that can be run by a processor are stored on the computer-readable storage medium 70. When the program instructions 80 are executed by the processor, the methods described in any of the above embodiments are implemented.

[0121] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0122] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0123] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically separately for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0124] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0125] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A speech translation method, characterized in that: include: Acquire input audio from at least one end, determine the language to be translated of the input audio, use the input audio from each end as the audio to be translated, and determine the initial recognition text corresponding to the audio to be translated; wherein the initial recognition text matches at least one recognition language; Based on the language to be translated, the initial recognized text and the recognized language corresponding thereto, obtaining a prompt text matching the audio to be translated; wherein the prompt text includes a conversion language matching the recognized language, and the prompt text is used to prompt the audio to be translated to be translated into a translation text matching the corresponding conversion language; Acquire a reference vocabulary matching the initial recognition text from a candidate vocabulary library; Based on the initial recognition text, the prompt text and the reference vocabulary, obtaining a translation text corresponding to the audio to be translated; Wherein, the obtaining of the translation text corresponding to the audio to be translated based on the initial recognition text, the prompt text and the reference vocabulary comprises: encoding the target audio corresponding to the audio to be translated as the first prompt information; and, based on encoding the prompt text, obtaining the second prompt information; and, in response to the reference translation vocabulary matching each candidate vocabulary included in the candidate vocabulary library, obtaining the reference translation vocabulary corresponding to the reference vocabulary, and obtaining the third prompt information based on the reference vocabulary and the reference translation vocabulary matching it; and inputting the first prompt information, the second prompt information and the third prompt information into the intelligent analysis model to obtain the translation text output by the intelligent analysis model; wherein, the intelligent analysis model is a large language model with data analysis capabilities.

2. The method according to claim 1, characterized in that The initial recognition text matches a plurality of the recognition languages, and obtaining a prompt text matching the audio to be translated based on the language to be translated, the initial recognition text and the corresponding recognition language, comprises: Based on the recognition language, the initial recognition text is divided into a plurality of text segments; wherein each of the text segments is matched with the same recognition language, and adjacent text segments are matched with different recognition languages; Based on all the text segments and the identified languages ​​matched thereto, determining the conversion language matched by each of the text segments; Based on the text segments and the matching conversion languages, the prompt texts matched by the respective audio segments in the audio to be translated are obtained; wherein the audio segments correspond one to one with the text segments.

3. The method according to claim 1, characterized in that The initial recognition text matches a plurality of the recognition languages, and obtaining a prompt text matching the audio to be translated based on the language to be translated, the initial recognition text and the corresponding recognition language, comprises: Acquire recognition characters matching different recognition languages ​​in the initial recognition text; Determining a translation strategy based on the number of the recognized characters matched by each of the recognized languages; The prompt text is obtained based on the recognized language and the translation strategy.

4. The method according to claim 1, characterized in that: The candidate word library includes a plurality of candidate words, and the step of obtaining a reference word matching the initial recognition text from the candidate word library includes: Obtaining a target audio code corresponding to the audio to be translated; and Obtaining a reference code corresponding to each of the candidate words in the candidate word library; Based on the target audio code and the reference code, obtaining a similarity score between the audio to be translated and each of the candidate words; The reference vocabulary is determined from all the candidate vocabulary based on the similarity scores.

5. The method according to claim 4, characterized in that The step of determining the reference vocabulary from all the candidate vocabulary based on the similarity score comprises: Based on the similarity score, determining at least one first screening vocabulary from all the candidate vocabulary; and, Obtaining a matching score between the initial recognition text and each of the candidate words, and determining at least one second screening word from all the candidate words based on the matching score; The reference vocabulary is determined based on the first screening vocabulary and the second screening vocabulary.

6. The method according to claim 1, characterized in that The third prompt information is used to represent the conversion relationship between the reference vocabulary and the reference translation vocabulary.

7. The method according to claim 6, characterized in that The acquiring third prompt information based on the reference vocabulary and the reference translation vocabulary matched thereto includes: For each of the reference words, construct a reference translation phrase based on the reference word and the corresponding reference translation word; Obtain a coding feature corresponding to the reference translation phrase, and use the coding feature as the third prompt information.

8. A speech translation system, characterized in that: include: A recognition module, used to obtain input audio from at least one end, determine the language to be translated of the input audio, use the input audio from each end as the audio to be translated, and determine the initial recognition text corresponding to the audio to be translated; wherein the initial recognition text matches at least one recognition language; A first processing module is used to obtain a prompt text matching the audio to be translated based on the language to be translated, the initial recognition text and the corresponding recognition language; wherein the prompt text includes a conversion language matching the recognition language, and the prompt text is used to prompt the audio to be translated to be translated into a translation text matching the corresponding conversion language; A second processing module, configured to obtain a reference vocabulary matching the initial recognition text from a candidate vocabulary library; A translation module, used for obtaining a translation text corresponding to the audio to be translated based on the initial recognition text, the prompt text and the reference vocabulary; Wherein, the obtaining of the translation text corresponding to the audio to be translated based on the initial recognition text, the prompt text and the reference vocabulary comprises: encoding the target audio corresponding to the audio to be translated as the first prompt information; and, based on encoding the prompt text, obtaining the second prompt information; and, in response to the reference translation vocabulary matching each candidate vocabulary included in the candidate vocabulary library, obtaining the reference translation vocabulary corresponding to the reference vocabulary, and obtaining the third prompt information based on the reference vocabulary and the reference translation vocabulary matching it; and inputting the first prompt information, the second prompt information and the third prompt information into the intelligent analysis model to obtain the translation text output by the intelligent analysis model; wherein, the intelligent analysis model is a large language model with data analysis capabilities.

9. An electronic device, characterized in that: include: A memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Simultaneous translation method and device, smart car-mounted terminal and storage medium

    CN108595443A

  • Multi-user multi-language recognition and translation method and device

    CN113299276A