Speech translation method and related device

By performing language determination processing and speech dictation processing for translating audio, automatically identifying the language and realizing bilingual translation, the problem of cumbersome users manually selecting languages ​​in the prior art is solved, and the reliability and user experience of translation are improved.

CN120199231APending Publication Date: 2025-06-24IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417643.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing translation equipment requires users to manually select languages ​​during bilingual translation, which is cumbersome and can easily lead to translation failure.

Method used

By performing language determination processing on translating audio, the language is automatically recognized, and the speech dictation results are translated into the target language in the specified language pair.

Benefits of technology

It realizes automatic language recognition, simplifies user operation processes, reduces the possibility of translation failure, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199231A_ABST
    Figure CN120199231A_ABST
Patent Text Reader

Abstract

The invention discloses a speech translation method and a related device, and relates to the field of language processing, and the method comprises the steps: carrying out the language judgment processing of a to-be-translated audio when responding to a bilingual translation demand between a specified language pair, and automatically determining a source language of the to-be-translated audio and a target language which needs to be translated, and then the voice dictation result of the to-be-translated audio is translated, namely the first dictation result corresponding to the language judgment result, and then the first translation result obtained by translation is output, so that the bilingual translation task between the specified language pairs is finally realized. According to the method and the device, the purpose of automatically selecting the language is achieved through language judgment, the user does not need to know the language of the current speaker and select the language according to the language, the operation process of the user in a bilingual translation scene is simplified, the translation failure possibility caused by user operation errors is reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of language processing technologies, and in particular, to a speech translation method and related devices. Background Art

[0002] With the development of economic globalization, the demand for cross-language communication, especially bilingual communication, is increasing day by day. The bilingual translation function plays an important role in scenarios such as tourism and business.

[0003] When existing translation devices perform bilingual translation, they require the user to manually select the language of the current speaker before speaking to specify the translation direction of the language, and then the translation device performs directional recognition and translation accordingly. That is to say, whether the speaker is switched or the language needs to be switched, the user needs to re-specify the current language, and the operation is cumbersome; moreover, in the actual use process, it is easy to occur that the translation fails because the language selected by the user does not match the actual language of the speaker, which is not conducive to the user experience. Summary of the Invention

[0004] In view of the above problems, this application provides a speech translation method and related devices to achieve the speech translation task of automatically identifying languages.

[0005] The specific solutions are as follows:

[0006] The first aspect of this application provides a speech translation method, including:

[0007] In response to a bilingual translation requirement, perform language determination processing and speech dictation processing on the audio to be translated; the result of the speech dictation processing includes a first dictation result corresponding to the language determination result of the audio to be translated;

[0008] Perform a first translation processing on the first dictation result to obtain a first translation result; the target language of the first translation processing is the language in the language pair specified by the bilingual translation requirement that does not match the language determination result;

[0009] Output the first translation result as the translation result of the audio to be translated.

[0010] The second aspect of this application provides a speech translation method, including:

[0011] In response to the user's bilingual translation requirement, display a bilingual translation interface corresponding to the specified language pair; the specified language pair includes two languages specified by the user;

[0012] In response to the user's trigger operation on the bilingual translation button, start recording audio;

[0013] After the user stops recording, stop the recording and display the first dictation result of the recorded audio to be translated on the bilingual translation interface. The first dictation result is the dictation result corresponding to the language determination result obtained by performing speech dictation processing on the audio to be translated, and the language determination result is obtained by performing language determination processing on the audio to be translated.

[0014] Output the first translation result, which is obtained by performing first translation processing on the first dictation result. The target language of the first translation processing is the language in the specified language pair that does not match the language determination result.

[0015] A third aspect of the present application provides a voice translation device, including:

[0016] A voice processing unit, configured to, in response to a bilingual translation requirement, perform language determination processing and speech dictation processing on the audio to be translated; the result of the speech dictation processing includes a first dictation result corresponding to the language determination result of the audio to be translated; and perform first translation processing on the first dictation result to obtain a first translation result; the target language of the first translation processing is the language in the specified language pair that does not match the language determination result for the bilingual translation requirement.

[0017] A result output unit, configured to output the first translation result as the translation result of the audio to be translated.

[0018] A fourth aspect of the present application provides a voice translation device, including:

[0019] A first display module, configured to, in response to the user's bilingual translation requirement, display a bilingual translation interface corresponding to a specified language pair; the specified language pair includes two languages specified by the user.

[0020] A bilingual translation module, configured to, in response to the user's triggering operation on the bilingual translation button, start recording audio; and stop recording after the user stops recording.

[0021] A second display module, configured to, after stopping recording, display the first dictation result of the recorded audio to be translated on the bilingual translation interface; wherein, the first dictation result is the dictation result corresponding to the language determination result obtained by performing speech dictation processing on the audio to be translated, and the language determination result is obtained by performing language determination processing on the audio to be translated.

[0022] A result output module, configured to output a first translation result, which is obtained by performing first translation processing on the first dictation result. The target language of the first translation processing is the language in the specified language pair that does not match the language determination result.

[0023] The fifth aspect of the present application provides a translation device, including at least one processor and a memory connected to the processor, wherein:

[0024] The memory is used to store computer programs;

[0025] The processor is used to execute the computer programs to implement the speech translation method described in the first aspect or the second aspect above.

[0026] The sixth aspect of the present application provides a storage medium carrying one or more computer programs, which can enable an electronic device to implement the speech translation method described in the first aspect or the second aspect above when the one or more computer programs are executed by the electronic device.

[0027] The seventh aspect of the present application provides a computer program product, including computer-readable instructions, which can enable an electronic device to implement the speech translation method described in the first aspect or the second aspect above when the computer-readable instructions run on the electronic device.

[0028] By means of the above technical solutions, when the present application responds to the bilingual translation requirements between specified language pairs, through the language determination process for the audio to be translated, the source language of the audio to be translated and the target language to be translated are automatically determined, and then the speech dictation result of the audio to be translated, that is, the first dictation result corresponding to the language determination result, is translated, and then the translated first translation result is output, finally realizing the bilingual translation task between specified language pairs. The present application realizes the purpose of automatically selecting languages through language determination, without the user having to explicitly determine the language of the current speaker and select a language accordingly, simplifies the operation process of the user in the bilingual translation scenario, and reduces the possibility of translation failure caused by user operation errors, which helps to improve the user experience. Description of the Drawings

[0029] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0030] Figure 1 A schematic diagram of the bilingual translation process based on an existing translator is shown;

[0031] Figure 2 A schematic diagram of an implementation system architecture of the speech translation method provided by an embodiment of the present application;

[0032] Figure 3 A schematic diagram of the process of a speech translation method provided by the present application;

[0033] Figure 4 Illustrates a schematic diagram of the language determination process when both language pairs are languages supported by the large model;

[0034] Figure 5 Illustrates a schematic diagram of the language determination process when one of the language pairs is a language supported by the large model;

[0035] Figure 6 Shows a schematic diagram of a speech translation process;

[0036] Figure 7 Is a schematic structural diagram of a speech translation device provided by the present application;

[0037] Figure 8 Is a schematic flow diagram of another speech translation method provided by the present application;

[0038] Figure 9 Is a schematic structural diagram of another speech translation device provided by the present application;

[0039] Figure 10 Is a schematic hardware structure diagram of a translation device provided by the present application. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0041] The applicant of this case has found through research that when the existing translation machine realizes the bilingual translation function, the user needs to specify the language translation direction to enable the translation machine to perform directional dictation, and then translate according to the directional dictation result of the voice, and finally complete the translation task. Exemplarily, Figure 1 Taking Chinese - English translation as an example, a schematic diagram of the bilingual translation process based on the existing translation machine is shown. Combining Figure 1As shown in the figure, the translation interface displayed according to the languages of the two parties in the conversation selected by the user (i.e., Chinese and English) includes Chinese buttons and English buttons. On this basis, the user first clarifies the language of the current speaker, and then clicks the button corresponding to the language spoken by the current speaker, so that the translator can identify the voice content, determine the translation direction and perform translation output according to the language corresponding to the button clicked by the user. Suppose the language spoken by the current speaker is Chinese and a Chinese-to-English translation is required. Then the user needs to click the Chinese button so that the translator can determine Chinese as the source language and English as the target language for translation, and record the speech content of the current speaker. Then the translator identifies the recorded speech content according to the determined source language (i.e., Chinese). After obtaining the Chinese content, it performs a translation from the source language to the target language (i.e., Chinese to English), and outputs the corresponding English translation after the translation is completed. The reverse translation process, i.e., the English-to-Chinese translation process, is similar and can be referred to the above description.

[0042] Based on the above content, existing translators bind audio input and language selection to the user's button click operation. The user can start speaking while selecting a language by pressing a button, so that the translator can determine the language of the current speaker and the translation direction by identifying the button clicked by the user, and then decode according to the determined language of the speaker and translate according to the determined translation direction. However, during the conversation, if the user presses the wrong button, or the user presses the button for the other party but the other party does not speak, or the other party preempts the user's right to speak, etc., the language of the audio data received by the translator will be inconsistent with the language selected by the user, resulting in the abnormal progress of the current round of translation.

[0043] To solve the above problems, the present application provides a voice translation method and related device, which solves the problems of cumbersome process caused by user language selection and abnormal translation caused by incorrect user language selection by automatically identifying the language of the user input audio, thereby optimizing the user experience to a certain extent.

[0044] The present application provides a voice translation method, which can be applied to the system architecture as shown in Figure 2 The figure. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 2 Taking including one server as an example for illustration).

[0045] Either the terminal 100 or the server 200 can be used alone to execute the voice translation method provided by the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used in cooperation to execute the voice translation method provided in the embodiments of the present application.

[0046] The terminal 100 in the embodiments of the present application may be a translator, a mobile phone, a tablet computer, a notebook computer, etc., and the embodiments of the present application do not make any restrictions on this.

[0047] An embodiment of the present application provides a voice translation method. Taking the application of this method to a computer device as an example, the computer device may specifically be Figure 1 the terminal 100 in [description] or a system composed of the terminal 100 and the server 200.

[0048] Figure 3 is a flowchart of a voice translation method shown according to an embodiment of the present application. This method can be applied to a computer device, and the computer device may specifically be Figure 2 the terminal 100 in [description] or a system composed of the terminal 100 and the server 200. Combining Figure 3 as shown, this method may include the following steps:

[0049] Step S101, in response to a bilingual translation requirement, perform language determination processing and speech dictation processing on the audio to be translated.

[0050] The above-mentioned bilingual translation requirement may refer to obtaining the audio to be translated required for this translation after the user specifies the language range of bilingual translation. That is to say, in the bilingual translation scenario of a specified language range (i.e., language pair), during one round of translation, the user only needs to provide the audio data of the user's speech to the computer device implementing the voice translation solution, without the user manually selecting the language of the audio data. Subsequently, the language of the input audio data, that is, the speaker's language, is automatically determined through language determination processing, and the target language to be translated, that is, the other party's language, is determined. Based on this, in the bilingual translation scenario, the user can input voice by interacting with a fixed button, without considering who the current speaker is and without selecting the current language, such as selecting and clicking the button corresponding to the current language; that is to say, the current speaker's language is no longer restricted by the language corresponding to the button. With a single button, the user can use the computer device applying the solution of the present application more freely.

[0051] The above-mentioned speech dictation processing is the processing of converting audio data into text data; the result of the speech dictation processing may include a first dictation result corresponding to the language determination result of the audio to be translated, that is, the obtained dictation result may include the text data of the determined language, where the determined language is the language that matches the above-mentioned language determination result, that is, the source language of the audio to be translated determined.

[0052] Step S102, perform a first translation process on the first dictation result to obtain a first translation result.

[0053] Among them, the source language of the first translation process is the language determination result determined above, and the target language of the first translation process is the language in the language pair specified by the bilingual translation requirement that does not match the language determination result. In addition, the above-mentioned first translation process refers to text translation processing, that is, the obtained first translation result is text data.

[0054] Step S103: Output the first translation result as the translation result of the audio to be translated.

[0055] Exemplarily, the output of the above translation result may include: output of the translated text, such as displaying the first translation result on the screen. In another possible implementation, the output of the above translation result may further include: output of the translated audio converted from the translated text, such as playing the first audio data converted from the first translation result through a speaker. Based on this, after obtaining the first translation result, it may further include: synthesizing and broadcasting the first translation result. Optionally, the output of the above translation result may further include: output of the input text, such as displaying the first dictation result on the screen.

[0056] Based on the above content, in response to the bilingual translation requirement between the specified language pairs in this embodiment, by performing language determination processing on the audio to be translated, the source language of the audio to be translated and the target language to be translated are automatically determined, and then the speech dictation result of the audio to be translated, that is, the first dictation result corresponding to the language determination result, is translated, and then the obtained first translation result is output, finally realizing the bilingual translation task between the specified language pairs. This application achieves the purpose of automatically selecting a language through language determination, without the user having to explicitly identify the language of the current speaker and select a language accordingly, simplifies the user operation process in the bilingual translation scenario, and reduces the possibility of translation failure caused by user operation errors, which helps to improve the user experience.

[0057] Next, an exemplary description of the above language determination scheme is given.

[0058] In one or more embodiments provided by the present application, the above step S101 of performing language determination processing on the audio to be translated may include: calling a pre-trained language recognition model to determine the language of the audio to be translated.

[0059] Among them, the pre-trained language recognition model may be trained based on audio sample data of language pairs marked with corresponding language labels, that is, this model may be a recognition model suitable for implementing the language recognition task between specified language pairs.

[0060] The above solution uses a model applicable to this bilingual translation task to achieve language identification processing. In another possible implementation, calling a pre-trained language identification model to determine the language of the audio to be translated may include calling a large model to determine the language of the audio to be translated among the language pairs specified in the bilingual translation requirement.

[0061] In one or more embodiments provided by the present application, the above step S101 of performing language determination processing on the audio to be translated may include the following steps S201-S202:

[0062] Step S201: Call a large model to perform language identification processing on the audio to be translated to obtain the original language score of the audio to be translated.

[0063] Wherein, the original language score may include scores of at least one supported language, and the score of any supported language is positively correlated with the possibility that the audio to be translated belongs to the supported language. The above supported languages may refer to the languages supported by the large model in the language identification task. The large model can identify the corresponding language based on speech recognition and output it in the form of the possible language scores of the speech data. In one possible implementation, the large model can output the scores of all supported languages respectively; in another possible implementation, the large model can only output the scores of a preset number of supported languages with the highest possibility respectively; the present application does not make a limitation.

[0064] In addition, the format of the audio to be translated may match the audio format supported by the language identification function of the large model. For example, the audio data of the audio compression format speex7 of the audio to be translated, where 7 represents the compression quality.

[0065] Step S202: Determine the language determination result from the language pairs according to the original language score.

[0066] The above solution uses the general language identification ability of the large model to perform non-directional language identification without specifying the language range, and then, based on the output of the large model, combines the previously specified language pair range of the user to determine the language of the audio to be translated, and finally realizes the language identification processing applicable to the specified language range. This embodiment realizes the high-precision language determination task to a certain extent through the language identification ability of the large model and the further determination within the specified language range, providing a basis for ensuring the realization of this round of translation. In addition, the above solution does not require targeted improvement of the large model, saving the implementation cost of the solution.

[0067] In one or more embodiments provided by the present application, the above step S201 of calling a large model to perform language identification processing on the audio to be translated to obtain the original language score of the audio to be translated may include the following step A:

[0068] Step A: Extract an audio segment of at least a second preset duration from the audio to be translated, input the extracted audio segment into the large model for the large model to perform language identification, and use the original language score of the audio segment output by the large model as the original language score of the audio to be translated.

[0069] It should be noted that if the total duration of the audio to be translated is less than the second preset duration, the complete audio to be translated can be used as the extracted audio segment and input into the large model.

[0070] Exemplarily, the above-mentioned second preset duration can be set according to the language identification ability of the large model for audio data. Specifically, the second preset duration can be positively correlated with the language identification ability of the large model for audio data, and the audio data herein can specifically refer to the audio data of the language pair specified by the bilingual translation requirement.

[0071] In another possible implementation, the above-mentioned step S201, which calls the large model to perform language identification processing on the audio to be translated to obtain the original language score of the audio to be translated, may include the following step B:

[0072] Step B: Divide the audio to be translated into clauses, input at least one clause (such as the first clause) obtained from the division into the large model for the large model to perform language identification, and use the original language score of the at least one clause output by the large model as the original language score of the audio to be translated.

[0073] With the above solution, the data volume for the large model to perform language identification processing can be reduced by means of audio segment extraction or clause division. Based on this, to a certain extent, the problems of slow language determination speed and low efficiency caused by the too long duration of the audio to be translated are solved, which helps to improve the speech translation efficiency and thus optimize the user experience.

[0074] The following gives an exemplary description of the further language determination process.

[0075] In one or more embodiments provided by the present application, the above-mentioned step S202, which determines the language determination result from the language pair according to the original language score, may include the following step C:

[0076] Step C: When the language identification accuracy of the large model for the specified language audio data is higher than the preset threshold, determine the language determination result according to the score ranking of the languages constituting the language pair in the original language score.

[0077] It should be noted that the language recognition accuracy of the large model for the audio data of any supported language being higher than the corresponding threshold can indicate that the score of the supported language output by the large model is reliable. The specified language audio data is the audio data of any language in the language pair. Based on this, the recognition accuracy of the large model for the specified language audio data being higher than the preset threshold can indicate that the scores of the languages that make up the language pair in the original language scores output by the large model are reliable. In addition, different languages can correspond to the same threshold, such as 90%, or different thresholds, which is not limited in this case; the specific value of the preset threshold is also not limited in this application.

[0078] Since the score of any supported language is positively correlated with the possibility that the input audio belongs to that language, in this paper, \(r_i\) is used to represent the score ranking of language \(i\), where \(r_i\) is a positive integer, and the value of \(r_i\) indicates that among the language scores output by the large model, the score of language \(i\) ranks \(r_i\), that is, the possibility that the audio data input to the large model belongs to language \(i\) ranks \(r_i\). For example, if the score of a language ranks first in the original language scores, that is, the score ranking is 1, then within the range of supported languages of the large model, the possibility that the audio to be translated is the audio data of this language is the highest. On this basis, the score ranking of any supported language can represent the relative possibility of that language, that is, the language with a higher score and a smaller score ranking is more likely to be the language of the audio data input to the large model.

[0079] In one or more embodiments provided by this application, the above step C, determining the language determination result according to the score rankings of the languages that make up the language pair in the original language scores, may include the following steps S301 - S302:

[0080] Step S301, when both languages in the language pair are supported languages of the large model, determine the score rankings of the first language and the second language that make up the language pair according to the original language scores.

[0081] Among them, the first language satisfies that the recognition accuracy of the large model for the audio data of the first language is higher than the preset threshold, and the recognition accuracy of the large model for the audio data of the first language is higher than the recognition accuracy of the large model for the audio data of the second language. In other words, the first language is the language in the language pair that satisfies the accuracy condition and has a higher accuracy, and the accuracy condition is that the recognition accuracy of the large model is higher than the preset threshold; in addition, the second language may or may not satisfy the accuracy condition. The recognition accuracy of the large model for each language can be prior knowledge, obtained in advance through testing or other means.

[0082] In addition, since there are certain differences in the recognition accuracy of the large model for input audio of the same language with different durations, the language can be determined by combining the duration of the audio to be translated input into the large model and the ranking of the scores of the two languages determined in the above step S301. Specifically, Figure 4 illustrates a schematic diagram of the language determination process when both languages in the language pair are supported by the large model. The determination process will be described in combination with Figure 4 and the following content.

[0083] Step S302: When the duration of the audio to be translated input into the large model is not less than the first preset duration, if the ranking of the score of the first language is not greater than R1, and the ranking of the score of the first language is less than the ranking of the score of the second language, then the first language is determined as the language determination result; otherwise, the second language is determined as the language determination result.

[0084] On this basis, step C: Determining the language determination result according to the ranking of the scores of the languages constituting the language pair in the original language scores may further include the following step S303:

[0085] Step S303: When the duration of the audio to be translated input into the large model is less than the first preset duration, if the ranking of the score of the first language is not greater than R2, and the ranking of the score of the first language is less than the ranking of the score of the second language, then the first language is determined as the language determination result; otherwise, the second language is determined as the language determination result.

[0086] Among them, both R1 and R2 are set positive integers, and R1 < R2. It should be noted that the longer the audio duration recognized by the large model, the higher the recognition accuracy of the large model, and the higher the probability that the audio to be translated does not belong to the first language due to the larger ranking of the score of the first language output by the large model. Based on this, the score ranking threshold R1 when the audio duration input into the large model is longer is less than the score ranking threshold R2 when the audio duration input into the large model is shorter. Optionally, R1 is 1 and R2 is 2.

[0087] In addition, the above first preset duration can be set according to the recognition accuracy of the large model for audio data of different durations, especially according to the recognition accuracy of the large model for audio data of different durations of the first language. Optionally, for different first languages, the corresponding first preset durations can be the same or different, which is not limited in this application.

[0088] Exemplarily, the probability of accurately identifying the audio data of a certain language by a large model can represent the recognition accuracy of the large model for the audio data of that language. Suppose in a Chinese-Afro-foreign bilingual translation scenario, the probability that the score ranking of the Chinese audio output by the large model with a duration less than 1 second is no greater than 2 is 98%, and the probability that the score ranking of the Chinese audio output by the large model with a duration of no less than 1 second is no greater than 1 is 98%. Based on this, the first preset duration corresponding to Chinese can be 1 second. Similarly, suppose in an English-Afro-foreign bilingual translation scenario, the probability that the score ranking of the English audio output by the large model with a duration less than 0.5 second is no greater than 2 is 93%, the probability that the score ranking of the English audio output by the large model with a duration of no less than 0.5 second and less than 1.5 seconds is no greater than 2 is 96%, and the probability that the score ranking of the English audio output by the large model with a duration of no less than 1.5 seconds is no greater than 1 is 98%. Based on this, the first preset duration corresponding to English can be 1.5 seconds.

[0089] Based on the above, the present application utilizes the recognition advantage of the large model for the specified language audio data, compares the score rankings, especially compares the score ranking of the first language with the corresponding threshold, and compares the relative rankings of the first language and the second language, thereby improving the accuracy and efficiency of language determination to a certain extent.

[0090] In one or more embodiments provided by the present application, the above step C, determining the language determination result according to the score ranking of the language constituting the language pair in the original language score, may include the following steps S401-S402:

[0091] Step S401, when one of the languages in the language pair is a supported language of the large model, determining the score ranking of the first language constituting the language pair according to the original language score.

[0092] Among them, for the first language, the recognition accuracy of the large model for the audio data of the first language is higher than the preset threshold; correspondingly, the second language is the non-large model language in the language pair, that is, the language not supported by the large model.

[0093] It should be noted that since the second language is not a supported language of the large model, it is difficult to directly determine whether the audio to be translated belongs to the second language based on the output of the large model; however, the first language is a supported language of the large model and meets the accuracy condition. Based on this, the language determination can be performed by combining the duration of the audio to be translated input into the large model and the score ranking of the first language determined in the above step S401. Specifically, Figure 5 Illustrated is a schematic diagram of the language determination process when one of the languages in the language pair is a supported language of the large model. The determination process will be described in combination with Figure 5 and the following content.

[0094] Step S402: When the duration of the audio to be translated input to the large model is not less than the first preset duration, if the score ranking of the first language is not greater than R1, then determine the first language as the language determination result; otherwise, determine the second language as the language determination result.

[0095] On the basis of the above, step C: determining the language determination result according to the score ranking of the languages constituting the language pair in the original language scores may further include the following step S403:

[0096] Step S403: When the duration of the audio to be translated input to the large model is less than the first preset duration, if the score ranking of the first language is not greater than R2, then determine the first language as the language determination result; otherwise, determine the second language as the language determination result.

[0097] Among them, R1 < R2. It should be noted that the value of the score ranking threshold R1 described in step S302 and step S402 may be the same or different, and the value of the score ranking threshold R2 described in step S303 and step S403 may be the same or different. This solution does not make a limit and can be set according to the recognition accuracy of the large model for the first language, the kinship between the second language that does not belong to the languages supported by the large model and the languages supported by the large model, etc.

[0098] Based on the above content, the present application utilizes the recognition advantage of the large model for the audio data of the specified language, compares the score rankings, especially compares the score ranking of the first language with the corresponding threshold, and improves the accuracy and efficiency of language determination to a certain extent. Moreover, this solution can be applied to the bilingual translation scenario between the dominant language (i.e., the language that meets the accuracy condition) and the unsupported language, enabling this solution to meet the high-precision bilingual automatic determination requirements between the dominant language and a large number of languages, with strong applicability, and providing a basis for supporting the mutual translation tasks between the dominant language and a large number of languages in the bilingual automatic determination mode.

[0099] Assume that the large model has a high determination accuracy for Chinese and English audio data, Chinese and English are the dominant languages of the large model, and the recognition ability of the large model for Chinese audio data is better than that for English. The following will illustrate the language determination solution provided by the present application in combination with a specific bilingual translation scenario.

[0100] Combined with the above description and Figure 4As shown, in the bilingual translation scenario of Chinese and foreign language F1 (such as English) supported by the large model, Chinese meets the accuracy condition and the large model's recognition ability for Chinese audio data is better than that for English. Chinese is taken as the first language and foreign language F1 as the second language. If the duration of the audio to be translated input into the large model is less than the first preset duration corresponding to Chinese (i.e., 1 second), it is determined whether the score ranking of the first language (i.e., Chinese) is less than or equal to 2 (i.e., R2) and the score ranking of the second language (i.e., foreign language F1) is greater than the score ranking of Chinese. If so, the language determination result is determined to be Chinese; otherwise, the language determination result is determined to be foreign language F1. If the duration of the audio to be translated input into the large model is not less than 1 second, it is determined whether the score ranking of the first language (i.e., Chinese) is less than or equal to 1 (i.e., R1). If so, the language determination result is determined to be Chinese; otherwise, the language determination result is determined to be foreign language F1. Since the minimum score ranking of a language is 1, when R1 is 1, there is no need to compare the score rankings of the two languages.

[0101] Combined with the above description and Figure 5 As shown, in the bilingual translation scenario of Chinese and foreign language F2 not supported by the large model, Chinese meets the accuracy condition and the large model does not support foreign language F2. Chinese is taken as the first language and foreign language F2 as the second language. If the duration of the audio to be translated input into the large model is less than 1 second, it is determined whether the score ranking of the first language (i.e., Chinese) is less than or equal to 2. If so, the language determination result is determined to be Chinese; otherwise, the language determination result is determined to be foreign language. If the duration of the audio to be translated input into the large model is not less than 1 second, it is determined whether the score ranking of the first language (i.e., Chinese) is less than or equal to 1. If so, the language determination result is determined to be Chinese; otherwise, the language determination result is determined to be foreign language F2.

[0102] Combined with the above description and Figure 4 As shown, in the bilingual translation scenario of English and non-Chinese foreign language F3 supported by the large model, English meets the accuracy condition and the large model's recognition ability for English audio data is better than that for foreign language F3. English is taken as the first language and foreign language F3 as the second language. If the duration of the audio to be translated input into the large model is less than the first preset duration corresponding to English (i.e., 1.5 seconds), it is determined whether the score ranking of the first language (i.e., English) is less than or equal to 2 and the score ranking of the second language (i.e., foreign language F3) is greater than the score ranking of English. If so, the language determination result is determined to be English; otherwise, the language determination result is determined to be foreign language F3. If the duration of the audio to be translated input into the large model is not less than 1.5 seconds, it is determined whether the score ranking of the first language (i.e., English) is less than or equal to 1. If so, the language determination result is determined to be English; otherwise, the language determination result is determined to be foreign language F3.

[0103] For the language determination scheme in the bilingual translation scenario of English and a foreign language not supported by the large model, reference can be made to the Chinese and foreign language F2 bilingual translation scenario described above, which will not be elaborated here.

[0104] In one or more embodiments provided by the present application, the above step S101 of performing speech dictation processing on the audio to be translated may include:

[0105] Dictate the audio to be translated according to the two languages that make up the language pair respectively.

[0106] Based on the above, two dictation results can be obtained simultaneously through dictation processing. Among them, the language of the first dictation result matches the language determination result and serves as the display path for translation display; the second dictation result that does not match the language determination result can be used as the cache path and cached locally for switching translation in the case of abnormal language determination.

[0107] In addition, the above solution performs dual-channel speech dictation, so that the speech dictation processing does not need to wait for the result of the language determination processing. That is, in step S101, the language determination and the two-channel speech dictation can be started simultaneously. Based on this, the processing speed can be improved to a certain extent, which helps to achieve an efficient bilingual translation task and thus optimize the user experience.

[0108] On the basis of the above, after outputting the first translation result as the translation result of the audio to be translated, the following steps S104 - S105 may further be included:

[0109] Step S104: In response to the user's switching request, perform a second translation process on the second dictation result to obtain a second translation result.

[0110] Among them, the second dictation result is the dictation result different from the first dictation result among the speech dictation results of the two languages obtained through the speech dictation processing. It should be noted that after the user performs a switching operation and issues a switching request, the source language of the audio to be translated determined is the language in the language pair that does not match the language determination result. Correspondingly, the target language of the second translation process is the language in the language pair that matches the language determination result. That is to say, the second translation process is opposite to the first translation process in the translation direction.

[0111] Step S105: Output the second translation result as the translation result of the audio to be translated.

[0112] In addition, the description of the second translation process, the second translation result, and the output of the second translation result can refer to the above description and will not be elaborated here.

[0113] Based on the above, in this embodiment, when the user determines that the language determination result is abnormal, the output translation result is quickly switched, enabling the user to obtain accurate translation results quickly, even without delay, thereby improving the usability and reliability of the computer device applying the solution of this embodiment.

[0114] In a possible implementation, capabilities such as language determination processing, speech dictation processing, and translation processing (specifically including first translation processing and second translation processing) can be integrated on the terminal 100.

[0115] To reduce the cost of the terminal, in a possible implementation, language determination processing, speech dictation processing, and translation processing (specifically including first translation processing and second translation processing) can be implemented by the server 200, and the server 200 provides capabilities such as language determination processing, speech dictation processing, and translation processing for multiple terminals 100.

[0116] Optionally, the server 200 can include an AI capability platform and a cloud platform, and the cloud platform deploys a large model. Exemplarily, Figure 6 shows a schematic diagram of a speech translation process. Combining Figure 6 as shown, specifically:

[0117] The terminal 100 (such as a translation device) starts recording under the trigger of the user. After the recording ends, it sends the audio to be translated and language pair information indicating the language pair (constituted by language A and language B) to the AI capability platform.

[0118] The AI capability platform calls the cloud platform for three-way parallel processing, specifically including: calling the large model deployed on the cloud platform to perform language recognition on the audio to be translated, and calling two cloud engines to perform speech dictation on the audio to be translated according to language A and language B respectively.

[0119] The AI capability platform receives the dictation results of language A, the dictation results of language B, and the language recognition result (expressed as the original language score) returned by the cloud platform, determines the language of the audio to be translated based on the original language score, and obtains the language determination result. Assume it is language A (or language B). Figure 6 Taking language A as an example for illustration;

[0120] The AI capability platform calls the cloud platform to perform text translation processing on the dictation result of the language A (or language B) corresponding to the language determination result, and the target language of the text translation processing is the language B (or language A) that does not match the language determination result; and receives the translation result of language B (or language A) fed back by the cloud platform.

[0121] The AI capability platform feeds back the dictation result of the audio to be translated to the terminal 100, including the dictation result in language A (or language B), and feeds back the translation result in language B (or language A);

[0122] Finally, the terminal 100 outputs the translation result of the audio to be translated to the user.

[0123] Specifically, the dictation result of the audio to be translated fed back by the AI capability platform to the terminal 100 may also include the dictation result in language B (or language A). After receiving the dictation result in language B (or language A) that does not match the language determination result, the terminal 100 caches it to support subsequent switching operations.

[0124] On the basis of the above, when the terminal 100 outputs the translation result of the audio to be translated to the user, specifically, after outputting the translation result in language B (or language A) that does not match the language determination result, the voice translation process may further include:

[0125] Under the trigger of the user, the terminal 100 switches the output translation result, which may specifically include:

[0126] The terminal 100 sends the dictation result in language B (or language A) to the AI capability platform so that the AI capability platform performs text translation processing, and the target language of the text translation processing is language A (or language B) that does not match the language determination result;

[0127] The AI capability platform calls the cloud platform to perform text translation processing on the dictation result in language B (or language A), and receives the translation result in language A (or language B) fed back by the cloud platform;

[0128] The AI capability platform feeds back the translation result in language A (or language B) to the terminal 100, and then the terminal 100 outputs the translation result of the switched audio to be translated to the user.

[0129] In addition, the terminal 100 may be composed of a bilingual service and a translation service. The bilingual service is used to interact with the user in a bilingual translation scenario, such as receiving user triggers or outputting to the user; the translation service is used to receive the trigger of the bilingual service and interact with the server 200 (especially the AI capability platform), such as sending data / requests to the AI capability platform or receiving the processing results fed back by the AI capability platform, such as language determination results, voice dictation results, translation results, etc.

[0130] Next, the voice translation device provided in the embodiments of the present application will be described. The voice translation device described below can be correspondingly referred to the voice translation method described above.

[0131] Figure 7 It is a schematic structural diagram of a voice translation device disclosed in the embodiments of the present application. AsFigure 7 As shown, the device may include: a voice processing unit 11 and a result output unit 12.

[0132] The voice processing unit 11 is configured to perform a language determination process and a speech dictation process on the audio to be translated in response to a bilingual translation requirement, and the result of the speech dictation process includes a first dictation result corresponding to the language determination result of the audio to be translated;

[0133] The voice processing unit 11 is further configured to perform a first translation process on the first dictation result to obtain a first translation result; the target language of the first translation process is the language in the language pair specified by the bilingual translation requirement that does not match the language determination result;

[0134] The result output unit 12 is configured to output the first translation result as the translation result of the audio to be translated.

[0135] In one or more embodiments provided by the present application, the process of the voice processing unit 11 performing a language determination process on the audio to be translated may include:

[0136] Invoking a large model to perform a language recognition process on the audio to be translated to obtain an original language score of the audio to be translated; the original language score includes scores of at least one supported language, and the score of any supported language is positively correlated with the possibility that the audio to be translated belongs to the supported language;

[0137] Determining the language determination result from the language pair according to the original language score.

[0138] In one or more embodiments provided by the present application, the process of the voice processing unit 11 determining the language determination result from the language pair according to the original language score may include:

[0139] In the case where the language recognition accuracy of the large model for the specified language audio data is higher than a preset threshold, determining the language determination result according to the score ranking of the languages constituting the language pair in the original language score; wherein, the specified language audio data is the audio data of any language in the language pair.

[0140] In one or more embodiments provided by the present application, the process of the voice processing unit 11 determining the language determination result according to the score ranking of the languages constituting the language pair in the original language score may include:

[0141] When both languages in the language pair are supported languages of the large model, determine the score rankings of the first language and the second language that make up the language pair respectively according to the original language scores; the recognition accuracy of the large model for the audio data of the first language is higher than the preset threshold and higher than the recognition accuracy of the large model for the audio data of the second language;

[0142] When the duration of the audio to be translated input to the large model is not less than the first preset duration, if the score ranking of the first language is not greater than R1 and the score ranking of the first language is less than the score ranking of the second language, then determine the first language as the language determination result; otherwise, determine the second language as the language determination result;

[0143] When the duration of the audio to be translated input to the large model is less than the first preset duration, if the score ranking of the first language is not greater than R2 and the score ranking of the first language is less than the score ranking of the second language, then determine the first language as the language determination result; otherwise, determine the second language as the language determination result; where R1 < R2.

[0144] In one or more embodiments provided in the present application, the process by which the speech processing unit 11 determines the language determination result according to the score rankings of the languages that make up the language pair in the original language scores may include:

[0145] When one of the languages in the language pair is a supported language of the large model, determine the score ranking of the first language that makes up the language pair according to the original language scores; the recognition accuracy of the large model for the audio data of the first language is higher than the preset threshold; the first language and the second language form the language pair;

[0146] When the duration of the audio to be translated input to the large model is not less than the first preset duration, if the score ranking of the first language is not greater than R1, then determine the first language as the language determination result; otherwise, determine the second language as the language determination result;

[0147] When the duration of the audio to be translated input to the large model is less than the first preset duration, if the score ranking of the first language is not greater than R2, then determine the first language as the language determination result; otherwise, determine the second language as the language determination result; where R1 < R2.

[0148] In one or more embodiments provided in the present application, the process by which the speech processing unit 11 calls a large model to perform language recognition processing on the audio to be translated to obtain the original language scores of the audio to be translated may include:

[0149] Extract an audio segment of at least a second preset duration from the audio to be translated, and input the extracted audio segment into the large model for the large model to perform language identification, and use the original language score of the audio segment output by the large model as the original language score of the audio to be translated.

[0150] In one or more embodiments provided by the present application, the process of the speech processing unit 11 performing speech dictation processing on the audio to be translated may include:

[0151] Dictate the audio to be translated according to the two languages that make up the language pair respectively.

[0152] On the basis of the above, the speech processing unit 11 may also be used for:

[0153] After outputting the first translation result as the translation result of the audio to be translated, in response to the user's switching request, perform a second translation process on the second dictation result to obtain a second translation result; wherein, the second dictation result is the dictation result different from the first dictation result among the speech dictation results of the two languages obtained through the speech dictation processing, and the target language of the second translation process is the language in the language pair that matches the language determination result;

[0154] The result output unit 12 may also be used to output the second translation result as the translation result of the audio to be translated.

[0155] Next, another speech translation solution provided by the present application will be described. The speech translation solution described below may be referred to in correspondence with the speech translation solution above.

[0156] Figure 8 is a schematic flowchart of another speech translation method shown according to an embodiment of the present application. This method may be applied to a translation device. Combining Figure 8 as shown, this method may include the following steps:

[0157] Step S501, in response to the user's bilingual translation requirement, display a bilingual translation interface corresponding to a specified language pair.

[0158] The above-mentioned user's bilingual translation requirement may refer to that the user selects the bilingual translation function and specifies the language range of the bilingual translation. Among them, the specified language pair may include two languages specified by the user. For the way for the user to specify the language, the language specification method in the bilingual translation scenario of the existing translation machine may be referred to.

[0159] Step S502, in response to the user's trigger operation on the bilingual translation button, start recording audio.

[0160] The above bilingual translation button is associated with the bilingual translation function. Exemplarily, it can be a virtual control on the bilingual translation interface or a physical button set on the translation device, which is not limited in this application.

[0161] Optionally, while starting to record audio, a prompt message indicating that recording is in progress can also be displayed on the bilingual translation interface.

[0162] Step S503: After the user stops recording, stop the recording and display the first dictation result of the recorded audio to be translated on the bilingual translation interface.

[0163] The above user's stop of recording can be that the user triggers the above bilingual translation button or an end button for controlling the end of recording, or it can be that the recording automatically stops after the recording duration reaches the specified maximum recording duration. Except for the way of starting recording, the specific recording process can refer to the recording scheme of existing translation machines, which will not be elaborated in this application.

[0164] Wherein, the first dictation result is the dictation result corresponding to the language determination result obtained by performing speech dictation processing on the audio to be translated, and the language determination result is obtained by performing language determination processing on the audio to be translated.

[0165] Step S504: Output the first translation result.

[0166] Wherein, the first translation result is obtained by performing first translation processing on the first dictation result, and the target language of the first translation processing is the language in the specified language pair that does not match the language determination result.

[0167] Based on the above content, in this embodiment, the bilingual translation task is realized by a single button. Without the user manually selecting the language of the current speaker or selecting the recording button corresponding to the language of the current speaker, but by automatically identifying and determining the language of the input audio data, the intelligence level of the translation device is increased, the usage fluency of the translation device is improved, and the situation of the failure of the current round of translation caused by the mismatch between the language corresponding to the recording button clicked by the user and the language of the current speaker can be avoided, thereby improving the usability of the translation device and helping to enhance the user experience.

[0168] In one or more embodiments provided by this application, the above step S504: Output the first translation result may include:

[0169] Display the first translation result on the bilingual translation interface and / or play the first audio data.

[0170] Wherein, the first audio data is obtained by converting the first translation result.

[0171] In one or more embodiments provided by the present application, after outputting the first translation result, it may further include:

[0172] Step S505: In response to the user's triggering operation on the switching button, display the second dictation result on the bilingual translation interface and output the second translation result.

[0173] Among them, the speech dictation process includes: dictating the audio to be translated according to the two languages that make up the specified language pair; the second dictation result is the dictation result different from the first dictation result among the speech dictation results of the two languages obtained through the speech dictation process, and the second translation result is obtained by performing a second translation process on the second dictation result, and the target language of the second translation process is the language in the specified language pair that matches the language determination result.

[0174] In addition, the above-mentioned switching button may be a virtual control on the bilingual translation interface or a physical button provided on the translation device, and the present application does not make a limitation.

[0175] On the basis of the above, the output of the second translation result may include:

[0176] Display the second translation result on the bilingual translation interface and / or play the second audio data. Among them, the second audio data is obtained by converting the second translation result.

[0177] Optionally, other descriptions of the speech translation method provided in this embodiment may be referred to the above description, and will not be elaborated here.

[0178] Next, a speech translation device provided by an embodiment of the present application will be described, and the speech translation device described below may be correspondingly referred to the speech translation method described above.

[0179] Figure 9 It is a schematic structural diagram of a speech translation device disclosed in an embodiment of the present application. As Figure 9 shown, the device may include:

[0180] A first display module 21, configured to display a bilingual translation interface corresponding to a specified language pair in response to the user's bilingual translation requirement; the specified language pair includes two languages specified by the user;

[0181] A bilingual translation module 22, configured to start recording audio in response to the user's triggering operation on the bilingual translation button; and stop recording after the user stops recording;

[0182] The second display module 23 is configured to display, after the recording is stopped, a first dictation result of the recorded audio to be translated on the bilingual translation interface; wherein, the first dictation result is a dictation result corresponding to the language determination result obtained by performing speech dictation processing on the audio to be translated, and the language determination result is obtained by performing language determination processing on the audio to be translated;

[0183] The first result output module 24 is configured to output a first translation result, which is obtained by performing a first translation process on the first dictation result, and the target language of the first translation process is the language in the specified language pair that does not match the language determination result.

[0184] In one or more embodiments provided by the present application, the first result output module 24 may include a third display module and / or an audio broadcast module;

[0185] The third display module is configured to display the first translation result on the bilingual translation interface;

[0186] The audio broadcast module is configured to play first audio data, wherein the first audio data is obtained by converting the first translation result.

[0187] In one or more embodiments provided by the present application, the device may further include a fourth display module, which is configured to display a second dictation result on the bilingual translation interface in response to a triggering operation of the user on the switching button; wherein, the speech dictation processing includes: dictating the audio to be translated according to the two languages that make up the specified language pair; the second dictation result is the dictation result different from the first dictation result among the speech dictation results of the two languages obtained by the speech dictation processing.

[0188] On the basis of the above, the device may further include a second result output module, which is configured to output a second translation result in response to a triggering operation of the user on the switching button; the second translation result is obtained by performing a second translation process on the second dictation result, and the target language of the second translation process is the language in the specified language pair that matches the language determination result.

[0189] In one or more embodiments provided by the present application, the device may further include: a translation service module 25;

[0190] The translation service module 25 is configured to perform speech dictation processing on the audio to be translated and perform language determination processing on the audio to be translated;

[0191] The translation service module 25 is further configured to perform a first translation process on the first dictation result.

[0192] Optionally, for a detailed description and an extended description of the device, reference may be made to the above description, which will not be elaborated here.

[0193] An embodiment of the present application also provides a translation device. Refer to Figure 10 As shown, it shows a schematic structural diagram of a translation device suitable for implementing the translation device in the embodiments of the present application. The translation device in the embodiments of the present application may include, but is not limited to, such as a mobile phone, a tablet computer, a translator, etc. Figure 10 The translation device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0194] As Figure 10 shown, the electronic device may include a processing device (such as a central processing unit, etc.) 1, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 2 or the program loaded from the storage device 8 into the random access memory (RAM) 3, so as to implement the voice translation method in the foregoing embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 3. The processing device 1, the ROM 2, and the RAM 3 are connected to each other through a bus 4. The input / output (I / O) interface 5 is also connected to the bus 4.

[0195] Generally, the following devices may be connected to the I / O interface 5: input devices 6 including, for example, a touch screen, a microphone, etc.; output devices 7 including, for example, a liquid crystal display (LCD), a speaker, etc.; storage devices 8 including, for example, a memory card, a hard disk, etc.; and a communication device 9. The communication device 9 may allow the translation device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 10 a translation device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and alternatively, more or fewer devices may be implemented or had.

[0196] An embodiment of the present application also provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any one of the voice translation methods provided in the embodiments of the present application.

[0197] An embodiment of the present application also provides a computer-readable storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can be enabled to implement any one of the voice translation methods provided in the embodiments of the present application.

[0198] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0199] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or dedicated circuits. However, for this application, software program implementation is a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0200] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0201] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, a computer, a training device, or a data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0202] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0203] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0204] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech translation method, characterized in that: include: In response to bilingual translation needs, language determination and voice dictation processing are performed on the audio to be translated; The result of the speech dictation processing includes a first dictation result corresponding to the language determination result of the audio to be translated; Performing a first translation process on the first dictation result to obtain a first translation result; The target language of the first translation processing is a language in the language pair specified by the bilingual translation requirement that does not match the language determination result; The first translation result is output as the translation result of the audio to be translated.

2. The speech translation method according to claim 1, characterized in that: The language determination process of the audio to be translated includes: Calling the large model to perform language recognition processing on the audio to be translated to obtain an original language score of the audio to be translated; the original language score includes a score of at least one supported language, and the score of any of the supported languages ​​is positively correlated with the possibility that the audio to be translated belongs to the supported language; The language determination result is determined from the language pair according to the original language score.

3. The speech translation method according to claim 2, characterized in that: Determining the language determination result from the language pair according to the original language score includes: When the language recognition accuracy of the large model for the audio data of the specified language is higher than a preset threshold, the language determination result is determined according to the ranking of the scores of the languages ​​constituting the language pair in the original language scores; The specified language audio data is audio data of any language in the language pair.

4. The speech translation method according to claim 3, characterized in that: Determining the language determination result according to the ranking of the scores of the languages ​​constituting the language pair in the original language scores includes: In the case where both of the language pairs are supported languages ​​of the large model, the score rankings of the first language and the second language constituting the language pair are determined according to the original language scores; the recognition accuracy of the large model for the audio data of the first language is higher than the preset threshold, and higher than the recognition accuracy of the large model for the audio data of the second language; When the duration of the audio to be translated input into the large model is not less than a first preset duration, if the score ranking of the first language is not greater than R1, and the score ranking of the first language is less than the score ranking of the second language, the first language is determined as the language determination result, otherwise the second language is determined as the language determination result, where R1 is a set positive integer.

5. The speech translation method according to claim 4, characterized in that: Determining the language determination result according to the ranking of the scores of the languages ​​constituting the language pair in the original language scores also includes: When the duration of the audio to be translated input into the large model is less than the first preset duration, if the score ranking of the first language is not greater than R2, and the score ranking of the first language is less than the score ranking of the second language, the first language is determined as the language determination result, otherwise the second language is determined as the language determination result; wherein R1<R2.

6. The speech translation method according to claim 3, characterized in that: Determining the language determination result according to the ranking of the scores of the languages ​​constituting the language pair in the original language scores includes: In the case where one of the languages ​​in the language pair is a supported language of the large model, the score ranking of the first language is determined according to the original language score; the recognition accuracy of the large model for the audio data of the first language is higher than the preset threshold; the first language and the second language constitute the language pair; When the duration of the audio to be translated input into the large model is not less than a first preset duration, if the score ranking of the first language is not greater than R1, the first language is determined as the language determination result, otherwise the second language is determined as the language determination result, where R1 is a set positive integer.

7. The speech translation method according to claim 6, characterized in that: Determining the language determination result according to the ranking of the scores of the languages ​​constituting the language pair in the original language scores also includes: When the duration of the audio to be translated input into the large model is less than the first preset duration, if the score ranking of the first language is not greater than R2, the first language is determined as the language determination result, otherwise the second language is determined as the language determination result; wherein R1<R2.

8. The speech translation method according to any one of claims 2 to 7, characterized in that: The calling of the large model to perform language recognition processing on the audio to be translated to obtain the original language score of the audio to be translated includes: An audio segment of at least a second preset time length is extracted from the audio to be translated, and the extracted audio segment is input into the large model for the large model to perform language recognition, and the original language score of the audio segment output by the large model is used as the original language score of the audio to be translated.

9. The speech translation method according to any one of claims 1 to 7, characterized in that: The voice dictation processing of the audio to be translated includes: transcribe the audio to be translated respectively in two languages ​​constituting the language pair; After outputting the first translation result as the translation result of the audio to be translated, the method further includes: In response to a user's switching request, performing a second translation process on the second dictation result to obtain a second translation result; wherein the second dictation result is a dictation result different from the first dictation result in the speech dictation results of the two languages ​​obtained by the speech dictation process, and the target language of the second translation process is the language in the language pair that matches the language determination result; The second translation result is output as the translation result of the audio to be translated.

10. A speech translation method, characterized in that: include: In response to the user's bilingual translation requirements, display a bilingual translation interface corresponding to the specified language pair; The specified language pair includes two languages ​​specified by the user; In response to a user triggering an operation on a bilingual translation button, starting to record audio; After the user stops recording, stop recording and display a first dictation result of the recorded audio to be translated on the bilingual translation interface, where the first dictation result is a dictation result corresponding to a language determination result obtained by performing voice dictation processing on the audio to be translated, and the language determination result is obtained by performing language determination processing on the audio to be translated; A first translation result is outputted, where the first translation result is obtained by performing a first translation process on the first dictation result, where a target language of the first translation process is a language in the designated language pair that does not match the language determination result.

11. The speech translation method according to claim 10, characterized in that: The outputting of the first translation result comprises: The first translation result is displayed on the bilingual translation interface, and / or first audio data is played, where the first audio data is converted from the first translation result.

12. The speech translation method according to claim 10 or 11, characterized in that: After outputting the first translation result, the method further includes: In response to a user triggering operation on the switch button, displaying the second dictation result on the bilingual translation interface and outputting the second translation result; Among them, the speech dictation processing includes: dictating the audio to be translated respectively according to the two languages ​​constituting the specified language pair; the second dictation result is a dictation result different from the first dictation result among the speech dictation results in the two languages ​​obtained by the speech dictation processing, the second translation result is obtained by performing a second translation processing on the second dictation result, and the target language of the second translation processing is the language in the specified language pair that matches the language determination result.

13. A speech translation device, characterized in that: include: A speech processing unit, used to respond to bilingual translation needs and perform language determination and speech dictation processing on the audio to be translated; The result of the speech dictation processing includes a first dictation result corresponding to the language determination result of the audio to be translated; and performing a first translation process on the first dictation result to obtain a first translation result; The target language of the first translation processing is a language in the language pair specified by the bilingual translation requirement that does not match the language determination result; The result output unit is used to output the first translation result as the translation result of the audio to be translated.

14. A speech translation device, characterized in that: include: A first display module, configured to display a bilingual translation interface corresponding to a specified language pair in response to a user's bilingual translation requirement; The specified language pair includes two languages ​​specified by the user; A bilingual translation module, used to start recording audio in response to a user triggering an operation on a bilingual translation button; And stop recording after the user stops recording; A second display module is used to display a first dictation result of the recorded audio to be translated on the bilingual translation interface after the recording is stopped; wherein the first dictation result is a dictation result corresponding to a language determination result obtained by performing voice dictation processing on the audio to be translated, and the language determination result is obtained by performing language determination processing on the audio to be translated; The result output module is used to output a first translation result, where the first translation result is obtained by performing a first translation process on the first dictation result, where a target language of the first translation process is a language in the specified language pair that does not match the language determination result.

15. The speech translation device according to claim 14, characterized in that: The device also includes: a translation service module; The translation service module is used to perform voice dictation processing on the audio to be translated and to perform language determination processing on the audio to be translated; The translation service module is further used to perform a first translation process on the first dictation result.

16. A translation device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the translation device can implement the speech translation method according to any one of claims 1 to 9, or implement the speech translation method according to any one of claims 10 to 12.

17. A storage medium, characterized in that: The storage medium carries one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the speech translation method as described in any one of claims 1 to 9, or the speech translation method as described in any one of claims 10 to 12.

18. A computer program product, characterized in that The invention comprises computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the speech translation method as claimed in any one of claims 1 to 9, or implement the speech translation method as claimed in any one of claims 10 to 12.