Method and device for translating voice data, equipment and medium

By receiving and processing source voice data of multiple tones, and using machine learning models to generate voice translations with different tones, the problem of difficulty in distinguishing multiple user tones in the prior art is solved, and efficient simultaneous translation is achieved.

CN120048249APending Publication Date: 2025-05-27BYTEDANCE TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510202592.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing natural language translation model is difficult to distinguish the different tones of multiple users in the simultaneous translation scenario, resulting in the inability to accurately distinguish the speaker.

Method used

By receiving source speech data represented in the source language, the machine learning model is used to determine the corresponding speech translation to ensure that the translation results are consistent with the input data in the tone, and support multi-tone speech translation.

Benefits of technology

It realizes the accurate distinction between different tones in simultaneous translation, and generates speech translation results with different tones, improving the efficiency of language communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048249A_ABST
    Figure CN120048249A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device, equipment and a medium for translating voice data. In one method, source speech data represented in a source language is received, the source speech data including a first speech segment having a first timbre and a second speech segment having a second timbre. A speech translation corresponding to the source speech data is determined using a machine learning model, the speech translation being represented in a target language, and the speech translation comprising a first speech translation corresponding to the first speech segment and a second speech translation corresponding to the second speech segment. The first speech translation has a first timbre and the second speech translation has a second timbre. By utilizing the implementation mode of the invention, different voice segments with different timbres expressed in a source language can be accurately distinguished, and voice translations with different timbres expressed in a target language can be respectively generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Implementations of the present disclosure generally relate to natural language translation, and particularly to methods, apparatuses, devices, and computer-readable storage media for translating speech data from a source language to a destination language. Background Art

[0002] Machine learning techniques have been widely used in natural language translation, and machine learning models dedicated to simultaneous translation have been developed. During simultaneous translation, input data is received in a streaming manner, and the machine learning model gradually outputs the translation result corresponding to the newly received part. The simultaneous translation scenario may involve conversations among multiple users. However, existing translation models usually only support translation of a single voice tone and cannot distinguish the different voice tones of multiple users. At this time, it is desirable to provide voice translation that supports multiple voice tones while ensuring low time delay and high accuracy of simultaneous translation. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for translating speech data is provided. In this method, source speech data represented in a source language is received, and the source speech data includes a first speech segment having a first voice tone and a second speech segment having a second voice tone. Using a machine learning model, a speech translation corresponding to the source speech data is determined. The speech translation is represented in a destination language, and the speech translation includes a first speech translation corresponding to the first speech segment and a second speech translation corresponding to the second speech segment. The first speech translation has the first voice tone, and the second speech translation has the second voice tone.

[0004] In a second aspect of the present disclosure, an apparatus for translating speech data is provided. The apparatus includes: a receiving module configured to receive source speech data represented in a source language, where the source speech data includes a first speech segment having a first voice tone and a second speech segment having a second voice tone; and a determining module configured to use a machine learning model to determine a speech translation corresponding to the source speech data. The speech translation is represented in a destination language, and the speech translation includes a first speech translation corresponding to the first speech segment and a second speech translation corresponding to the second speech segment. The first speech translation has the first voice tone, and the second speech translation has the second voice tone.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to execute the method according to the first aspect of the present disclosure.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement the method according to the first aspect of the present disclosure.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the present disclosure.

[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the implementation manners of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the following, in conjunction with the drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various implementation manners of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0010] Figure 1 A block diagram of an application environment according to an implementation manner of the present disclosure is shown;

[0011] Figure 2 A block diagram of a system for translating speech data according to some implementation manners of the present disclosure is shown;

[0012] Figure 3 A block diagram of determining speech translation based on previous data according to some implementation manners of the present disclosure is shown;

[0013] Figure 4 A block diagram of a prompt for invoking a language model according to some implementation manners of the present disclosure is shown;

[0014] Figure 5 A block diagram of a system for translating speech data using a machine learning model according to some implementation manners of the present disclosure is shown;

[0015] Figure 6 A block diagram of a reference sample for training a machine learning model according to some implementation manners of the present disclosure is shown;

[0016] Figure 7A flowchart of a method for translating speech data according to some implementations of the present disclosure is shown;

[0017] Figure 8 A block diagram of an apparatus for translating speech data according to some implementations of the present disclosure is shown; and

[0018] Figure 9 A block diagram of a device capable of implementing multiple implementations of the present disclosure is shown. Detailed Implementations

[0019] Implementations of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some implementations of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. On the contrary, these implementations are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and implementations of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0020] In the description of the implementations of the present disclosure, the term "including" and its like should be understood as an open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". There may also be other explicit and implicit definitions below. As used herein, the term "model" may represent the association relationship between various data. For example, the above association relationship can be obtained based on various technical solutions known currently and / or to be developed in the future.

[0021] It can be understood that the data involved in the technical solution of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.

[0023] For example, in response to receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving an active request from a user, the way of sending a prompt message to the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It can be understood that the above notification and user authorization acquisition process are only illustrative and do not limit the implementation of the present disclosure. Other ways that comply with relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0026] The term "in response to" used herein represents a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the execution timing of subsequent actions performed in response to the event or condition and the time when the event occurs or the condition is established are not necessarily strongly correlated. For example, in some cases, the subsequent action can be immediately executed when the event occurs or the condition is established; while in other cases, the subsequent action can be executed after a period of time after the event occurs or the condition is established.

[0027] Example environment

[0028] Currently, machine learning models dedicated to simultaneous interpretation have been developed. During simultaneous interpretation, input data is received in a streaming manner, and the machine learning model gradually outputs the translation result corresponding to the newly received part. See Figure 1 Describe the environment of simultaneous interpretation, the Figure 1 shows a block diagram 100 of an application environment according to an implementation of the present disclosure. As Figure 1 shown, the simultaneous interpretation scenario may involve conversations among multiple users. User 110 and User 120 can provide speech data expressed in the source language and expect to translate the speech data into a speech translation result expressed in the target language. For the sake of description, in the context of the present disclosure, only Chinese is used as an example of the source language and English is used as an example of the target language to describe the translation process. Alternatively and / or additionally, the source language and the target language can include other natural languages.

[0029] As Figure 1As shown, between time points T0 and T1, user 120 can provide voice data 122; between time points T1 and T2, user 110 can provide voice data 112; and between time points T2 and T3, user 120 can provide voice data 124. At this time, the collected voice data includes interleaved voice segments from multiple people. A translation model can be used to translate the voice data represented in Chinese into a voice translation result represented in English. However, existing translation models usually only support translation from voice to text, or only support translation of a single voice tone (e.g., the voice tone synthesized by a machine), and cannot distinguish the different voice tones of multiple users. At this time, it is desirable to provide voice translation that supports multiple voice tones while ensuring low time delay and high accuracy for simultaneous interpretation.

[0030] Summary of speech translation

[0031] To at least partially address the deficiencies in the prior art, according to one implementation of the present disclosure, a method for translating voice data is proposed. Refer to Figure 2 Describe the overview according to one implementation of the present disclosure, Figure 2 FIG. 200 shows a block diagram for translating voice data according to some implementations of the present disclosure. As Figure 2 shown, a method for translating voice data is provided. Specifically, source voice data 210 represented in a source language can be received. The source voice data 210 includes a first voice segment 212 having a first voice tone and a second voice segment 214 having a second voice tone.

[0032] Subsequently, a machine learning model 220 can be used to determine a voice translation 230 corresponding to the source voice data 210, and the voice translation 230 is represented in a target language. The voice translation 230 is represented in an audio format and includes a first voice translation 232 corresponding to the first voice segment 212 and a second voice translation 234 corresponding to the second voice segment 214. Here, the first voice translation 232 has the first voice tone, and the second voice translation 234 has the second voice tone.

[0033] For ease of description, the voice translation process is described only by taking a meeting as an example. The meeting may involve a first user Alice and a second user Bob. The two users talk in Chinese and expect to translate the Chinese conversation content into English. Specifically, the first voice segment 212 comes from the first user Alice with a sweet voice tone, and the second voice segment comes from the second user Bob with a deep voice tone. At this time, the generated first voice translation can have a sweet voice tone, and the second voice translation can have a deep voice tone.

[0034] Using the implementations of the present disclosure, different speech segments with different timbres represented in the source language can be accurately distinguished, and speech translations with different timbres represented in the target language can be generated respectively. In this way, listeners can distinguish the speakers corresponding to different speech translations in a more explicit manner, thereby improving the efficiency of language communication.

[0035] Detailed process of speech translation

[0036] An overview of some implementations according to the present disclosure has been described. Hereinafter, more information about speech translation will be provided. According to some implementations of the present disclosure, historical data before the source speech data can be used as context data for a machine learning model. Specifically, the source speech data is a part of the source speech data sequence. At this time, historical data can be extracted from the source speech sequence. Specifically, in the process of determining the speech translation corresponding to the source speech data, the previous text data corresponding to the previous speech data in the source speech data sequence can be determined, and the previous speech data is located before the source speech data; the previous text translation corresponding to the previous text data can be obtained; and using the machine learning model, based on the previous text data and the previous text translation, the speech translation corresponding to the source speech data can be determined. Using some implementations of the present disclosure, simultaneous translation can be provided in the context of a multi-person conversation, thereby improving the accuracy of machine translation.

[0037] See Figure 3 for more details of the translation process, which Figure 3 shows a block diagram 300 for determining a speech translation based on previous data according to some implementations of the present disclosure. As Figure 3 shown, the source speech data sequence 340 can have a relatively long time length. For example, the speech data collected in real time during a meeting. The source speech data 210 to be translated and the previous speech data 310 located before the source speech data 210 can be extracted from the source speech data sequence 340. The previous speech data 310 can be used as context data for the translation process and input into the machine learning model 220, and then the speech translation 230 for the source speech data 210 can be determined. In this way, a more accurate speech translation 230 can be obtained under the constraint of the previous speech data 310.

[0038] According to some implementations of the present disclosure, in the process of determining a speech translation corresponding to source speech data based on previous text data and previous text translations, a machine learning model can be utilized to respectively determine a first text translation corresponding to a first speech segment and a second text translation corresponding to a second speech segment based on the previous text data and the previous text translations. Further, the machine learning model can be utilized to respectively determine a first speech translation corresponding to the first text translation and a second speech translation corresponding to the second text translation. Subsequently, the machine learning model is utilized to combine the first speech translation and the second speech translation to determine the speech translation. In the context of the present disclosure, the text translation is represented in text format, and the speech translation is represented in audio format.

[0039] It should be understood that the machine learning model herein can be an end-to-end machine learning model with multiple capabilities, and the model can include multiple network modules to respectively perform corresponding functions. For example, the machine learning model can include: a translation module configured to respectively determine a first text translation corresponding to the first speech segment and a second text translation corresponding to the second speech segment based on the previous text data and the previous text translations; and a speech generation module configured to respectively determine a first speech translation corresponding to the first text translation and a second speech translation corresponding to the second text translation, and combine the first speech translation and the second speech translation to determine the speech translation.

[0040] According to some implementations of the present disclosure, the translation module herein can be implemented based on a language model. Specifically, in the process of respectively determining the first text translation and the second text translation based on the previous text data and the previous text translations, a machine learning model can be utilized to determine the text translation corresponding to the source speech data. Herein, the text translation can include a switching marker represented in text format, and the switching marker corresponds to the switching point between the first speech segment and the second speech segment. Continuing to refer to Figure 3 , the previous text data 320 corresponding to the previous speech data 310 can be determined. Herein, the previous text data 320 can be determined based on an automatic speech recognition (abbreviated as ASR) model, and the previous text translation 330 can be determined in a previous translation step.

[0041] According to some implementations of the present disclosure, the translation method of the present disclosure can be performed iteratively. For example, in step 1, the previous speech data 310, the previous text data 320, and the previous text translation 330 are all empty. At this time, although the context data is empty, the machine learning model 220 can still provide a text translation corresponding to the source speech data. In step 2, the source speech data in step 1 can be used as the previous speech data 310, the previous text data 320 can be determined by the ASR model, and the previous text translation 330 can be determined using the translation module in the machine learning model. At this time, the context data is non-empty, and thus the translation module can determine a text translation and a speech translation corresponding to the source speech data in the current step in the context specified by the context data. In subsequent steps, the operation can be performed in a similar manner until there is no more untranslated content in the source speech data sequence 340.

[0042] According to some implementations of the present disclosure, the source speech data can be received in a streaming manner. At this time, the source speech data sequence can be a real-time received audio stream. In this way, newly received content can be processed in real time, so as to provide an accurate translation result. Specifically, in the scenario of simultaneous interpretation, the speaker can keep speaking, and data segments can be determined in a predetermined manner. In multiple steps, newly received data segments can be continuously detected and processed using the process described above. In this way, a translation service can be provided in real time in an accurate and efficient manner.

[0043] According to some implementations of the present disclosure, over time, the source speech data can span a relatively long period. At this time, the source speech data sequence can be divided into speech data blocks with a predetermined length to ensure that the amount of data to be translated processed in one translation step is not too high. For example, the predetermined time length can be set to 20 seconds, 30 seconds (or other time lengths). Specifically, during the process of receiving the source speech data, speech data blocks can be determined from the source speech data sequence according to the predetermined time length, and further the source speech data can be determined from the speech data blocks. Here, the newly received real-time speech data can be used as the source speech data. Alternatively and / or additionally, the maximum length of the source speech data can be set. In this way, the translation process can be performed in real time, thereby achieving the purpose of simultaneous interpretation.

[0044] According to some implementations of the present disclosure, the previous speech data and the source speech data may be located within the same speech data block. In other words, within the speech data block, the portion before the source speech data can be used as the previous speech data. By using some implementations of the present disclosure, the context data can be limited within a finite range. In this way, on the one hand, the most recently obtained historical data can be utilized as much as possible as the context data, thereby avoiding the interference of early historical data on the translation process. On the other hand, it can be avoided that the context data is too long, resulting in an excessive workload of the machine learning model and a long delay. In this way, the translation performance of the machine learning model can be further improved.

[0045] According to some implementations of the present disclosure, the text translation can be divided into a first text translation and a second text translation based on a switching marker in the text translation. It should be understood that here the switching marker can correspond to the switching point between the first reference speech segment and the second reference speech segment, that is, the switching point between the first tone color and the second tone color. Assuming that the first reference speech segment in the source speech data is between time points T0 and T1, and the second speech segment is between time points T1 and T2, then the switching marker can correspond to T1. The switching marker can be represented in text format. For example, using the format <spk-chg> <time-xx>to indicate a switching flag, herein <spk-chg>Keywords indicating switching markers, and <time-xx>Indicates the position of the switching point in the source speech data. For example, <spk-chg> <time-t1>It may be represented that the switching point between the first reference speech segment and the second reference speech segment is located at the time point T1 in the source speech data. It should be understood that the above switching markers are merely exemplary. Alternatively and / or additionally, other formats may be used to represent the switching markers. For example, different keywords and time formats may be set. Another example of the switching marker may be represented as: <switch_point> <time:xx>。

[0046] After the first text translation and the second text translation have been obtained from the translation module, the speech generation module in the machine learning model can be utilized to respectively determine a first speech translation corresponding to the first text translation and a second speech translation corresponding to the second text translation. Subsequently, the speech generation module can combine the first speech translation and the second speech translation to determine the speech translation. Here, the speech module is implemented based on a text-to-speech machine learning model, and the timbre of the speech translation to be generated can be specified. For example, it can be specified that the first speech translation has a first timbre and the second speech translation has a second timbre.

[0047] According to some implementations of the present disclosure, the translation process can be performed based on a language model. For example, a prompt can be input to the machine learning model, and a response from the machine learning model to the prompt can be received. Refer to Figure 4 for more details regarding the prompt, the Figure 4 shows a block diagram 400 of a prompt for invoking a language model according to some implementations of the present disclosure. As Figure 4 shown, the prompt 410 can include multiple parts. For example, the field 420 can be in text format and is used to specify the previous text data 320 of the stored previous speech data 310. The field 422 can be in text format and is used to specify the previous text translation 330 of the previous text data 320. The field 424 can be in audio format and is used to specify the source speech data 210 to be translated. The instruction 426 can specify the task performed by the machine learning model, such as extracting the transcribed text from the speech and translating it into English.

[0048] Alternatively and / or additionally, the instruction 426 can further specify the format of the output translation: [Current ASR] <seperator>[Current ST]<end_time_token>. Here, "[Current ASR]" represents the text to be translated extracted from the source speech data 210 to be translated, " <seperator>" indicates a separator (e.g., colon ":", equal sign "=", or other separator), "[Current ST]" indicates the English translation text corresponding to the text to be translated, and "<end_time_token>" indicates the end time of the source speech data 210. It should be understood that although the above prompts are written in the English language, alternatively and / or additionally, based on the processing capabilities of the machine learning model, the prompts can be written in other languages. For example, the instruction 426 in the prompt can include: "Extract the transcribed text from the speech and translate it into English."

[0049] In the following, refer to Figure 5 Describe more details in the overall translation process, which Figure 5 shows a block diagram 500 for translating speech data using a machine learning model according to some implementations of the present disclosure. As Figure 5 shown, the prompt 410 can be input to the machine learning model 220. The previous text data 320 can be filled in the field 420 in the prompt 410, the previous text translation 330 can be filled in the field 422, and the source speech data 210 can be filled in the field 424. Here, the source speech data 210 is audio data represented in Chinese, and the corresponding Chinese text is "Well, what basic functions does it have."

[0050] The speech-to-text network 510 (corresponding to the translation module) in the machine learning model 220 can receive the prompt and output a text translation. The text translation can be represented as a text string, for example, "what basic function it has." <spk-chg><time-4.00>Alice, the function”. The text translation can include multiple parts: the first text translation 530, the switching marker 532, and the second text translation 534. At this time, the switching marker 532 can represent the switching point between the two text translations. It should be understood that the speech-to-text network 510 here is trained using a reference sample with a switching marker. Therefore, when a voice tone switch appears in the source voice data 210, the switching marker is output in the text translation. Table 1 shows the specific data of each field. The first voice segment ranges from the start time point 0.00 to the time point 4.00, and the second voice segment ranges from the time point 4.00 to the end time point 6.00.

[0051] Data in the fields of Table 1

[0052] Field name Data Memory ASR First, well, what I said before was to help you understand Memory ST First,what I said before was to help you understand Current ASR Well, what are its basic functions? <spk-chg><time-4.00>Alice, this function Current ST En,what basic function it has. <spk-chg><time-4.00>Alice, the function end_time_token <time-6.00>

[0053] Furthermore, the text-to-speech network 520-1 (corresponding to the voice generation module) in the machine learning model 220 can be used to generate a first voice translation 550 corresponding to the first text translation 530 and having the first voice segment 540 with the first voice tone. The text-to-speech network 520-2 can be used to generate a second voice translation 552 corresponding to the second text translation 534 and having the second voice segment 542 with the second voice tone. Subsequently, the first voice translation 550 and the second voice translation 552 can be combined to generate the voice translation 230.

[0054] It should be understood that although Figure 5 The translation process has been described by taking the source voice data 210 with voice switching as an example. Alternatively and / or additionally, the machine learning model of the present disclosure can process source voice data without voice switching. At this time, if the model finds that there is no voice tone switch in the source voice data, the switching marker is not output in the text translation. In this way, a unified machine learning model can be used to perform the translation process, thereby simplifying the management complexity of the translation process.

[0055] According to some implementations of the present disclosure, an end-to-end machine learning model can be used to perform the translation process. Compared with the existing technical solutions that perform voice translation in a cascaded manner, the data communication between the various network modules in the end-to-end manner is carried out in the feature space and does not need to be converted into an explicit format recognizable by humans. Therefore, the transmission delay can be reduced and the translation efficiency can be improved. Further, the loss function can be determined and the various network modules in the machine learning model can be trained in an overall manner. Thus, the accuracy of the machine learning model can be improved as a whole. In this way, the end-to-end machine learning model can have lower latency and higher accuracy.

[0056] It should be understood that the machine learning model 220 herein can be a pre-trained model. Reference samples can be collected and the machine learning model can be trained. According to some implementations of the present disclosure, any suitable translation model can be used as the base model for the translation module in the machine learning model 220. Herein, the translation module can translate source speech data represented in a source language into text translation represented in a target language. Different from a conventional speech-to-text translation model, the text translation herein can include multiple text translations respectively corresponding to multiple speech segments in the source speech data, and switching markers corresponding to switching points between the multiple speech segments.

[0057] In other words, the translation module can be trained to perform multiple functions: translating speech data represented in a source language into text data represented in a target language, and identifying switching points between different speech segments with different timbres in the speech data. By using some implementations of the present disclosure, the machine learning model can be made to provide more abundant functions, thereby providing accurate timbre switching points for the subsequent process of generating speech translation.

[0058] According to some implementations of the present disclosure, the translation model is determined based on the following: obtaining reference previous text data and reference previous text translation represented in a source language, the reference previous text data being represented in the source language and the reference previous text translation being represented in the target language; obtaining reference source speech data represented in the source language, the reference source speech data including a first reference speech segment having a first timbre and a second reference speech segment having a second timbre; obtaining a first reference text translation corresponding to the first reference speech segment and a second reference text translation corresponding to the second reference speech segment; and updating the translation model based on the reference previous text data, the reference previous text translation, the reference source audio data, the first reference text translation, and the second reference text translation.

[0059] In the context of the present disclosure, the reference previous speech data and the reference source speech data are represented in the source language and are successive audio data extracted from a known speech data stream. Herein, the reference sample can be represented as: "[Memory ASR][Memory ST][Current Audio]Transcribe the speech and then translate it to Chinese:[Current ASR] <seperator>[Current ST]<end_time_token>”.

[0060] See Figure 6 for more details regarding a reference sample that Figure 6 shows a block diagram 600 of a reference sample for training a machine learning model according to some implementations of the present disclosure. As Figure 6 shown, a publicly available translation dataset can be utilized to construct the reference sample 610. The process of training a machine learning model is described by way of example only for Chinese-to-English translation. The reference sample 610 can include multiple fields: field 611 corresponds to reference previous text data represented in the source language, field 612 corresponds to a reference previous text translation represented in the target language (i.e., the English translation of the reference previous text data), field 613 corresponds to reference source speech data in a speech format represented in the source language, field 614 corresponds to reference text data (i.e., the text recognized from the reference source speech data), field 615 corresponds to a reference text translation represented in the target language (i.e., the English translation of the reference text data), and field 618 corresponds to an end marker.

[0061] According to some implementations of the present disclosure, the first reference text translation includes a reference switch marker corresponding to a switch point between a first reference speech segment and a second reference speech segment. At this time, the end of field 614 includes field 616, corresponding to the switch point in the reference text data; the end of field 615 includes field 617, corresponding to the switch point in the reference text data. In this way, the translation module can accurately learn relevant knowledge about identifying speech switch points. Table 2 below shows examples of each field in the reference sample.

[0062] Table 2 Examples of the reference sample

[0063]

[0064]

[0065] In the above reference sample, the speech segment corresponding to "I will first summarize the status of three projects" has a first tone color and a time range of 0.00 to 3.00 seconds; and the speech segment corresponding to "Bob, sorry, there is a problem" has a second tone color and a time range of 3.00 to 5.00 seconds.

[0066] According to some implementations of the present disclosure, the translation module can be updated using the reference sample 610. At this time, fields 611, 612, and 613 correspond to the data part in the reference sample, and fields 614 and 615 correspond to the annotation part in the reference sample. Specifically, fields 611, 612, and 613 can be input into the translation module to be updated, and predictions of fields 614 and 615 from the translation module can be received. Further, based on the differences between fields 614 and 615 and the corresponding predictions, the parameters of the translation module can be updated. For example, the parameters of the translation module can be updated in a direction that minimizes the difference. According to some implementations of the present disclosure, the translation module can be updated iteratively and using a large number of reference samples to obtain a translation module that can both provide translations and identify voice tone switching points.

[0067] According to some implementations of the present disclosure, a text-to-speech (TTS) network can be used to implement the voice module. For example, a text-to-speech model with a voice tone cloning function can be used to construct the voice module. The voice tone of the original audio can be combined with the text in the target language to generate an audio corresponding to the text in the target language with the voice tone of the original audio. According to some implementations of the present disclosure, the text-to-speech network 520 can be trained separately. Alternatively and / or additionally, different reference samples can be used to train the text-to-speech network 520 together with the speech-to-text network 510. In this way, the machine learning model 220 can acquire knowledge about various aspects of performing voice translation, thereby improving the performance of the machine learning model 220.

[0068] It should be understood that although the translation process is only described above by taking Chinese as the source language and English as the target language as an example. Alternatively and / or additionally, the source language and the target language can include other natural languages, such as Japanese, French, German, Spanish, and so on. For the convenience of description, only the audio format is taken as an example of the voice data to be translated. Alternatively and / or additionally, the voice data can be represented, for example, in the audio format or the video format.

[0069] Using the implementations of the present disclosure, the machine learning model 220 can accurately distinguish different voice segments with different voice tones expressed in the source language and generate voice translations with different voice tones expressed in the target language respectively. Further, the end-to-end machine learning model 220 can more efficiently determine the voice translation. Thus, the listener can distinguish the speakers corresponding to different voice translations in a more definite manner, thereby improving the language communication efficiency.

[0070] Example process

[0071] Figure 7 FIG. 700 is a flowchart of a method for translating speech data according to some implementations of the present disclosure. At block 710, source speech data represented in a source language is received, the source speech data including a first speech segment having a first timbre and a second speech segment having a second timbre. At block 720, a speech translation corresponding to the source speech data is determined using a machine learning model, the speech translation being represented in a target language and including a first speech translation corresponding to the first speech segment and a second speech translation corresponding to the second speech segment, the first speech translation having the first timbre and the second speech translation having the second timbre.

[0072] According to some implementations of the present disclosure, the source speech data is part of a sequence of source speech data, and determining the speech translation corresponding to the source speech data includes: determining previous text data corresponding to previous speech data in the sequence of source speech data, the previous speech data being located before the source speech data; obtaining a previous text translation corresponding to the previous text data; and using a machine learning model, based on the previous text data and the previous text translation, determining the speech translation corresponding to the source speech data.

[0073] According to some implementations of the present disclosure, determining the speech translation corresponding to the source speech data based on the previous text data and the previous text translation includes: using a machine learning model, based on the previous text data and the previous text translation, respectively determining a first text translation corresponding to the first speech segment and a second text translation corresponding to the second speech segment; using a machine learning model, respectively determining a first speech translation corresponding to the first text translation and a second speech translation corresponding to the second text translation; and using a machine learning model, combining the first speech translation and the second speech translation to determine the speech translation.

[0074] According to some implementations of the present disclosure, respectively determining the first text translation and the second text translation based on the previous text data and the previous text translation includes: using a machine learning model, determining a text translation corresponding to the source speech data; and based on a switching marker in the text translation, dividing the text translation into the first text translation and the second text translation, the switching marker corresponding to a switching point between a first reference speech segment and a second reference speech segment.

[0075] According to some implementations of the present disclosure, receiving the source speech data includes: extracting a speech data block from the sequence of source speech data according to a predetermined time length; and determining the source speech data from the speech data block.

[0076] According to some implementations of the present disclosure, the previous speech data and the source speech data are within the speech data block.

[0077] According to some implementations of the present disclosure, a machine learning model includes: a translation model, and a translation module for translating source speech data represented in a source language into a text translation represented in a target language, the text translation including a plurality of text translations respectively corresponding to a plurality of speech segments in the source speech data, and a transition marker corresponding to a transition point between the plurality of speech segments.

[0078] According to some implementations of the present disclosure, the translation model is determined based on: obtaining reference previous text data and a reference previous text translation represented in the source language, the reference previous text data being represented in the source language and the reference previous text translation being represented in the target language; obtaining reference source speech data represented in the source language, the reference source speech data including a first reference speech segment having a first timbre and a second reference speech segment having a second timbre; obtaining a first reference text translation corresponding to the first reference speech segment and a second reference text translation corresponding to the second reference speech segment; and updating the translation model based on the reference previous text data, the reference previous text translation, the reference source audio data, the first reference text translation, and the second reference text translation.

[0079] According to some implementations of the present disclosure, the first reference text translation includes a reference transition marker corresponding to a transition point between the first reference speech segment and the second reference speech segment.

[0080] According to some implementations of the present disclosure, the source speech data is received in a streaming manner.

[0081] Example device and equipment

[0082] Figure 8 A block diagram of an apparatus 800 for translating speech data according to some implementations of the present disclosure is shown. The apparatus 800 includes: a receiving module 810 configured to receive source speech data represented in a source language, the source speech data including a first speech segment having a first timbre and a second speech segment having a second timbre; and a determining module 820 configured to use a machine learning model to determine a speech translation corresponding to the source speech data, the speech translation being represented in a target language and the speech translation including a first speech translation corresponding to the first speech segment and a second speech translation corresponding to the second speech segment, the first speech translation having the first timbre and the second speech translation having the second timbre.

[0083] According to some implementations of the present disclosure, the source speech data is part of a source speech data sequence, and the determining module is further configured to: determine previous text data corresponding to previous speech data in the source speech data sequence, where the previous speech data is located before the source speech data; obtain a previous text translation corresponding to the previous text data; and use a machine learning model to determine a speech translation corresponding to the source speech data based on the previous text data and the previous text translation.

[0084] According to some implementations of the present disclosure, the determining module 820 is further configured to: use a machine learning model to respectively determine a first text translation corresponding to a first speech segment and a second text translation corresponding to a second speech segment based on the previous text data and the previous text translation; use a machine learning model to respectively determine a first speech translation corresponding to the first text translation and a second speech translation corresponding to the second text translation; and use a machine learning model to combine the first speech translation and the second speech translation to determine a speech translation.

[0085] According to some implementations of the present disclosure, the determining module 820 is further configured to: use a machine learning model to determine a text translation corresponding to the source speech data; and divide the text translation into a first text translation and a second text translation based on a switching marker in the text translation, where the switching marker corresponds to a switching point between a first reference speech segment and a second reference speech segment.

[0086] According to some implementations of the present disclosure, the receiving module 810 is further configured to: extract speech data blocks from the source speech data sequence according to a predetermined time length; and determine the source speech data from the speech data blocks.

[0087] According to some implementations of the present disclosure, the previous speech data and the source speech data are located within the speech data block.

[0088] According to some implementations of the present disclosure, the machine learning model includes: a translation model, where the translation module is configured to translate source speech data represented in a source language into a text translation represented in a target language, and the text translation includes multiple text translations respectively corresponding to multiple speech segments in the source speech data and switching markers corresponding to switching points between the multiple speech segments.

[0089] According to some implementations of the present disclosure, a translation model is determined based on: obtaining reference prior text data and a reference prior text translation represented in a source language, the reference prior text data being represented in the source language and the reference prior text translation being represented in a target language; obtaining reference source speech data represented in the source language, the reference source speech data including a first reference speech segment having a first voice tone and a second reference speech segment having a second voice tone; obtaining a first reference text translation corresponding to the first reference speech segment and a second reference text translation corresponding to the second reference speech segment; and updating the translation model based on the reference prior text data, the reference prior text translation, the reference source audio data, the first reference text translation, and the second reference text translation.

[0090] According to some implementations of the present disclosure, the first reference text translation includes a reference switching marker corresponding to a switching point between the first reference speech segment and the second reference speech segment.

[0091] According to some implementations of the present disclosure, the source speech data is received in a streaming manner.

[0092] Figure 9 A block diagram of a device 900 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that Figure 9 the illustrated computing device 900 is merely exemplary and should not constitute any limitation on the functions and scope of the implementations described herein. Figure 9 The illustrated computing device 900 can be used to implement the methods described above.

[0093] As Figure 9 shown, the computing device 900 is in the form of a general-purpose computing device. The components of the computing device 900 may include, but are not limited to, one or more processors 910, a memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processor 910 may be an actual or virtual processor and is capable of performing various processes according to programs stored in the memory 920. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 900.

[0094] Computing device 900 generally includes multiple computer storage media. Such media can be any available media accessible to computing device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 920 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 930 can be removable or non-removable media and can include machine-readable media such as a flash drive, a magnetic disk, or any other media that can be capable of storing information and / or data (such as training data for training) and can be accessed within computing device 900.

[0095] Computing device 900 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 9 a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. Memory 920 can include a computer program product 925 having one or more program modules that are configured to execute various methods or actions of various implementations of the present disclosure.

[0096] Communication unit 940 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of computing device 900 can be implemented in a single computing cluster or multiple computing machines that are capable of communicating via a communication connection. Thus, computing device 900 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0097] Input device 950 can be one or more input devices such as a mouse, a keyboard, a trackball, etc. Output device 960 can be one or more output devices such as a display, a speaker, a printer, etc. Computing device 900 can also communicate with one or more external devices (not shown) as needed via communication unit 940, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with computing device 900, or communicate with any device that enables computing device 900 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0098] According to an implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above. According to an implementation of the present disclosure, there is provided a computer program product having a computer program stored thereon, and when the program is executed by a processor, the method described above is implemented.

[0099] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0100] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that when these instructions are executed by the processor of the computer or other programmable data processing apparatus, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing apparatus, and / or other devices to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured article, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0101] The computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operation steps are executed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0103] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field of this technology without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled artisans in the field of this technology to understand the various implementation manners disclosed herein.< / seperator> < / seperator> < / seperator> < / time:xx> < / spk-chg> < / spk-chg>

Claims

1. A method for translating speech data, comprising: receiving source speech data represented in a source language, the source speech data comprising a first speech segment having a first timbre and a second speech segment having a second timbre; as well as Using a machine learning model, a speech translation corresponding to the source speech data is determined, the speech translation is expressed in a target language, and the speech translation includes a first speech translation corresponding to the first speech segment and a second speech translation corresponding to the second speech segment, the first speech translation has the first timbre, and the second speech translation has the second timbre.

2. The method of claim 1, wherein the source speech data is a portion of a sequence of source speech data, and determining the speech translation corresponding to the source speech data comprises: Determining previous text data corresponding to previous speech data in the source speech data sequence, the previous speech data being located before the source speech data; obtaining a previous text translation corresponding to the previous text data; as well as The speech translation corresponding to the source speech data is determined based on the previous text data and the previous text translation using the machine learning model.

3. The method of claim 2, wherein determining the speech translation corresponding to the source speech data based on the previous text data and the previous text translation comprises: Determining, using the machine learning model, a first text translation corresponding to the first speech segment and a second text translation corresponding to the second speech segment based on the previous text data and the previous text translation, respectively; Determine, using the machine learning model, a first speech translation corresponding to the first text translation and a second speech translation corresponding to the second text translation; as well as The first speech translation and the second speech translation are combined using the machine learning model to determine the speech translation.

4. The method of claim 3, wherein determining the first text translation and the second text translation based on the previous text data and the previous text translation, respectively, comprises: Determining, using the machine learning model, a text translation corresponding to the source speech data; as well as The text translation is divided into the first text translation and the second text translation based on a switching mark in the text translation, wherein the switching mark corresponds to a switching point between the first reference speech segment and the second reference speech segment.

5. The method according to claim 2, wherein receiving the source voice data comprises: Extracting speech data blocks from the source speech data sequence according to a predetermined time length; as well as The source voice data is determined from the voice data block. The method according to claim 5 , wherein the previous speech data and the source speech data are located within the speech data block.

7. The method of claim 1, wherein the machine learning model comprises: A translation model, wherein the translation module is used to translate source speech data represented in the source language into a text translation represented in the target language, wherein the text translation includes a plurality of text translations corresponding to a plurality of speech segments in the source speech data, and a switching mark corresponding to a switching point between the plurality of speech segments.

8. The method of claim 7, wherein the translation model is determined based on: obtaining reference previous text data in the source language and a reference previous text translation, the reference previous text data being in the source language and the reference previous text translation being in a target language; Acquire reference source voice data represented in the source language, the reference source voice data comprising a first reference voice segment having the first timbre and a second reference voice segment having the second timbre; Acquire a first reference text translation corresponding to the first reference speech segment and a second reference text translation corresponding to the second reference speech segment; as well as The translation model is updated based on the reference previous text data, the reference previous text translation, the reference source audio data, the first reference text translation, and the second reference text translation. 9 . The method of claim 8 , wherein the first reference text translation includes a reference switching mark corresponding to a switching point between the first reference speech segment and the second reference speech segment.

10. The method according to claim 1, wherein the source voice data is received in a streaming manner.

11. A device for translating speech data, comprising: A receiving module, configured to receive source voice data represented in a source language, wherein the source voice data includes a first voice segment with a first timbre and a second voice segment with a second timbre; as well as A determination module is configured to determine a speech translation corresponding to the source speech data using a machine learning model, wherein the speech translation is represented in a target language and includes a first speech translation corresponding to the first speech segment and a second speech translation corresponding to the second speech segment, wherein the first speech translation has the first timbre and the second speech translation has the second timbre.

12. An electronic device comprising: at least one processor; as well as At least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processor.

13. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, cause the processor to implement the method according to any one of claims 1 to 10.

14. A computer instruction product comprising computer instructions, wherein the computer instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.