Voice simultaneous transmission method, related device, equipment and storage medium

By monitoring the audio accumulation time and switching the summary switch status, refining the unsynthesised text into the summary text and synthesizing it, the delay problem caused by the accumulation of translated content in the voice simultaneous transmission system is solved, and the consistency and delay of voice simultaneous transmission are achieved.

CN120340458APending Publication Date: 2025-07-18ANHUI IFLYREC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510516771.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing voice simultaneous transmission system, the delay problem caused by the accumulation of translated content cannot be effectively solved, especially when the audio length of the target language is long, resulting in auditory incoherence.

Method used

By monitoring the accumulated duration of the audio to be played, switching the summary switch state, refining the text to be synthesized that has not been performed in the on state as the summary text, and replacing it with the original text for synthesis, directly performing speech synthesis in the off state, monitoring the accumulated duration to determine whether to start the summary strategy.

Benefits of technology

It effectively alleviates the accumulation of translated content, reduces the accumulation delay of voice simultaneous transmission, and improves the coherence of voice simultaneous transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340458A_ABST
    Figure CN120340458A_ABST
Patent Text Reader

Abstract

The invention discloses a simultaneous voice transmission method, a related device, equipment and a storage medium, and the method comprises the steps: obtaining a to-be-synthesized text set, and obtaining a to-be-played audio set; determining whether to switch the current state of the abstract switch based on the accumulated duration of the to-be-played audio set; in response to the fact that the current state is the open state, abstract texts extracted from the to-be-synthesized texts which are not subjected to speech synthesis in the to-be-synthesized text set are selected to serve as new to-be-synthesized texts to replace the to-be-synthesized texts which are not subjected to speech synthesis, and speech synthesis is sequentially executed based on the to-be-synthesized text set. Adding the to-be-played audio into a to-be-played audio set; and in response to the current state being the closed state, sequentially executing speech synthesis based on the to-be-synthesized text set, and adding the speech synthesis to the to-be-played audio set as to-be-played audio. According to the scheme, continuous accumulation of the translation content can be relieved, so that accumulation delay of simultaneous voice transmission is reduced as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and particularly relates to a speech simultaneous translation method and related devices, equipment, and storage media. Background Art

[0002] Currently, speech simultaneous translation systems mainly rely on speech-to-text technology to transcribe the source language speech into text, then use machine translation technology to convert the text into the target language, and finally output the audio data of the target language through speech synthesis technology.

[0003] However, under the premise of expressing the same content, the audio lengths of the source language and the target language are not exactly the same. Especially when the audio length of the target language is relatively long, the latency of the speech simultaneous translation system will gradually increase as the translated content accumulates, resulting in discontinuous hearing. Although there are improvement techniques for the time difference in the translation link in the prior art, they cannot solve the simultaneous translation latency caused by the accumulation of translated content. In view of this, how to alleviate the continuous accumulation of translated content to minimize the cumulative latency of speech simultaneous translation has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by the present application is to provide a speech simultaneous translation method and related devices, equipment, and storage media, which can alleviate the continuous accumulation of translated content to minimize the cumulative latency of speech simultaneous translation.

[0005] To solve the above technical problem, in the first aspect of the present application, a speech simultaneous translation method is provided, including: obtaining a set of texts to be synthesized and obtaining a set of audio to be played; wherein, the set of texts to be synthesized contains several texts to be synthesized obtained by sequentially recognizing and translating the source language speech; determining whether to switch the current state of the summary switch based on the cumulative duration of the set of audio to be played; in response to the current state being the on state, selecting the summary text refined from the texts to be synthesized that have not been subjected to speech synthesis in the set of texts to be synthesized as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to speech synthesis, and sequentially performing speech synthesis on the set of texts to be synthesized as the audio to be played and adding it to the set of audio to be played; in response to the current state being the off state, sequentially performing speech synthesis on the set of texts to be synthesized as the audio to be played and adding it to the set of audio to be played.

[0006] To solve the above technical problems, a second aspect of the present application provides a voice simultaneous interpretation device, including: a set acquisition module, a state determination module, a first response module, and a second response module. The set acquisition module is configured to acquire a set of texts to be synthesized and a set of audio to be played; wherein, the set of texts to be synthesized includes a number of texts to be synthesized obtained by sequentially recognizing and translating the source language speech. The state determination module is configured to determine whether to switch the current state of the summary switch based on the cumulative duration of the set of audio to be played. The first response module is configured to, in response to the current state being the on state, select the summary text refined from the texts to be synthesized in the set of texts to be synthesized that have not been subjected to voice synthesis as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to voice synthesis, and sequentially perform voice synthesis based on the set of texts to be synthesized, and add the synthesized audio to the set of audio to be played. The second response module is configured to, in response to the current state being the off state, sequentially perform voice synthesis based on the set of texts to be synthesized, and add the synthesized audio to the set of audio to be played.

[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, at least including a memory and a processor coupled to each other. The memory stores at least program instructions, and the processor is configured to execute the program instructions to implement the voice simultaneous interpretation method in the first aspect above.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the voice simultaneous interpretation method in the first aspect above.

[0009] In the above solution, a set of texts to be synthesized is obtained, and a set of audio to be played is obtained. The set of texts to be synthesized contains several texts to be synthesized obtained by sequentially recognizing and translating the source language speech. Based on the cumulative duration of the set of audio to be played, it is determined whether to switch the current state of the summary switch. In response to the current state being the on state, the summary text refined from the texts to be synthesized in the set of texts to be synthesized that have not been subjected to speech synthesis is selected as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to speech synthesis, and speech synthesis is sequentially performed based on the set of texts to be synthesized and added as the audio to be played to the set of audio to be played. In response to the current state being the off state, speech synthesis is sequentially performed based on the set of texts to be synthesized and added as the audio to be played to the set of audio to be played. Therefore, during the speech simultaneous translation process, the cumulative duration of the set of audio to be played can be monitored to determine whether to start the summary. Thus, when the summary is turned off, speech synthesis can be directly performed, and when the summary is started, the texts to be synthesized in the set of texts to be synthesized that have not been subjected to speech synthesis are refined into summary texts and speech synthesis is performed accordingly. Furthermore, compared with the prior art without a summary strategy, the target language text to be synthesized can be shortened, and correspondingly, the target language audio to be played can also be reduced. Therefore, it is possible to alleviate the continuous accumulation of translation content and reduce the cumulative delay of speech simultaneous translation as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a flowchart of an embodiment of the speech simultaneous translation method of the present application; Figure 2 is a schematic diagram of the process of an embodiment of the speech simultaneous translation device of the present application; Figure 3 is a schematic diagram of the framework of an embodiment of the electronic device of the present application; Figure 4 is a schematic diagram of the framework of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0012] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0013] The terms "system" and "network" are often used interchangeably in this article. The term " / and" in this article only describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the front and rear associated objects. In addition, "multiple" in this article means two or more than two.

[0014] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the voice simultaneous interpretation method of the present application. Specifically, it may include the following steps: Step S11: Obtain the set of texts to be synthesized and obtain the set of audio to be played.

[0015] In the embodiments of the present disclosure, the set of texts to be synthesized includes several texts to be synthesized obtained by sequentially recognizing and translating the source language speech. It should be noted that during the voice simultaneous interpretation process, the source language speaker may continuously or intermittently output the source language speech, and then the source language speech can be recognized and translated to obtain the texts to be synthesized in the target language, which are sequentially added to the set of texts to be synthesized. For the specific implementation of recognition and translation, reference can be made to the technical details related to speech recognition and translation, which will not be elaborated here. In addition, the specific types of the source language and the target language are not limited here. For example, the source language can be Chinese and the target language can be English, and no further examples will be given here.

[0016] In an implementation scenario, taking the source language as Chinese and the target language as English as an example, the source language speech may include, but is not limited to: "Distinguished guests, dear partners: Good day! In this business wave full of opportunities and challenges, we gather here today to jointly explore the path forward. Standing here today, I am deeply honored and full of expectations for the future...". Correspondingly, the set of texts to be synthesized may include the following texts to be synthesized: "Distinguished guests and dear partners", "Good day!", "In this era full of opportunities and challenges in the business world, we are gathered here today to jointly explore the path ahead", "Standing here, I feel deeply honored and filled with anticipation for the future.". For the convenience of distinction, the above four texts to be synthesized can be numbered as "Text to be synthesized 1", "Text to be synthesized 2", "Text to be synthesized 3", and "Text to be synthesized 4" in sequence. Of course, the above example is only a possible example of the set of texts to be synthesized in the actual application process, and the specific content of the set of texts to be synthesized is not limited here, and no further examples will be given.

[0017] In an implementation scenario, the audio in the audio set to be played can be obtained by text-to-speech synthesis of the text in the text set to be synthesized. It should be noted that for the specific process of text-to-speech synthesis, the technical details of text-to-speech synthesis can be referred to and will not be elaborated here. In addition, the audio in the audio set to be synthesized can be played sequentially during the process of simultaneous interpretation. For example, if there is unplayed audio in the audio set to be played, it can be continuously played.

[0018] Step S12: Based on the cumulative duration of the audio set to be played, determine whether to switch the current state of the summary switch.

[0019] In an implementation scenario, the summary switch can be defaulted to the off state. That is to say, generally, automatic summary extraction is not performed by default. Instead, according to the cumulative duration of the audio set to be played, the current state of the summary switch is triggered to be switched to the on state to perform summary extraction.

[0020] In an implementation scenario, it can be detected whether the cumulative duration of the audio set to be played exceeds a first threshold. In response to the cumulative duration having exceeded the first threshold, it can be determined that the current state of the summary switch is switched to the on state, and in response to the cumulative duration not having exceeded the first threshold, it can be determined that the current state of the summary switch is switched to the off state. By the above method, determining whether to switch the current state of the summary switch to the on state or the off state according to whether the cumulative duration exceeds the first threshold can switch to the on state in time when the cumulative duration is too long to start summary extraction, thereby shortening the simultaneous interpretation delay, and maintaining the off state when the cumulative duration is short to maintain normal synthesis.

[0021] In a specific implementation scenario, as described above, the audio set to be played can contain at least one audio. In the case of only one audio, the cumulative duration can be the duration of the unplayed part of the audio. In addition, in the case of multiple audios, if one audio is being played and the rest of the audios are in the to-be-played state, the cumulative duration can be the total duration of all the audios in the to-be-played state, or the cumulative duration can also be the sum of the total duration of all the audios in the to-be-played state and the duration of the unplayed part of the audio being played. Of course, the above examples are only several possible calculation methods of the cumulative duration, and the calculation method of the cumulative duration is not limited here, nor will they be exemplified one by one.

[0022] In a specific implementation scenario, the first threshold can be set according to actual application requirements. For example, when the tolerance for simultaneous translation delay is relatively high, the first threshold can be set appropriately larger; or when the tolerance for simultaneous translation delay is relatively low, the first threshold can be set appropriately smaller. As a special example, the first threshold can be set to 15 seconds. At this time, when it is detected that the cumulative duration has exceeded 15 seconds, the current state of the summary switch can be switched to the on state. Conversely, when it is detected that the cumulative duration has not exceeded 15 seconds, the current state of the summary switch can be maintained in the off state. When the first threshold is set to other values, the same principle can be applied, and no further examples will be given here.

[0023] Step S13: In response to the current state being the on state, select the summary text refined from the to-be-synthesized text that has not been subjected to speech synthesis in the to-be-synthesized text set as the new to-be-synthesized text to replace the to-be-synthesized text that has not been subjected to speech synthesis, and sequentially perform speech synthesis based on the to-be-synthesized text set, and add it as the to-be-played audio to the to-be-played audio set.

[0024] In one implementation scenario, all the text before the last character in the text set to be synthesized can be selected as the target text. On this basis, the abstract text refined from the target text can be selected as the new text to be synthesized to replace the target text, so as to update the text set to be synthesized. Still taking the aforementioned text set to be synthesized as an example, if the speech synthesis of the current "Text to be Synthesized 1" and "Text to be Synthesized 2" has been executed successively, then the remaining "Text to be Synthesized 3" and "Text to be Synthesized 4" (there may also be newly generated texts to be synthesized, which will not be further exemplified here) are left in the text set to be synthesized at this time. Assuming that the current state of the abstract switch is determined to be switched to the on state according to the aforementioned method, the last character in the text set to be synthesized can be determined first. In this example, the last character is the English full stop ".", so all the text before this (i.e., "Text to be Synthesized 3" and "Text to be Synthesized 4") can be used as the target text, and based on this, the target abstract can be refined, such as: "In this business era of opportunities and challenges, we gather to explore the future path, and I feel honored and expectant.", or "Amid business opportunities and challenges, we meet to chart the way forward. I'm honored and full of anticipation.", or "Amidst business opportunities and hurdles, we assemble to seek the future course. I feel honored and anticipatory.", and so on. Of course, the above examples are only several possible examples of the abstract text in the actual application process. The specific content of the abstract text is not limited here, and no further examples will be given. The above method, which selects all the text before the last character in the text set to be synthesized as the target text, and selects the abstract text refined from the target text as the new text to be synthesized to replace the target text to update the text set to be synthesized, can update the text set to be synthesized in a timely manner to assist in abstract refinement.

[0025] In a specific implementation scenario, after the target text is selected, the extraction of the abstract from the target text can be immediately triggered.

[0026] In a specific implementation scenario, after the target text is selected, as another possible implementation method, the summary extraction may not be performed immediately. Instead, it is first detected whether the target text exceeds a second threshold. In response to the target text having exceeded the second threshold, a summary text can be extracted based on the target text, and the summary text is selected as the new text to be synthesized to replace the target text. Conversely, in response to the target text not having exceeded the second threshold, the source language speech recognition and translation can continue to be awaited, and the step of detecting whether the target text exceeds the second threshold is returned. It should be noted that during the continuous waiting process, as the source language speech recognition and translation are performed, new texts to be synthesized are added to the set of texts to be synthesized. At this time, the text length of all the texts before the last character in the set of texts to be synthesized (i.e., the target text) also increases accordingly. Therefore, after returning to the step of detecting whether the target text exceeds the second threshold, it is possible that the target text has exceeded the second threshold, in which case the summary extraction can be performed, or it is also possible that the target text still has not exceeded the second threshold, in which case the waiting can continue. In this way, by repeating this process, the summary extraction can be triggered when a target text of sufficient length is awaited, and the subsequent speech synthesis can be triggered after the summary extraction is completed. That is to say, during the waiting process, the speech synthesis is paused. In addition, the second threshold can be set according to actual application needs. For example, in a situation where the tolerance for simultaneous interpretation delay is relatively high, the second threshold can be set appropriately larger, while in a situation where the tolerance for simultaneous interpretation delay is relatively low, the second threshold can be set appropriately smaller. As a special example, the second threshold can be set to 100 words. In the above method, by detecting whether the target text exceeds the second threshold, in response to the target text having exceeded the second threshold, a summary text is extracted based on the target text, and the summary text is selected as the new text to be synthesized to replace the target text. In response to the target text not having exceeded the second threshold, the source language speech recognition and translation are continued to be awaited, and the step of detecting whether the target text exceeds the second threshold is returned. This can further detect the text length of the target text when the current state of the summary switch is in the on state, trigger the summary extraction when the text length is too long, and continue to wait when the text length is still short, which helps to trigger the summary extraction as timely as possible while appropriately alleviating the frequent triggering of the summary extraction.

[0027] In an implementation scenario, in order to extract the summary text, the compression degree can be determined based on the cumulative duration, and the compression degree is positively correlated with the cumulative duration. Then, the text to be synthesized that has not been subjected to speech synthesis is extracted based on the compression degree to obtain the summary text. In the above method, by determining the compression degree that is positively correlated with the cumulative duration according to the cumulative duration, and then performing the summary extraction based on the compression degree, a summary text with a length that matches can be adaptively extracted following the cumulative duration, which helps to further alleviate the simultaneous interpretation delay.

[0028] In a specific implementation scenario, a mapping relationship between the compression degree and the cumulative duration can be pre-constructed, so that after obtaining the cumulative duration of the audio set to be played, the compression degree can be determined based on this mapping relationship. It should be noted that the mapping relationship can be a linear relationship or a non-linear relationship, which is not limited here.

[0029] In a specific implementation scenario, to improve the efficiency of abstract extraction, a first prompt can be constructed based on the compression degree. It should be noted that the first prompt is used to instruct the large language model to extract an abstract from the text to be synthesized that has not been subjected to speech synthesis according to the compression degree. On this basis, the output content of the large language model in response to the first prompt can be obtained as the abstract text. Exemplarily, still taking the aforementioned target text as an example, the first prompt may include, but is not limited to, the following content: "Please extract an abstract from the following text, with the requirement that the compression degree is [fill in the determined compression degree here]. The following is the text content for which the abstract needs to be extracted: In this era full of opportunities and challenges in the business world, we are gathered here today to jointly explore the path ahead. Standing here, I feel deeply honored and filled with anticipation for the future." In addition, the large language model can be an open-source large model such as Llama, Bloom, etc., or the large language model can also be obtained by fine-tuning the parameters of the open-source large model based on a specific corpus, or the large language model can also be a custom large model. The specific source of the large language model is not limited here. The above method constructs a first prompt based on the compression degree, and the first prompt is used to instruct the large language model to extract an abstract from the text to be synthesized that has not been subjected to speech synthesis according to the compression degree, and then obtains the output content of the large language model in response to the first prompt as the abstract text, which can utilize the general understanding ability of the large language model for abstract extraction and helps to improve the efficiency of abstract extraction. Of course, in the actual application process, it is not necessarily limited to using the large language model to perform abstract extraction. For example, a pre-trained language model such as BERT (Bidirectional Encoder Representation of Transformer, bidirectional encoding representation based on Transformer) can be used for abstract extraction. For the technical details of pre-trained language models such as BERT, please refer to the relevant content, which will not be elaborated here.

[0030] In another implementation scenario, to extract the summary text, different from the foregoing implementation, in addition to determining the compression degree based on the cumulative duration, the context text can also be obtained. It should be noted that the context text includes at least one of the following: the text to be synthesized for which speech synthesis has been performed, and the text to be synthesized newly generated by recognizing and translating the source language speech. For example, the text to be synthesized for which speech synthesis has been performed can specifically be the text to be synthesized for which speech synthesis has been performed before the foregoing target text, and the text to be synthesized newly generated by recognizing and translating the source language speech can specifically be the text to be synthesized newly generated after the foregoing target text. On this basis, the text to be synthesized for which speech synthesis has not been performed can be refined based on the context text and the compression degree to obtain the summary text. In the above manner, in addition to further referring to the context text in the summary extraction process, the accuracy of summary extraction can be improved.

[0031] In a specific implementation scenario, the context text can include the text to be synthesized for which speech synthesis has been performed before the target text; or, the context text can include the text to be synthesized newly generated by recognizing and translating the source language speech after the target text; or, the context text can include: the text to be synthesized for which speech synthesis has been performed before the target text, and the text to be synthesized newly generated by recognizing and translating the source language speech after the target text.

[0032] In a specific implementation scenario, still taking the example of using a large language model to perform abstract extraction, a second prompt can be constructed based on the context text and the compression level, and the second prompt is used to instruct the large language model to refer to the context text and perform abstract extraction on the text to be synthesized that has not been subjected to speech synthesis according to the compression level. Exemplarily, still taking the aforementioned target text as an example, the second prompt may include, but is not limited to, the following content: "Please refer to the following text snippet [fill in the context text here] for abstract extraction, with the required compression level being [fill in the determined compression level here]. The following is the text content for which the abstract needs to be extracted: In this era full of opportunities and challenges in the business world, we are gathered here today to jointly explore the path ahead. Standing here, I feel deeply honored and filled with anticipation for the future.". Based on this, the output content of the large language model in response to the second prompt can be obtained as the abstract text. In addition, the possible sources of the large language model can be referred to the aforementioned relevant descriptions and will not be elaborated here. Of course, as mentioned before, pre-trained language models such as BERT can also be used to perform abstract extraction, which is not limited here. The above method, by constructing a second prompt based on the context text and the compression level, and then obtaining the output content of the large language model in response to the second prompt as the abstract text, can further refer to the context text in addition to the compression level during the abstract extraction process, and can improve the accuracy of abstract extraction.

[0033] In an implementation scenario, after extracting the abstract text from the target text and replacing the target text with the abstract text as the new text to be synthesized, speech synthesis can be sequentially performed based on the set of texts to be synthesized and added as the audio to be played to the set of audio to be played. It should be noted that taking the example of selecting all the text before the last character in the set of texts to be synthesized as the target text, after replacing the target text with the abstract text as the new text to be synthesized, the first text to be synthesized in the set of texts to be synthesized at this time is the abstract text. Of course, it is not excluded that after the first text to be synthesized in the set of texts to be synthesized, there are still new texts to be synthesized generated by the recognition and translation of the source language speech. In addition, the specific process of speech synthesis can refer to the technical details of speech synthesis and will not be elaborated here.

[0034] Step S14: In response to the current state being the closed state, sequentially perform speech synthesis based on the set of texts to be synthesized and add it as the audio to be played to the set of audio to be played.

[0035] In an implementation scenario, when the current state of the summary switch is off, voice synthesis can be directly performed sequentially based on the text set to be synthesized to obtain audio data as the audio to be played, which is then added to the audio set to be played. It should be noted that during this process, as long as there is still unplayed audio in the audio set to be played, continuous playback can be carried out without waiting for other steps. In addition, as long as the source language voice is still being continuously output, recognition and translation can be continuously performed without waiting for other steps. In other words, audio playback and recognition and translation can be regarded as continuously executed background default steps.

[0036] In an implementation scenario, as described above, the audios to be played in the audio set to be played can be played sequentially. Therefore, whether after the execution of "performing voice synthesis sequentially based on the text set to be synthesized and adding it to the audio set to be played" in the previous step S13 or after the execution of "performing voice synthesis sequentially based on the text set to be synthesized and adding it to the audio set to be played" in step S14 here, the previous step S12 "determining whether to switch the current state of the summary switch based on the cumulative duration of the audio set to be played" can be returned. Exemplarily, if the current state of the summary switch switched to the on state during the previous detection of the cumulative duration, when returning to the previous step S12, the current state of the summary switch may switch to the off state due to a significant reduction in the cumulative duration. In this case, voice synthesis can be directly performed sequentially based on the text set to be synthesized without performing summary extraction again; or, if the current state of the summary switch switched to the off state during the previous detection of the cumulative duration, when returning to the previous step S12, the current state of the summary switch may switch to the on state due to a significant increase in the cumulative duration. In this case, summary extraction can be performed to shorten the cumulative delay of simultaneous interpretation. Thus, it can be seen that by detecting the cumulative duration to trigger the switching of the current state of the summary switch to the on state or the off state, it is possible to adaptively decide to perform summary extraction at an appropriate time to reduce the cumulative delay of simultaneous interpretation.

[0037] In the above solution, a set of texts to be synthesized is obtained, and a set of audio to be played is obtained. The set of texts to be synthesized contains a number of texts to be synthesized obtained by sequentially recognizing and translating the source language speech. Based on the cumulative duration of the set of audio to be played, it is determined whether to switch the current state of the summary switch. In response to the current state being the on state, the summary text refined from the texts to be synthesized in the set of texts to be synthesized that have not been subjected to speech synthesis is selected as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to speech synthesis, and speech synthesis is sequentially performed based on the set of texts to be synthesized and added to the set of audio to be played as the audio to be played. In response to the current state being the off state, speech synthesis is sequentially performed based on the set of texts to be synthesized and added to the set of audio to be played as the audio to be played. Therefore, during the speech simultaneous translation process, the cumulative duration of the set of audio to be played can be monitored to determine whether to start the summary. Thus, when the summary is turned off, speech synthesis is directly performed, and when the summary is started, the texts to be synthesized in the set of texts to be synthesized that have not been subjected to speech synthesis are refined into summary texts and speech synthesis is performed accordingly. Furthermore, compared with the prior art without a summary strategy, the target language text to be synthesized can be shortened, and correspondingly, the target language audio to be played can also be reduced. Therefore, the continuous accumulation of translation content can be alleviated, and the cumulative delay of speech simultaneous translation can be reduced as much as possible.

[0038] Please refer to Figure 2 , Figure 2 which is a schematic framework diagram of an embodiment of the speech simultaneous translation device of the present application. The speech simultaneous translation device 20 includes: a set acquisition module 21, a state determination module 22, a first response module 23, and a second response module 24. The set acquisition module 21 is configured to obtain a set of texts to be synthesized and a set of audio to be played; wherein, the set of texts to be synthesized contains a number of texts to be synthesized obtained by sequentially recognizing and translating the source language speech. The state determination module 22 is configured to determine whether to switch the current state of the summary switch based on the cumulative duration of the set of audio to be played. The first response module 23 is configured to, in response to the current state being the on state, select the summary text refined from the texts to be synthesized in the set of texts to be synthesized that have not been subjected to speech synthesis as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to speech synthesis, and sequentially perform speech synthesis based on the set of texts to be synthesized and add it to the set of audio to be played as the audio to be played. The second response module 24 is configured to, in response to the current state being the off state, sequentially perform speech synthesis based on the set of texts to be synthesized and add it to the set of audio to be played as the audio to be played.

[0039] In the above solution, the speech simultaneous translation device 20 obtains a set of texts to be synthesized and a set of audio to be played. The set of texts to be synthesized includes several texts to be synthesized obtained by sequentially recognizing and translating the source language speech. Based on the cumulative duration of the set of audio to be played, it is determined whether to switch the current state of the summary switch. In response to the current state being the on state, the summary text refined from the texts to be synthesized that have not been subjected to speech synthesis in the set of texts to be synthesized is selected as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to speech synthesis, and speech synthesis is sequentially performed based on the set of texts to be synthesized and added to the set of audio to be played as the audio to be played. In response to the current state being the off state, speech synthesis is sequentially performed based on the set of texts to be synthesized and added to the set of audio to be played as the audio to be played. Therefore, during the speech simultaneous translation process, the cumulative duration of the set of audio to be played can be monitored to determine whether to start the summary. Thus, when the summary is turned off, speech synthesis can be directly performed. When the summary is started, the texts to be synthesized that have not been subjected to speech synthesis in the set of texts to be synthesized are refined into summary texts and speech synthesis is performed accordingly. Furthermore, compared with the prior art without a summary strategy, the target language text to be synthesized can be shortened, and correspondingly, the target language audio to be played can also be shortened. Therefore, it is possible to alleviate the continuous accumulation of translation content and reduce the cumulative delay of speech simultaneous translation as much as possible.

[0040] In some disclosed embodiments, the state determination module 22 includes a duration detection sub-module for detecting whether the cumulative duration exceeds a first threshold; the state determination module 22 includes a first determination sub-module for determining that the current state of the summary switch is the on state in response to the cumulative duration having exceeded the first threshold; the state determination module 22 includes a second determination sub-module for determining that the current state of the summary switch is the off state in response to the cumulative duration not having exceeded the first threshold.

[0041] In some disclosed embodiments, the first response module 23 includes a text selection sub-module for selecting all the text before the last character in the set of texts to be synthesized as the target text; the first response module 23 includes a summary refinement sub-module for selecting the summary text refined from the target text as the new text to be synthesized to replace the target text to update the set of texts to be synthesized.

[0042] In some disclosed embodiments, the text selection sub-module includes a length detection unit for detecting whether the target text exceeds a second threshold; the text selection sub-module includes a first response unit for, in response to the target text having exceeded the second threshold, refining the summary text based on the target text and selecting the summary text as the new text to be synthesized to replace the target text; the text selection sub-module includes a second response unit for, in response to the target text not having exceeded the second threshold, continuing to wait for the source language speech to be recognized and translated and returning to the step of detecting whether the target text exceeds the second threshold.

[0043] In some disclosed embodiments, the speech simultaneous translation device 20 includes a degree determination module for determining a compression degree based on the cumulative duration, where the compression degree is positively correlated with the cumulative duration; the speech simultaneous translation device 20 includes a text summary module for refining the text to be synthesized that has not been subjected to speech synthesis based on the compression degree to obtain a summary text.

[0044] In some disclosed embodiments, the text summary module includes a first construction sub-module for constructing a first prompt based on the compression degree, where the first prompt is used to instruct the large language model to perform summary refinement on the text to be synthesized that has not been subjected to speech synthesis according to the compression degree; the text summary module includes a first processing sub-module for obtaining the output content of the large language model in response to the first prompt as the summary text.

[0045] In some disclosed embodiments, the speech simultaneous translation device 20 includes a context acquisition module for acquiring context text, where the context text includes at least one of the following: the text to be synthesized for which speech synthesis has been performed, the text to be synthesized newly generated by recognizing and translating the source language speech; the text summary module is specifically configured to refine the text to be synthesized that has not been subjected to speech synthesis based on the context text and the compression degree to obtain a summary text.

[0046] In some disclosed embodiments, the text summary module includes a second construction sub-module for constructing a second prompt based on the context text and the compression degree, where the second prompt is used to instruct the large language model to refer to the context text and perform summary refinement on the text to be synthesized that has not been subjected to speech synthesis according to the compression degree; the text summary module includes a second processing sub-module for obtaining the output content of the large language model in response to the second prompt as the summary text.

[0047] In some disclosed embodiments, the speech simultaneous translation device 20 includes a loop iteration module for returning a step of determining whether to switch the current state of the summary switch based on the cumulative duration of the set of audio to be played.

[0048] Please refer to Figure 3 , Figure 3 which is a schematic framework diagram of an embodiment of the electronic device of the present application. The electronic device 30 at least includes a memory 31 and a processor 32 that are coupled to each other. At least program instructions are stored in the memory 31, and the processor 32 is configured to execute the program instructions to implement the steps in any of the above-mentioned speech simultaneous translation method embodiments. Specifically, reference can be made to the foregoing disclosed embodiments, which will not be elaborated herein. It should be noted that the electronic device 30 may include, but is not limited to, devices such as learning machines, tablet computers, laptop computers, smart large screens, translation machines, servers, etc., and the specific type of the electronic device 30 is not limited herein.

[0049] Specifically, the processor 32 is used to control itself and the memory 31 to implement the steps in any of the above-described speech simultaneous translation method embodiments. The processor 32 may also be referred to as a CPU (Central Processing Unit). The processor 32 may be an integrated circuit chip with signal processing capabilities. The processor 32 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 32 may be implemented jointly by integrated circuit chips.

[0050] In the above solution, the electronic device 30 obtains a set of texts to be synthesized and a set of audio to be played. The set of texts to be synthesized includes several texts to be synthesized obtained by sequentially recognizing and translating the source language speech. Based on the cumulative duration of the set of audio to be played, it is determined whether to switch the current state of the summary switch. In response to the current state being the on state, the summary text refined from the texts to be synthesized in the set of texts to be synthesized that have not been subjected to speech synthesis is selected as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to speech synthesis, and speech synthesis is sequentially performed based on the set of texts to be synthesized and added as the audio to be played to the set of audio to be played. In response to the current state being the off state, speech synthesis is sequentially performed based on the set of texts to be synthesized and added as the audio to be played to the set of audio to be played. Therefore, during the speech simultaneous translation process, the cumulative duration of the set of audio to be played can be monitored to determine whether to start the summary. Thus, in the case of turning off the summary, speech synthesis is directly performed, and in the case of starting the summary, the texts to be synthesized in the set of texts to be synthesized that have not been subjected to speech synthesis are refined into summary texts and speech synthesis is performed accordingly. Furthermore, compared with the prior art without a summary strategy, the target language text to be synthesized can be shortened, and correspondingly, the target language audio to be played can also be reduced. Therefore, it is possible to alleviate the continuous accumulation of translation content and reduce the cumulative delay of speech simultaneous translation as much as possible.

[0051] Please refer to Figure 4 , Figure 4 FIG. is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 40 stores program instructions 41 that can be run by the processor. The program instructions 41 are used to implement the steps in any of the above-described speech simultaneous translation method embodiments.

[0052] In the above solution, the computer-readable storage medium 40 obtains the text set to be synthesized and the audio set to be played. The text set to be synthesized includes several texts to be synthesized obtained by sequentially recognizing and translating the source language speech. Based on the cumulative duration of the audio set to be played, it is determined whether to switch the current state of the summary switch. In response to the current state being the on state, the summary text refined from the texts to be synthesized in the text set to be synthesized that have not been subjected to speech synthesis is selected as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to speech synthesis, and speech synthesis is sequentially performed based on the text set to be synthesized and added as the audio to be played to the audio set to be played. In response to the current state being the off state, speech synthesis is sequentially performed based on the text set to be synthesized and added as the audio to be played to the audio set to be played. Therefore, during the speech simultaneous translation process, the cumulative duration of the audio set to be played can be monitored to determine whether to start the summary. Thus, in the case of turning off the summary, speech synthesis is directly performed, and in the case of starting the summary, the texts to be synthesized in the text set to be synthesized that have not been subjected to speech synthesis are refined into summary texts and speech synthesis is performed accordingly. Furthermore, compared with the prior art without a summary strategy, the target language text to be synthesized can be shortened, and correspondingly, the target language audio to be played can also be shortened. Therefore, it is possible to alleviate the continuous accumulation of translation content and reduce the cumulative delay of speech simultaneous translation as much as possible.

[0053] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0054] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0055] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.

[0056] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0057] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0058] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0059] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A simultaneous interpretation method for speech, characterized in that, Including: Obtain a set of texts to be synthesized, and obtain a set of audio to be played; wherein, the set of texts to be synthesized contains a number of texts to be synthesized obtained by sequentially recognizing and translating the source language speech. Based on the cumulative duration of the set of audio to be played, determine whether to switch the current state of the summary switch. In response to the current state being the on state, select the summary text refined from the texts to be synthesized that have not been subjected to speech synthesis in the set of texts to be synthesized as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to speech synthesis, and sequentially perform the speech synthesis based on the set of texts to be synthesized, and add the synthesized audio as the audio to be played to the set of audio to be played. In response to the current state being the off state, sequentially perform the speech synthesis based on the set of texts to be synthesized, and add the synthesized audio as the audio to be played to the set of audio to be played.

2. The method according to claim 1, wherein The determining whether to switch the current state of the summary switch based on the cumulative duration of the set of audio to be played includes: Detect whether the cumulative duration exceeds a first threshold. In response to the cumulative duration having exceeded the first threshold, determine to switch the current state of the summary switch to the on state. In response to the cumulative duration not having exceeded the first threshold, determine to switch the current state of the summary switch to the off state.

3. The method according to claim 1, wherein The selecting the summary text refined from the texts to be synthesized that have not been subjected to speech synthesis in the set of texts to be synthesized as the new text to be synthesized to replace the texts to be synthesized that have not been subjected to speech synthesis includes: Select all the text before the last character in the set of texts to be synthesized as the target text. Select the summary text refined from the target text as the new text to be synthesized to replace the target text, so as to update the set of texts to be synthesized.

4. The method according to claim 3, wherein The selecting the summary text refined from the target text as the new text to be synthesized to replace the target text includes: Detect whether the target text exceeds a second threshold. In response to the target text having exceeded the second threshold, refine the summary text based on the target text, and select the summary text as the new text to be synthesized to replace the target text. In response to the target text not having exceeded the second threshold, continue to wait for the source language speech to perform the recognition and translation, and return to the step of detecting whether the target text exceeds the second threshold.

5. The method according to claim 1, wherein The refining step of the summary text includes: Based on the cumulative duration, determine the compression degree; wherein, the compression degree is positively correlated with the cumulative duration. Refine the texts to be synthesized that have not been subjected to speech synthesis based on the compression degree to obtain the summary text.

6. The method according to claim 5, wherein The refining the texts to be synthesized that have not been subjected to speech synthesis based on the compression degree to obtain the summary text includes: Based on the compression degree, construct a first prompt; wherein, the first prompt is used to instruct the large language model to perform summary refinement on the texts to be synthesized that have not been subjected to speech synthesis according to the compression degree. Obtain the output content of the large language model in response to the first prompt as the summary text.

7. The method according to claim 5, wherein Before refining the to-be-synthesized text that has not been subjected to speech synthesis based on the compression degree to obtain the summary text, the method further includes: Obtaining context text; wherein, the context text includes at least one of the to-be-synthesized text that has already undergone speech synthesis and the to-be-synthesized text newly generated by the recognition and translation of the source language speech; The refining the to-be-synthesized text that has not been subjected to speech synthesis based on the compression degree to obtain the summary text includes: Refining the to-be-synthesized text that has not been subjected to speech synthesis based on the context text and the compression degree to obtain the summary text.

8. The method according to claim 7, wherein The refining the to-be-synthesized text that has not been subjected to speech synthesis based on the context text and the compression degree to obtain the summary text includes: Constructing a second prompt based on the context text and the compression degree; wherein, the second prompt is used to instruct the large language model to refer to the context text and perform summary refining on the to-be-synthesized text that has not been subjected to speech synthesis according to the compression degree; Obtaining the output content of the large language model in response to the second prompt as the summary text.

9. The method according to claim 1, wherein After the to-be-played audio in each of the to-be-played audio sets is played sequentially, and after performing the speech synthesis on the to-be-synthesized text set in sequence and adding the resulting to-be-played audio to the to-be-played audio set, the method further includes: Returning to the step of determining whether to switch the current state of the summary switch based on the cumulative duration of the to-be-played audio set.

10. A simultaneous interpretation device for speech, characterized in that, including: A set acquisition module, configured to acquire a to-be-synthesized text set and a to-be-played audio set; wherein, the to-be-synthesized text set contains a number of to-be-synthesized texts obtained by sequentially performing recognition and translation on the source language speech; A state determination module, configured to determine whether to switch the current state of the summary switch based on the cumulative duration of the to-be-played audio set; A first response module, configured to, in response to the current state being the on state, select the summary text refined from the to-be-synthesized text that has not been subjected to speech synthesis in the to-be-synthesized text set as the new to-be-synthesized text to replace the to-be-synthesized text that has not been subjected to speech synthesis, and perform the speech synthesis on the to-be-synthesized text set in sequence and add the resulting to-be-played audio to the to-be-played audio set; A second response module, configured to, in response to the current state being the off state, perform the speech synthesis on the to-be-synthesized text set in sequence and add the resulting to-be-played audio to the to-be-played audio set.

11. An electronic device, characterized in that, At least including a memory and a processor coupled to each other, at least program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the speech simultaneous translation method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, Stored with program instructions that can be run by the processor, and the program instructions are used to implement the speech simultaneous translation method according to any one of claims 1 to 9.