Voice translation simultaneous interpretation method and device, translation machine and storage medium

By distributing voice data to multiple servers in real time during the voice translation process for text recognition, translation, and speech synthesis, and employing streaming and semantic segmentation technologies, the problem of long user waiting time is solved, enabling instant display and broadcasting, and improving the user experience.

CN119811386BActive Publication Date: 2025-11-18IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411971835.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-11-18
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

In existing speech translation technologies, users need to wait for speech recognition, translation, and synthesis to complete before they can hear the translation result, which increases the time required for users' sensory experience and affects the user experience.

Method used

By distributing voice data to multiple servers for real-time processing, including text recognition, translation, and speech synthesis, streaming and semantic segmentation technologies are employed to reduce data transmission and processing time.

Benefits of technology

This feature allows users to see the translation results and hear the synthesized audio immediately after lifting the button, significantly reducing user waiting time and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811386B_ABST
    Figure CN119811386B_ABST
Patent Text Reader

Abstract

The application provides a speech translation simultaneous interpretation method and device, a translation machine and a storage medium. The method comprises the following steps: receiving speech data to be recognized, sending the speech data to a first server, obtaining recognized text and displaying it in real time; the first server performs text recognition on the speech data, obtains the recognized text and returns it in real time; the recognized text is sent to a second server, and a translation result is obtained and displayed in real time; the second server translates the recognized text, obtains the translation result and returns it in real time; the translation result is sent to a third server, and synthesized audio is obtained and broadcast in real time; the third server performs speech synthesis on the translation result, obtains the synthesized audio and returns it in real time. The method displays the translation result in real time and broadcasts the synthesized audio in real time, reduces the time length for which the user waits for the translation result and the synthesized audio, thereby reducing the sensory time consumption of the user, enabling the user to see the translation result and hear the synthesized audio broadcast as soon as the user lifts the key, and further improving the use experience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a speech translation and simultaneous interpretation method, apparatus, translator, and storage medium. Background Technology

[0002] With the rapid iteration and updates of AI-powered voice translation technology, a plethora of intelligent translation products are emerging, bringing great convenience to people's cross-language and multi-scenario communication.

[0003] Currently, speech translation scenarios based on AI capability cascading, i.e., speech-to-speech translation scenarios, are implemented by cascading AI capabilities for speech recognition, text translation, and speech synthesis. For example, user A speaks Chinese, which is then recognized as Chinese text. Next, the Chinese text is translated into English text, and then the English text is synthesized into audio using speech AI capabilities for playback, enabling cross-language communication between users A and B.

[0004] However, existing technologies require the user to finish speaking before acquiring all recognition results and performing translation and synthesis. The perceived time for voice translation is the time from when the user lifts the button to when they hear the start of the corresponding audio playback. This perceived time is affected by the length of both the translation and synthesis processes. Furthermore, when the user's spoken audio is long, the translation time increases, and the synthesis time also increases due to the length of the synthesized text. This increased time from both aspects results in a visible waiting period after the user lifts the button, which is when the translation device finishes recording, before seeing the translation result and hearing the corresponding synthesized audio playback, significantly impacting the user experience. Summary of the Invention

[0005] This invention provides a voice translation and simultaneous interpretation method, device, translator, and storage medium to solve the problem that in the existing voice translation scenario, it is necessary to wait for the user to finish speaking before acquiring all recognition results and then performing translation and synthesis. The user's sensory time for voice translation is the time between when the user lifts the button and when they hear the start of the playback of the corresponding audio of the translation result. Such sensory time is affected by the length of translation and synthesis time, thereby reducing the user's experience.

[0006] This invention provides a speech translation and simultaneous interpretation method, comprising the following steps:

[0007] The system receives voice data to be recognized and sends the voice data to a first server to obtain the recognized text and display it in real time; the first server is used to perform text recognition on the voice data, obtain the recognized text, and return it in real time.

[0008] The identified text is sent to a second server to obtain a translation result, which is then displayed in real time. The second server is used to translate the identified text, obtain the translation result, and return it in real time.

[0009] The translation result is sent to a third server to obtain synthesized audio, which is then played back in real time. The third server is used to perform speech synthesis on the translation result, obtain the synthesized audio, and return it in real time.

[0010] According to a speech translation simultaneous interpretation method provided by the present invention, the step of sending the recognized text to a second server to obtain the translation result and display it in real time includes:

[0011] The identified text is semantically segmented to obtain multiple segmented clauses, and the multiple segmented clauses are sent to the second server in the form of a stream to obtain multiple translation results and display them in real time; the second server is used to translate the multiple segmented clauses, obtain the multiple translation results, and return them in real time.

[0012] Accordingly, sending the translation result to a third server to obtain the synthesized audio and broadcast it in real time includes:

[0013] The multiple translation results are concatenated to obtain multiple concatenated translation results. The multiple concatenated translation results are then semantically segmented to obtain multiple target segmentation results. The multiple target segmentation results are then sent to the third server in the form of a stream to obtain multiple synthesized audio files, which are then played back in real time. The third server is used to perform speech synthesis on the multiple target segmentation results to obtain the multiple synthesized audio files, which are then returned in real time.

[0014] According to a speech translation simultaneous interpretation method provided by the present invention, the step of semantically segmenting the multiple concatenated translation results to obtain multiple target segmentation results includes:

[0015] Determine the speech synthesis duration corresponding to the multiple target segmentation results, and the playback duration of the synthesized audio corresponding to the multiple target segmentation results;

[0016] Based on the speech synthesis duration and the broadcast duration, the splitting conditions are determined, and based on the splitting conditions, the multiple splicing translation results are semantically segmented to obtain the multiple target segmentation results.

[0017] According to a speech translation and simultaneous interpretation method provided by the present invention, the splitting condition is that the playback duration of the synthesized audio corresponding to the first segmentation result among the plurality of target segmentation results is greater than the speech synthesis duration of the second segmentation result among the plurality of target segmentation results.

[0018] According to a speech translation simultaneous interpretation method provided by the present invention, the recognized text is semantically segmented to obtain multiple segmented clauses, including:

[0019] Based on the semantic information of each word in the identified text, as well as the context information of each word, the correlation degree between each word and its adjacent words is determined. The correlation degree is used to characterize the degree of consistency between the semantics expressed by each word and its adjacent words.

[0020] If the correlation between any word segment and its adjacent word segment is less than a preset threshold, the word segment and its adjacent word segment are split to obtain the multiple segmented clauses.

[0021] According to the speech translation and simultaneous interpretation method provided by the present invention, the second server is specifically used for:

[0022] The identified text is translated based on language information, and the translation result is returned in real time; the language information is carried by the speech data.

[0023] The present invention also provides a voice translation and simultaneous interpretation device, comprising the following units:

[0024] The first real-time display unit is used to receive the voice data to be recognized, send the voice data to the first server, obtain the recognized text, and display it in real time; the first server is used to perform text recognition on the voice data, obtain the recognized text, and return it in real time.

[0025] The second real-time display unit is used to send the identified text to the second server, obtain the translation result, and display it in real time; the second server is used to translate the identified text, obtain the translation result, and return it in real time.

[0026] The third real-time broadcasting unit is used to send the translation result to the third server to obtain the synthesized audio and broadcast it in real time; the third server is used to perform speech synthesis on the translation result to obtain the synthesized audio and return it in real time.

[0027] The present invention also provides a translator, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the speech translation simultaneous interpretation method as described above.

[0028] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech translation simultaneous interpretation method as described above.

[0029] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the speech translation simultaneous interpretation method as described above.

[0030] The present invention provides a speech translation and simultaneous interpretation method, apparatus, translator, and storage medium. The method receives speech data to be recognized and sends it to a first server to obtain recognized text, which is then displayed in real time. The first server performs text recognition on the speech data, obtaining the recognized text and returning it in real time. The recognized text is then sent to a second server to obtain a translation result, which is also displayed in real time. The second server translates the recognized text, obtaining a translation result and returning it in real time. Finally, the translation result is sent to a third server to obtain synthesized audio, which is then played back in real time. The third server performs speech synthesis on the translation result, obtaining synthesized audio and returning it in real time. This method reduces the time users spend waiting for translation results and synthesized audio by displaying the translation results and playing back the synthesized audio in real time, thereby reducing user sensory input time. Users can see the translation result and hear the synthesized audio as soon as they lift a button, further improving the user experience. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0032] Figure 1 This is a schematic diagram of the key translation provided by the present invention.

[0033] Figure 2 This is the process of calling voice dictation and translation in existing technologies.

[0034] Figure 3 This is one of the flowcharts of the speech translation and simultaneous interpretation method provided by the present invention.

[0035] Figure 4 This is the second flowchart of the speech translation and simultaneous interpretation method provided by the present invention.

[0036] Figure 5 This is a schematic diagram illustrating the data interaction between the translation machine client and server provided by this invention.

[0037] Figure 6 This is a schematic diagram of the speech translation and simultaneous interpretation device provided by the present invention.

[0038] Figure 7 This is a schematic diagram of the translator provided by the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0040] Terms such as "first" and "second" in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category.

[0041] Figure 1 is a schematic diagram of key translation provided by the present invention, Figure 2 is the call process of voice dictation translation in the prior art, such as Figure 1 、 Figure 2 As shown, in the current voice translation scenario implemented based on AI ability cascading, that is, the voice-to-voice translation scenario, here, taking the Chinese-to-English voice translation scenario as an example, its implementation process is: when the user presses a key, the user then emits a voice containing "Hello", and this voice is translated into "Hello" through text. After the translation is completed, the obtained English text "Hello" is subjected to voice synthesis, converting the English text into the corresponding voice signal, and finally generating the English voice audio corresponding to "Hello" and outputting it.

[0042] The implementation method of the above voice-to-voice translation scenario is: cascading the AI capabilities of speech recognition, text translation, and voice synthesis. For example: User A speaks Chinese and is recognized as Chinese recognition text, then all the recognition texts are aggregated. Secondly, the Chinese recognition text is translated into English text, and then the English text is synthesized into an audio through the voice AI ability for broadcast, realizing cross-language communication between User A and B.

[0043] Based on the above problems, the present invention provides a method for simultaneous interpretation of voice translation, Figure 3 is one of the schematic flowcharts of the method for simultaneous interpretation of voice translation provided by the present invention, such as Figure 1 As shown, the method includes step 110, step 120 and step 130.

[0044] Step 110: Receive the voice data to be recognized and send the voice data to the first server to obtain the recognized text and display it in real time; the first server is used to perform text recognition on the voice data, obtain the recognized text and return it in real time.

[0045] Specifically, the execution subject of this method can be a translation machine, i.e., a client, or a server. The following explanation combines the two execution subjects.

[0046] When the executor of this method is a translator, the translator receives the speech data to be recognized. The speech data to be recognized is the speech data that needs to be translated and interpreted in the future. The speech data can come from the user's direct speech or can be obtained through a sound pickup device. Here, the sound pickup device can be a smartphone, a tablet computer, or a smart appliance, such as a speaker, a television, or an air conditioner. After the sound pickup device obtains the speech data through a microphone array, it can also amplify and reduce noise. This embodiment of the invention does not specifically limit this.

[0047] Specifically, the translator can be equipped with a voice button. When the user presses the voice button, the translator can start recording. When the user releases the voice button, the translator can stop recording, thereby obtaining the voice data directly spoken by the user.

[0048] Here, the speech data to be identified can be Chinese speech data, English speech data, Spanish speech data, etc., and the embodiments of the present invention do not specifically limit it.

[0049] After receiving the voice data, the translation device can send it to a first server to obtain the recognized text and display it in real time. Here, the first server is used to perform text recognition on the voice data, obtain the recognized text, and return it in real time.

[0050] Here, the first server can be a recognition engine. It is understood that the text recognition obtained by performing text recognition on speech data can be achieved by a recognition engine that deploys a large language model (LLM). The large language model can be a pre-trained language model such as IFlytek Spark, XLNet (extreme multi-label learning network), or Robustly Optimized BERT approach. This embodiment of the invention does not specifically limit this.

[0051] It should be noted that the translator and the first server can interact via the GRPC (Google Remote Procedure Call) communication protocol or the REST (Representational State Transfer) communication protocol, etc. This embodiment of the invention does not specifically limit the interaction.

[0052] When the execution subject of this method is the server, the server receives the voice data to be recognized, and then sends the voice data to the first server to obtain the recognized text and display it in real time.

[0053] It should be noted that the server and the first server can interact with each other through the GPRC communication protocol or the REST communication protocol, etc. This embodiment of the invention does not make specific limitations on this.

[0054] Understandably, the voice data could be "The weather is so nice today, where shall we go?", and the recognized text could be "The weather is so nice today". The real-time display of the recognized text, that is, displaying "The weather is so nice today" on the translation device, shows that there is no need to aggregate all the recognition results for display, which greatly reduces the time spent by the user's senses, eliminates the waiting time, and improves the user experience.

[0055] Step 120: The identified text is sent to the second server to obtain the translation result and display it in real time; the second server is used to translate the identified text, obtain the translation result and return it in real time.

[0056] Specifically, after obtaining the identified text, the translation machine can send the identified text to a second server to obtain the translation result and display it in real time. Here, the second server is used to translate the identified text, obtain the translation result, and return it in real time.

[0057] Here, the second server can be a translation engine. It is understood that the translation of the recognized text can be achieved by deploying a translation engine with a large language model. The large language model can be a pre-trained language model such as the iFlytek Xinghuo model, XLNet model, or ROBERTa model. This embodiment of the invention does not specifically limit this.

[0058] Here, the second server can translate the identified text based on the language information to obtain the translation result. It should be understood that the language information includes the source language type corresponding to the identified text and the target language type corresponding to the translation result.

[0059] It should be noted that the translator and the second server can interact via the GPRC communication protocol or the REST communication protocol, etc. This embodiment of the invention does not specifically limit the interaction.

[0060] When the execution subject of this method is the server, the server can send the recognized text to the second server to obtain the translation result and display it in real time.

[0061] It should be noted that the server and the second server can interact with each other through the GPRC communication protocol or the REST communication protocol, etc. This embodiment of the invention does not make specific limitations on this.

[0062] Understandably, displaying translation results in real time, without aggregating all translation results for display, greatly reduces the time users spend on the process, eliminates waiting time, and improves the user experience.

[0063] Step 130: The translation result is sent to a third server to obtain synthesized audio and broadcast it in real time; the third server is used to perform speech synthesis on the translation result to obtain the synthesized audio and return it in real time.

[0064] Specifically, after obtaining the translation result, the translation machine can send the translation result to a third server to obtain the synthesized audio and broadcast it in real time. Here, the third server is used to perform speech synthesis on the translation result, obtain the synthesized audio, and return it in real time.

[0065] Considering that the translation results may have defects such as excessively long sentences and excessively long speech synthesis time, the translation results can be concatenated, and then semantic segmentation can be performed based on the concatenated translation results to obtain more reasonable target segmentation results, and then the target segmentation results can be sent to a third server.

[0066] Here, the third server can be a synthesis engine. It is understood that the translation result can be processed to produce synthesized audio, which can be achieved by a synthesis engine that deploys a large language model. The large language model can be a pre-trained language model such as the iFlytek Starfire model, XLNet model, or ROBERTa model. This embodiment of the invention does not specifically limit this.

[0067] It should be noted that the translator and the third server can interact via the GPRC communication protocol or the REST communication protocol, etc. This embodiment of the invention does not specifically limit the interaction.

[0068] When the execution subject of this method is the server, the server can send the recognized text to a third server to obtain the translation result and display it in real time.

[0069] It should be noted that the server and the third server can interact with each other through the GPRC communication protocol or the REST communication protocol, etc. This embodiment of the invention does not make specific limitations on this.

[0070] Understandably, real-time playback of synthesized audio reduces the time users have to wait for the synthesized audio to be played, thereby reducing the time users spend on sensory input. This allows users to hear the synthesized audio start playing as soon as they lift a button, further improving the user experience.

[0071] It should be noted that speech-to-speech translation scenarios typically exist in multiple speech interactions between user A and user B. Therefore, multiple speech interactions correspond to the repeated execution of steps 110 to 130 above.

[0072] Understandably, the first server can encapsulate the recognized text and then return the encapsulated structured data to the translator or server, allowing the translator or server to obtain the recognized text through parsing. Similarly, the second server can encapsulate the translation result and then return the encapsulated structured data to the translator or server, allowing the translator or server to obtain the translation result through parsing. Likewise, the third server can encapsulate the synthesized audio and then return the encapsulated structured data to the translator or server, allowing the translator or server to obtain the synthesized audio through parsing.

[0073] It should be noted that the first server, the second server, and the third server in the above process can be different servers or the same server, that is, the text recognition, translation, and speech synthesis functions are concentrated in one server. This embodiment of the invention does not make specific limitations in this regard.

[0074] The method provided in this invention receives speech data to be recognized and sends the speech data to a first server to obtain recognized text, which is then displayed in real time. The first server performs text recognition on the speech data, obtains the recognized text, and returns it in real time. The recognized text is then sent to a second server to obtain a translation result, which is also displayed in real time. The second server translates the recognized text, obtains a translation result, and returns it in real time. Finally, the translation result is sent to a third server to obtain synthesized audio, which is then played back in real time. The third server performs speech synthesis on the translation result, obtains the synthesized audio, and returns it in real time. This method reduces the time users spend waiting for the translation result and synthesized audio by displaying the translation result and playing the synthesized audio in real time, thereby reducing the time users spend on sensory input. Users can see the translation result and hear the synthesized audio as soon as they lift a button, further improving the user experience.

[0075] Based on the above embodiments, step 120 includes:

[0076] Step 121: Semantically segment the identified text to obtain multiple segmented clauses, and send the multiple segmented clauses to the second server in the form of a stream to obtain multiple translation results and display them in real time; the second server is used to translate the multiple segmented clauses, obtain the multiple translation results and return them in real time.

[0077] Accordingly, step 130 includes:

[0078] Step 131: Concatenate the multiple translation results to obtain multiple concatenated translation results, and perform semantic segmentation on the multiple concatenated translation results to obtain multiple target segmentation results. Send the multiple target segmentation results to the third server in the form of a stream to obtain multiple synthesized audio and broadcast it in real time. The third server is used to perform speech synthesis on the multiple target segmentation results to obtain the multiple synthesized audio and return it in real time.

[0079] Specifically, the identified text is semantically segmented to obtain multiple clauses, which are then streamed to a second server to generate multiple translation results, which are displayed in real time. Here, the second server translates the multiple clauses, generating multiple translation results and returning them in real time, thus achieving a progressive streaming translation display that allows users to quickly see the translation results.

[0080] Here, based on the semantic information of each word in the identified text and the context information of each word, the degree of correlation between each word and its adjacent words can be determined. Then, based on the degree of correlation between each word and its adjacent words, the identified text can be segmented to obtain multiple segmented clauses.

[0081] Understandably, based on the semantic information of each word in the identified text, as well as the contextual information of each word, the correlation between each word and its adjacent words is determined. Then, based on the correlation between each word and its adjacent words, the identified text is segmented to obtain multiple segmented clauses. This is equivalent to semantic segmentation of the identified text based on word part of speech, word structure, etc. It can also combine the punctuation marks at the end of complete clauses in different languages, thereby achieving the goal of cutting long sentences into short sentences according to semantics, improving the accuracy and reliability of semantic segmentation.

[0082] It's important to note that sending multiple segmentation clauses to the second server as a stream involves transmitting these clauses as data blocks, dividing them into smaller chunks, rather than transmitting the entire data block at once. This method reduces bandwidth requirements, especially in network transmission. Furthermore, streaming allows for gradual data processing without loading the entire dataset into memory, thus conserving computational resources.

[0083] Because data packets corresponding to multiple segmented clauses are transmitted in smaller chunks, data streams of any size can be processed without being limited by a fixed size. This is particularly effective for processing real-time generated data or data streams of uncertain size. Through streaming, data packets corresponding to multiple segmented clauses can be continuously transmitted from the translator or server to a second server without interruption or pause.

[0084] Accordingly, considering that if the concatenated translation results are too long, the speech synthesis time will increase accordingly, which may cause the translator to stutter during playback. Therefore, multiple translation results can be concatenated to obtain multiple concatenated translation results. Semantic segmentation can then be performed on these multiple concatenated translation results to obtain multiple target segmentation results. These target segmentation results are then sent to a third server in the form of a stream to obtain multiple synthesized audio files, which are then played back in real time. Here, the third server is used to perform speech synthesis on the multiple target segmentation results, obtaining multiple synthesized audio files and returning them in real time.

[0085] The method provided in this embodiment of the invention takes into account that if the concatenated translation results are too long, the speech synthesis time will increase accordingly. It concatenates multiple translation results to obtain multiple concatenated translation results, performs semantic segmentation on the multiple concatenated translation results to obtain multiple target segmentation results, and sends the multiple target segmentation results to a third server in the form of a stream to obtain multiple synthesized audio and broadcast them in real time. As a result, the user's waiting time for audio broadcast is reduced, the user's sensory time is reduced, and the user can hear the synthesized audio start broadcast as soon as they lift the button.

[0086] Based on the above embodiments, step 131, which involves semantic segmentation of the multiple concatenated translation results to obtain multiple target segmentation results, includes:

[0087] Step 1311: Determine the speech synthesis duration corresponding to the multiple target segmentation results, and the playback duration of the synthesized audio corresponding to the multiple target segmentation results;

[0088] Step 1312: Based on the speech synthesis duration and the broadcast duration, determine the splitting conditions, and based on the splitting conditions, perform semantic segmentation on the multiple splicing translation results to obtain the multiple target segmentation results.

[0089] Specifically, considering that the segmentation conditions for semantic segmentation of multiple spliced ​​translation results are related to their respective speech synthesis duration and the playback duration of the synthesized audio, the speech synthesis duration corresponding to multiple target segmentation results and the playback duration of the synthesized audio corresponding to multiple target segmentation results are determined first.

[0090] Based on the speech synthesis duration and broadcast duration, the segmentation conditions are determined, and based on the segmentation conditions, the multiple splicing translation results are semantically segmented to obtain multiple target segmentation results.

[0091] Here, the splitting condition is that the playback duration of the synthesized audio corresponding to the first segmentation result in the multiple target segmentation results is greater than the speech synthesis duration of the second segmentation result in the multiple target segmentation results.

[0092] Understandably, this splitting condition is mainly to prevent the translator from finishing playing the synthesized audio segment returned in the previous session before the next synthesized audio segment has been returned in real time, which would cause the translator to stutter when playing the audio.

[0093] The method provided in this embodiment of the invention determines the speech synthesis duration corresponding to multiple target segmentation results and the playback duration of the synthesized audio corresponding to multiple target segmentation results. Based on the speech synthesis duration and playback duration, it determines the splitting conditions and performs semantic segmentation on multiple spliced ​​translation results based on the splitting conditions to obtain multiple target segmentation results. This avoids the situation where the translator finishes playing the synthesized audio segment returned in the previous time, but the next synthesized audio segment has not yet been returned in real time, which would cause the translator to play audio with stuttering.

[0094] Based on the above embodiments, step 121 involves semantic segmentation of the identified text to obtain multiple segmented clauses, including:

[0095] Step 1211: Based on the semantic information of each word in the identified text and the context information of each word, determine the degree of association between each word and its adjacent words. The degree of association is used to characterize the degree of consistency between the semantics expressed by each word and its adjacent words.

[0096] Step 1212: If the correlation between any word segment and its adjacent word segment is less than a preset threshold, the word segment and its adjacent word segment are split to obtain the multiple segmented clauses.

[0097] Specifically, based on the semantic information of each word in the identified text and the contextual information of each word, the degree of association between each word and its adjacent words is determined. Here, the degree of association is used to characterize the consistency of the semantics expressed by each word and its adjacent words.

[0098] Semantic information refers to the specific meaning or significance carried by each word in the identified text. It is not limited to the literal meaning of the words, but also includes the meaning of the words in a specific context, culture, or field. Contextual information refers to the specific position of each word in the identified text and the other words or sentences surrounding it.

[0099] Specifically, firstly, the identified text needs to be segmented and labeled with parts of speech. Segmentation divides the continuous identified text into individual words or phrases, while part-of-speech tagging assigns a part-of-speech label (such as noun, verb, adjective, etc.) to each word. Next, the semantic information of each segment needs to be extracted. This can be achieved using pre-trained word vector models, such as word2vec or BERT (Bidirectional Encoder Representations from Transformers). Then, a context embedding model is used to extract the context information corresponding to each segment. After extracting the semantic and context information of the segments, various methods can be used to calculate the correlation between each segment and its neighboring segments. For example, the cosine similarity between the semantic and context information can be calculated as the correlation between each segment and its neighboring segments.

[0100] If the correlation between any word and its adjacent words is less than a preset threshold, then any word and its adjacent words are split to obtain multiple segmented clauses.

[0101] Understandably, by separating word segments with a relevance less than a preset threshold from adjacent word segments (i.e., word segments with low relevance), ambiguity between word segments can be reduced, making the meaning of each segmented clause clearer and more explicit, thus improving the accuracy and reliability of semantic segmentation.

[0102] Based on the above embodiments, the second server is specifically used for:

[0103] The identified text is translated based on language information, and the translation result is returned in real time; the language information is carried by the speech data.

[0104] Specifically, the second server translates the recognized text based on language information, obtains the translation result, and returns it in real time. This language information is carried within the speech data. Specifically, the language information includes the source language type of the recognized text and the target language type of the translated result.

[0105] Understandably, when speech data carries language information, the second server can immediately apply a large language model to translate the speech data from the source language to the target language. This avoids the need for additional time to identify the language during the translation process, thereby improving the real-time nature of the translation.

[0106] Furthermore, when voice data carries language information, it means users don't need to manually select the language or make any additional settings. This not only simplifies the process but also reduces translation errors caused by incorrect language selection, thereby improving the user experience.

[0107] Based on any of the above embodiments, the present invention provides a speech translation simultaneous interpretation method. Figure 4 This is the second flowchart of the speech translation and simultaneous interpretation method provided by this invention. Figure 5 This is a schematic diagram illustrating the data interaction between the translation machine client and server provided by the present invention, as shown below. Figure 4 , Figure 5 As shown, the method includes:

[0108] The first step is for the translation machine client to establish a long connection with the AI ​​server based on the HTTP2 (Hypertext Transfer Protocol version 2) communication custom protocol. Based on the asynchronous bidirectional streaming characteristics of multiplexing and binary framing on a single connection of the HTTP2 protocol, a session and data interaction are established.

[0109] The second step involves the AI ​​server establishing sessions with the recognition engine, translation engine, and synthesis engine, respectively. The AI ​​server interacts with the recognition engine, translation engine, and synthesis engine via the gRPC communication protocol.

[0110] The third step involves the translation client sending an audio stream to the AI ​​server in real time via HTTP / 2 client streaming.

[0111] The fourth step involves the AI ​​server receiving the audio stream in real time and sending it to the recognition engine. The recognition engine performs text recognition on the audio stream, obtains the recognized text, and returns it to the AI ​​server. The AI ​​server then sends the recognized text to the translation client in real time via an HTTP / 2 server stream, where the translation client displays the recognition result in real time. Furthermore, both the audio stream sent by the translation client and the recognized text stream sent by the AI ​​server are implemented within the same session on a single TCP (Transmission Control Protocol) connection.

[0112] The fifth step involves the AI ​​server acquiring the recognized text in real time and performing semantic segmentation of the recognized text into complete clauses of reasonable length, resulting in multiple segmented clauses. These multiple segmented clauses are then streamed into the translation engine, which translates them to obtain multiple translation results. The translation engine then sends these multiple translation results to the translation machine client in real time via HTTP2 server stream, where the translation machine client displays the translation results in real time.

[0113] The sixth step involves the AI ​​server concatenating multiple translation results acquired in real time, resulting in multiple concatenated results. These concatenated results are then semantically segmented. Here, the semantic segmentation algorithm splits the sentence length based on the AI ​​synthesis time to obtain multiple target segmentation results. These target segmentation results are then streamed into the synthesis engine to obtain a synthesized audio stream. This operation is primarily to prevent the translator from playing the previously returned synthesized audio segment before the next synthesized audio segment has been returned in real time, which would cause stuttering during audio playback.

[0114] Finally, the AI ​​server sends the synthesized audio stream to the translator client in real time via HTTP2 server stream. Since the translator client has already gathered the synthesized audio when the user lifts the button, the lift-to-play function can be realized.

[0115] Depend on Figure 4 It can be seen that the real-time audio stream sent by the translation machine client to the AI ​​server, as well as the AI ​​server sending the recognized text stream, the translation result stream, and the synthesized audio stream to the translation machine client, are all implemented in the same session on a single TCP connection. Figure 5 In the process, the translator client first sends Headers and multiple Data streams to the server. Headers typically contain metadata such as request type, content type, and authentication information. Data is the actual content to be transmitted, which may be text, images, audio, etc. When the translator client finishes sending data, it sends a "Data with end flag" identifier. This "Data with end flag" is an effective means of ensuring data integrity and accuracy by adding a special marker to the end of the data stream to indicate the end position, thus preventing data loss or errors. Of course, the server can send the translator client recognized text streams, translated result streams, and synthesized audio streams; that is, the server sends Headers and multiple Data streams to the translator client, thereby achieving data interaction.

[0116] In addition, when the translation client stops sending voice data to the server, it sends an RST_STREAM to the server. RST_STREAM is used to notify of abnormal termination of the stream. When one end wants to close a stream and no longer wants to send or receive data on that stream, it can send an RST_STREAM frame.

[0117] It should be noted that, Figure 5 The client and server establish a session and exchange data on a single TCP connection based on a customized HTTP / 2 communication protocol.

[0118] It should be noted that in the embodiments of the present invention, the AI ​​server and the server refer to the same object.

[0119] Understandably, the HTTP / 2 communication protocol allows multiple requests and responses to occur simultaneously over the same TCP connection. This multiplexing mechanism can significantly reduce latency and connection overhead. Since a new connection doesn't need to be established for each request, network resources can be utilized more efficiently, especially in environments with high network latency or limited connection limits.

[0120] HTTP / 2's customized communication protocol also improves transmission efficiency: by splitting HTTP protocol messages into smaller binary frames for transmission, HTTP / 2 makes parsing and processing more efficient. This reduces resource waste and complexity caused by text parsing, thereby improving overall transmission efficiency.

[0121] Compared with the prior art, the method provided by the embodiments of the present invention greatly reduces the time spent by the user's senses. The longer the user's audio is, the more obvious the superior performance of the method in quickly displaying the translation results and audio synthesis results on the screen. Users can immediately see the translation results and hear the synthesized audio broadcast as soon as they lift the button, without waiting time, which greatly improves the user experience.

[0122] The speech translation and simultaneous interpretation device provided by the present invention is described below. The speech translation and simultaneous interpretation device described below can be referred to in correspondence with the speech translation and simultaneous interpretation method described above.

[0123] Based on any of the above embodiments, the present invention provides a voice translation and simultaneous interpretation device. Figure 6 This is a schematic diagram of the speech translation and simultaneous interpretation device provided by the present invention, as shown below. Figure 6 As shown, the device includes:

[0124] The first real-time display unit 610 is used to receive the voice data to be recognized and send the voice data to the first server to obtain the recognized text and display it in real time; the first server is used to perform text recognition on the voice data, obtain the recognized text and return it in real time.

[0125] The second real-time display unit 620 is used to send the recognized text to the second server, obtain the translation result, and display it in real time; the second server is used to translate the recognized text, obtain the translation result, and return it in real time.

[0126] The third real-time broadcasting unit 630 is used to send the translation result to the third server to obtain the synthesized audio and broadcast it in real time; the third server is used to perform speech synthesis on the translation result to obtain the synthesized audio and return it in real time.

[0127] The apparatus provided in this invention receives speech data to be recognized and sends the speech data to a first server to obtain recognized text, which is then displayed in real time. The first server performs text recognition on the speech data, obtains the recognized text, and returns it in real time. The recognized text is then sent to a second server to obtain a translation result, which is also displayed in real time. The second server translates the recognized text, obtains a translation result, and returns it in real time. Finally, the translation result is sent to a third server to obtain synthesized audio, which is then played back in real time. The third server performs speech synthesis on the translation result, obtains the synthesized audio, and returns it in real time. This method reduces the time users spend waiting for the translation result and synthesized audio by displaying the translation result and playing the synthesized audio in real time, thereby reducing the time users spend on sensory input. Users can see the translation result and hear the synthesized audio as soon as they lift a button, further improving the user experience.

[0128] Based on any of the above embodiments, the second real-time display unit 620 is specifically used for:

[0129] The second real-time display subunit is used to perform semantic segmentation on the identified text to obtain multiple segmented clauses, and send the multiple segmented clauses to the second server in the form of a stream to obtain multiple translation results and display them in real time; the second server is used to translate the multiple segmented clauses, obtain the multiple translation results and return them in real time.

[0130] Accordingly, the third real-time broadcasting unit 630 is specifically used for:

[0131] The real-time broadcasting subunit is used to concatenate the multiple translation results to obtain multiple concatenated translation results, and to perform semantic segmentation on the multiple concatenated translation results to obtain multiple target segmentation results. The multiple target segmentation results are then sent to the third server in the form of a stream to obtain multiple synthesized audio and broadcast them in real time. The third server is used to perform speech synthesis on the multiple target segmentation results to obtain the multiple synthesized audio and return it in real time.

[0132] Based on any of the above embodiments, the real-time broadcasting subunit is specifically used for:

[0133] Determine the speech synthesis duration corresponding to the multiple target segmentation results, and the playback duration of the synthesized audio corresponding to the multiple target segmentation results;

[0134] Based on the speech synthesis duration and the broadcast duration, the splitting conditions are determined, and based on the splitting conditions, the multiple splicing translation results are semantically segmented to obtain the multiple target segmentation results.

[0135] Based on any of the above embodiments, the splitting condition is that the playback duration of the synthesized audio corresponding to the first segmentation result among the plurality of target segmentation results is greater than the speech synthesis duration of the second segmentation result among the plurality of target segmentation results.

[0136] Based on any of the above embodiments, the second real-time display subunit is specifically used for:

[0137] Based on the semantic information of each word in the identified text, as well as the context information of each word, the correlation degree between each word and its adjacent words is determined. The correlation degree is used to characterize the degree of consistency between the semantics expressed by each word and its adjacent words.

[0138] If the correlation between any word segment and its adjacent word segment is less than a preset threshold, the word segment and its adjacent word segment are split to obtain the multiple segmented clauses.

[0139] Based on any of the above embodiments, the second server is specifically used for:

[0140] The identified text is translated based on language information, and the translation result is returned in real time; the language information is carried by the speech data.

[0141] Figure 7 This is a schematic diagram of the translator provided by the present invention, as shown below. Figure 7 As shown, the translator may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a speech translation simultaneous interpretation method. This method includes: receiving speech data to be recognized and sending the speech data to a first server to obtain recognized text and display it in real time; the first server performing text recognition on the speech data to obtain the recognized text and returning it in real time; sending the recognized text to a second server to obtain a translation result and display it in real time; the second server translating the recognized text to obtain the translation result and returning it in real time; sending the translation result to a third server to obtain synthesized audio and broadcast it in real time; and the third server performing speech synthesis on the translation result to obtain the synthesized audio and returning it in real time.

[0142] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0143] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the speech translation simultaneous interpretation method provided by the above methods. The method includes: receiving speech data to be recognized and sending the speech data to a first server to obtain recognized text and display it in real time; the first server is used to perform text recognition on the speech data to obtain the recognized text and return it in real time; sending the recognized text to a second server to obtain a translation result and display it in real time; the second server is used to translate the recognized text to obtain the translation result and return it in real time; sending the translation result to a third server to obtain synthesized audio and broadcast it in real time; the third server is used to perform speech synthesis on the translation result to obtain the synthesized audio and return it in real time.

[0144] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the speech translation simultaneous interpretation method provided by the above methods. The method includes: receiving speech data to be recognized and sending the speech data to a first server to obtain recognized text and display it in real time; the first server performing text recognition on the speech data to obtain the recognized text and returning it in real time; sending the recognized text to a second server to obtain a translation result and display it in real time; the second server translating the recognized text to obtain the translation result and returning it in real time; sending the translation result to a third server to obtain synthesized audio and broadcast it in real time; and the third server performing speech synthesis on the translation result to obtain the synthesized audio and returning it in real time.

[0145] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for simultaneous interpretation of speech translation, characterized in that, include: Receive the voice data to be recognized and send the voice data to the first server to obtain the recognized text and display it in real time; The first server is used to perform text recognition on the voice data, obtain the recognized text, and return it in real time; The identified text is sent to a second server to obtain the translation result, which is then displayed in real time. The second server is used to translate the identified text, obtain the translation result, and return it in real time; The translation results are sent to a third server to obtain synthesized audio, which is then played in real time. The third server is used to perform speech synthesis on the translation result, obtain the synthesized audio, and return it in real time. The step of sending the identified text to the second server to obtain the translation result and display it in real time includes: The identified text is semantically segmented to obtain multiple segmented clauses, and the multiple segmented clauses are sent to the second server in the form of a stream to obtain multiple translation results and display them in real time; the second server is used to translate the multiple segmented clauses, obtain the multiple translation results, and return them in real time. Accordingly, sending the translation result to a third server to obtain the synthesized audio and broadcast it in real time includes: The multiple translation results are concatenated to obtain multiple concatenated translation results. The multiple concatenated translation results are then semantically segmented to obtain multiple target segmentation results. The multiple target segmentation results are then sent to the third server in the form of a stream to obtain multiple synthesized audio files, which are then played back in real time. The third server is used to perform speech synthesis on the multiple target segmentation results to obtain the multiple synthesized audio files, which are then returned in real time. The semantic segmentation of the multiple concatenated translation results yields multiple target segmentation results, including: Determine the speech synthesis duration corresponding to the multiple target segmentation results, and the playback duration of the synthesized audio corresponding to the multiple target segmentation results; Based on the speech synthesis duration and the broadcast duration, the splitting conditions are determined, and based on the splitting conditions, the multiple splicing translation results are semantically segmented to obtain the multiple target segmentation results.

2. The speech translation and simultaneous interpretation method according to claim 1, characterized in that, The splitting condition is that the playback duration of the synthesized audio corresponding to the first segmentation result among the multiple target segmentation results is greater than the speech synthesis duration of the second segmentation result among the multiple target segmentation results.

3. The speech translation and simultaneous interpretation method according to claim 1, characterized in that, The semantic segmentation of the identified text yields multiple segmented clauses, including: Based on the semantic information of each word in the identified text, as well as the context information of each word, the correlation degree between each word and its adjacent words is determined. The correlation degree is used to characterize the degree of consistency between the semantics expressed by each word and its adjacent words. If the correlation between any word segment and its adjacent word segment is less than a preset threshold, the word segment and its adjacent word segment are split to obtain the multiple segmented clauses.

4. The speech translation and simultaneous interpretation method according to any one of claims 1 to 3, characterized in that, The second server is specifically used for: The identified text is translated based on language information, and the translation result is returned in real time; the language information is carried by the speech data.

5. A voice translation and simultaneous interpretation device, characterized in that, include: The first real-time display unit is used to receive the voice data to be recognized, send the voice data to the first server, obtain the recognized text, and display it in real time. The first server is used to perform text recognition on the voice data, obtain the recognized text, and return it in real time; The second real-time display unit is used to send the recognized text to the second server to obtain the translation result and display it in real time; The second server is used to translate the identified text, obtain the translation result, and return it in real time; The third real-time broadcasting unit is used to send the translation result to the third server to obtain the synthesized audio and broadcast it in real time; the third server is used to perform speech synthesis on the translation result to obtain the synthesized audio and return it in real time. The second real-time display unit is specifically used for: The identified text is semantically segmented to obtain multiple segmented clauses, and the multiple segmented clauses are sent to the second server in the form of a stream to obtain multiple translation results and display them in real time; The second server is used to translate the multiple segmented clauses, obtain the multiple translation results, and return them in real time; The third real-time broadcasting unit specifically includes: The third real-time broadcasting subunit is used to concatenate the multiple translation results to obtain multiple concatenated translation results, and to perform semantic segmentation on the multiple concatenated translation results to obtain multiple target segmentation results. The multiple target segmentation results are then sent to the third server in the form of a stream to obtain multiple synthesized audio and broadcast it in real time. The third server is used to perform speech synthesis on the multiple target segmentation results to obtain the multiple synthesized audio and return it in real time. The third real-time broadcast subunit is specifically used for: Determine the speech synthesis duration corresponding to the multiple target segmentation results, and the playback duration of the synthesized audio corresponding to the multiple target segmentation results; Based on the speech synthesis duration and the broadcast duration, the splitting conditions are determined, and based on the splitting conditions, the multiple splicing translation results are semantically segmented to obtain the multiple target segmentation results.

6. A translation machine, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the speech translation and simultaneous interpretation method as described in any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech translation and simultaneous interpretation method as described in any one of claims 1 to 4.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech translation and simultaneous interpretation method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Speech translation method and device, and device for speech translation

    CN107632980A

  • Speech translation interaction method and system

    CN108228575A