Speech processing method for realizing streaming tts

By using the user terminal to detect punctuation and sentence breaks in uncertain text identified by the recognition engine and making comprehensive comparisons, the recognition results are determined in advance and sent in a streaming manner, which solves the problem of long waiting time in existing translation devices and achieves real-time translation.

WO2025241219A1PCT designated stage Publication Date: 2025-11-27SHENZHEN TIMEKETTLE TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/096970
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-23
Filing Date
2024-06-03
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing translation devices and software cannot provide real-time translation during continuous speaking, resulting in excessively long waiting times and impacting communication efficiency.

Method used

By performing punctuation and sentence segmentation detection and comprehensive comparison on the uncertain recognition text responded by the recognition engine through the user terminal, the recognition result is determined in advance and sent to the translation engine for translation and speech synthesis in a streaming manner, and the speech playback speed is dynamically adjusted.

Benefits of technology

It reduces the waiting time for users to hear the translated audio, achieving an effect similar to human simultaneous interpretation and improving communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024096970_27112025_PF_FP_ABST
    Figure CN2024096970_27112025_PF_FP_ABST
Patent Text Reader

Abstract

A method for realizing streaming speech translation. The method comprises: a user performing speech wake-up on a user terminal; the user terminal sending a streaming speech data packet to a recognition engine for speech recognition; the recognition engine constantly returning to the user terminal undetermined recognized text; the user terminal performing sentence segmentation detection on the undetermined recognized text to obtain determined recognized text in advance; the user terminal sending to a translation engine the determined recognized text that is determined in advance, so as to perform translation and complete speech synthesis; and finally, the user terminal playing synthesized speech. In the method, a partial recognition result is determined in advance by means of streaming speech recognition, and the partial recognition result is translated, so that the waiting time for a user to obtain translated speech is reduced, thereby achieving the effect of manual simultaneous interpretation.
Need to check novelty before this filing date? Find Prior Art

Description

Speech processing method for realizing streaming TTS TECHNICAL FIELD

[0001] The present application relates to the technical field of speech translation, in particular to a speech processing method for realizing streaming TTS. BACKGROUND

[0002] With the continuous development of internationalization, communication between different languages is very frequent, but the mutual understanding of languages becomes an obstacle to communication. Most people communicate smoothly by wearing a translator. This way is not only high in cost but also not convenient. At present, there are many translation devices and software on the market. The translation device generally translates after the speaker finishes a sentence, and then the other party hears the translation result and feeds back the communication. This way of waiting for translation is very low in efficiency and affects work efficiency.

[0003] There are also some so-called "simultaneous interpretation" devices. The device first continuously sends voice data packets to the recognition translation engine, and then the recognition translation engine continuously responds to unstable recognition results to the "simultaneous interpretation" device. Only when the duration of the user's pause (i.e. voice pause) reaches a certain time, the recognition translation engine recognizes a certain result from the previously received voice data packet, and then translates and synthesizes the determined recognition result into voice. Although this "simultaneous interpretation" device can achieve continuous recognition and translation result feedback, when the user continuously speaks, due to the continuous change of the recognized content, the user must wait for the recognition engine to return the final determined result before the complete recognized text can be synthesized and converted into voice, which cannot achieve the effect of artificial simultaneous interpretation, that is, the other party can hear the translated voice before the user finishes speaking. This will prolong the waiting time for the user to hear the translated voice, especially when the user continuously expresses long sentences, this problem will be more prominent. SUMMARY

[0004] The purpose of the present application is to overcome the defects and deficiencies of the prior art, and to provide a speech processing method for realizing streaming TTS applied to the field of translation. The method determines the recognition result in advance, and translates and synthesizes the determined recognition result in a streaming manner, reduces the waiting time for hearing the translated voice, and achieves the effect of artificial simultaneous interpretation.

[0005] In order to achieve the above purpose, the present application provides a speech processing method for realizing streaming TTS, which comprises:

[0006] A user starts to speak and voice wakes up a user terminal, the user terminal sends a voice recognition request to a recognition engine, after the voice recognition request is successful, the user terminal sends a streaming voice data packet to the recognition engine for voice recognition, the recognition engine continuously responds to the user terminal with uncertain recognition text,

[0007] The user terminal detects and recognizes the uncertain recognition text, specifically: the user terminal detects the characters of the uncertain recognition text each time the recognition engine responds, when the number of characters is detected for the first time is greater than or equal to n, n is (20, 50), at this time the uncertain recognition text is recorded as B1, the user terminal triggers the first punctuation and sentence detection for the uncertain recognition text B1, and in the subsequent latest response uncertain recognition text, each time n new characters are added based on the length of the uncertain recognition text B1, the punctuation and sentence detection is triggered once.

[0008] Each time the punctuation and sentence detection is successfully triggered, the user terminal performs a punctuation and sentence detection action, if the punctuation and sentence detection of the user terminal reaches a punctuation and sentence condition, it is determined that the sentence is successfully punctuated, and the part of the uncertain recognition text before the punctuation and sentence point this time is identified as the determined recognition text this time;

[0009] After each determination of the recognition text, the user terminal sends the determined recognition text to the translation engine for translation and completes the speech synthesis, and finally the user terminal plays the synthesized speech.

[0010] According to an embodiment of the present application, each time the punctuation and sentence detection is successfully triggered, it is determined whether the latest uncertain recognition text from the recognition engine contains one or more determined recognition texts obtained through successful punctuation before. If the latest uncertain recognition text does not contain the determined recognition text obtained through successful punctuation before, it is directly determined whether the latest uncertain recognition text reaches the punctuation and sentence condition; if the latest uncertain recognition text contains one or more determined recognition texts obtained through successful punctuation before, first, the similarity of the latest uncertain recognition text and the one or more determined recognition texts obtained through successful punctuation before is compared, the one or more determined recognition texts obtained before are ignored, the remaining uncertain recognition text is cut off, and the cut-off uncertain recognition text is subjected to punctuation and sentence detection to determine whether the cut-off uncertain recognition text reaches the punctuation and sentence condition.

[0011] According to one embodiment of the present application, the punctuation condition is that, when detecting the uncertain recognized text to be punctuated, if the uncertain recognized text to be punctuated contains the first type of punctuation symbol, then the punctuation is successful.

[0012] According to one embodiment of the present application, if the uncertain recognized text to be punctuated contains one first type of punctuation symbol, then the first type of punctuation symbol at this time is the punctuation point, and the part of the uncertain recognized text before the punctuation point is determined as the determined recognized text.

[0013] According to one embodiment of the present application, if the uncertain recognized text to be punctuated contains two or more first type of punctuation symbols, then the last first type of punctuation symbol is the punctuation point at this time, and the part of the uncertain recognized text before the punctuation point is determined as the determined recognized text.

[0014] According to one embodiment of the present application, if the uncertain recognized text to be punctuated contains no first type of punctuation symbol, then it is detected whether it contains the second type of punctuation symbol, and if the uncertain recognized text to be punctuated contains two or more second type of punctuation symbols, then the punctuation is successful at this time, the last second type of punctuation symbol is taken as the punctuation point at this time, and the part of the uncertain recognized text before the punctuation point is determined as the determined recognized text; otherwise, the punctuation is failed.

[0015] According to one embodiment of the present application, the first type of punctuation symbol includes period, exclamation mark, question mark, and semicolon.

[0016] According to one embodiment of the present application, the second type of punctuation symbol includes comma.

[0017] According to one embodiment of the present application, the determined recognized text obtained by punctuation is sent to the translation engine for translation each time, and the translated text is synthesized into speech. When the string length of the text that has been synthesized into speech but not played by the user terminal is greater than or equal to 40, the user terminal adjusts the speech playing speed to 1.5 times of the default value; when the string length of the text that has been synthesized into speech but not played by the user terminal is less than 40, the user terminal adjusts the speech playing speed to the default value.

[0018] Another object of the present application is to provide another speech processing method for realizing streaming TTS, which comprises:

[0019] A user starts to speak and voice wakes up the user terminal, the user terminal sends a voice recognition request to an identification engine, after the voice recognition request succeeds, the user terminal sends a streaming voice data packet to the identification engine for voice recognition, the identification engine continuously responds with uncertain recognition text to the user terminal,

[0020] The user terminal detects and identifies the uncertain recognition text, specifically, after the user terminal continuously receives the uncertain recognition text responded by the identification engine for multiple times, the user terminal comprehensively compares and judges the N times of uncertain recognition text recently responded, N is in the range of (2, 6), when the front M sentences of the N times of uncertain recognition text do not change, M is in the range of (1, 5), the front M sentences of the uncertain recognition text are determined as the determined recognition text this time;

[0021] The user terminal sends each determined recognition text to the translation engine for translation and completes voice synthesis, and finally the user terminal plays out the synthesized voice.

[0022] According to an embodiment of the present application, the method further comprises: when the N times of unstable recognition text recently responded by the user terminal all contain one or more determined recognition texts determined before, the N times of uncertain recognition text recently responded are compared with the one or more determined recognition texts determined before in similarity, the one or more determined recognition texts determined before are ignored, the remaining part of the N times of uncertain recognition text is cut out, and then the remaining part of the N times of uncertain recognition text is comprehensively compared and judged, when the front M sentences of the remaining part of the N times of uncertain recognition text do not change, the front M sentences of the remaining part of the N times of uncertain recognition text are determined as the determined recognition text this time.

[0023] According to an embodiment of the present application, the user terminal sends the determined recognition text to the translation engine for translation each time, and performs voice synthesis on the translated text, when the string length of the text that has completed voice synthesis but has not been played out by the user terminal is greater than or equal to 40, the user terminal adjusts the voice playing speed to 1.5 times of the default value; when the string length of the text that has completed voice synthesis but has not been played out by the user terminal is less than 40, the user terminal adjusts the voice playing speed to the default value.

[0024] Compared with the prior art, the present application has the following beneficial effects:

[0025] 1. The application determines the partially determined recognition result in advance through the user terminal punctuating and dividing the uncertain recognition text responded by the recognition engine, and sends the determined recognition result to the translation engine for translation and speech synthesis in a streaming manner, which realizes that the user terminal can obtain or determine the recognition text before the recognition engine responds the determined recognition text or the final recognition text to the user terminal, reduces the waiting time for the user to hear the translated speech, and achieves the effect of artificial simultaneous interpretation.

[0026] 2. The application also determines the partially determined recognition result in advance through the user terminal comprehensively comparing and judging the multiple uncertain recognition texts responded by the recognition engine, and sends the determined recognition result to the translation engine for translation and speech synthesis in a streaming manner, which also can reduce the waiting time for the user to hear the translated speech, and achieves the effect of artificial simultaneous interpretation.

[0027] 3. The application dynamically adjusts the TTS speech text playing speed through the string length judgment of the text which has completed speech synthesis but has not been played by the user terminal, and reduces the playing time of the translated speech. BRIEF DESCRIPTION OF DRAWINGS

[0028] Fig. 1 is a flow chart of the speech processing method for realizing streaming TTS of the application;

[0029] Fig. 2 is a schematic diagram of an embodiment of the application;

[0030] Fig. 3 is a flow chart of another speech processing method for realizing streaming TTS of the application. DETAILED DESCRIPTION

[0031] The application will be further described in detail below in combination with embodiments and drawings, but the embodiments of the application are not limited thereto.

[0032] As shown in FIG. 1, it is a flow diagram of a speech processing method for realizing streaming TTS according to the present application, the method comprises: a user starts to speak and wakes up a user terminal by voice, when the user terminal is woken up, the user terminal sends a speech recognition request to a recognition engine, when the speech recognition request is successful, the user terminal sends a streaming speech data packet to the recognition engine for speech recognition, the recognition engine continuously responds to the user terminal with uncertain recognition text according to the streaming speech data packet, the uncertain recognition text in the embodiment of the present application is actually an intermediate state recognition result sent by the recognition engine, since the intermediate state recognition result in the prior art is not considered as the final result sent by the recognition engine, i.e. the final result is not considered as the certain result sent by the recognition engine, only when the recognition engine detects a speech pause in the user's speaking process for a certain time, the final recognition result will be sent to the user terminal, at this time, the recognition result sent by the recognition engine to the user terminal is the certain recognition result, in addition, it needs to be explained that the uncertain recognition text responded from the recognition engine to the user terminal in the prior art is usually interrupted into different small sentences by punctuation marks, the punctuation marks include period, question mark, comma, exclamation mark, semicolon and the like.

[0033] Since the uncertain recognition text result responded from the translation engine to the user terminal is not sent to the translation engine for translation and TTS synthesis, the time for the user to wait for the final completed translation and synthesized speech will be relatively long, especially when the user speaks continuously, which makes the communication efficiency not high. Therefore, the improvement of the present application lies in that before the user terminal receives the final result sent by the recognition engine, the user terminal detects and recognizes the uncertain recognition text responded from the recognition engine, which is specifically shown as follows: the user terminal detects the characters of the uncertain recognition text responded from the recognition engine each time, when the number of characters is detected to be less than n all the time, the user terminal will not trigger punctuation and sentence detection, until when the number of characters is detected to be greater than or equal to n for the first time, n is in the range of (20, 50), in the embodiment of the present application, the value of n is determined according to the actual situation, and preferably, n is 30. In order to facilitate description, the uncertain recognition text at this time is recorded as B1, the first punctuation and sentence detection is triggered for the uncertain recognition text B1 by the user terminal, and each time n characters are newly added in the subsequent latest responded uncertain recognition text based on the length of the uncertain recognition text B1, the punctuation and sentence detection will be triggered once; and each time the punctuation and sentence detection is successfully triggered, the punctuation and sentence detection action will be executed by the user terminal once, if the punctuation and sentence detection of the user terminal reaches the punctuation and sentence condition, it is determined that the sentence breaking is successful, and the part of the uncertain recognition text before the punctuation and sentence point this time is identified as the certain recognition text this time;

[0034] After determining the recognized text each time, the user terminal sends the determined recognized text to the translation engine for translation and speech synthesis, and finally the user terminal plays the synthesized speech.

[0035] In the embodiment of the present application, each time punctuation and sentence breaking detection is successfully triggered, it is first determined whether the latest uncertain recognized text from the response of the recognition engine contains one or more determined recognized texts obtained through successful sentence breaking before. If the latest uncertain recognized text does not contain the determined recognized texts obtained through successful sentence breaking before, it is directly determined whether the latest uncertain recognized text meets the punctuation and sentence breaking condition; if the latest uncertain recognized text contains one or more determined recognized texts obtained through successful sentence breaking before, it is first compared with the similarity of the latest uncertain recognized text and the one or more determined recognized texts obtained through successful sentence breaking before, the one or more determined recognized texts obtained before are ignored, the remaining uncertain recognized text is cut out, and punctuation and sentence breaking detection is performed on the cut-out uncertain recognized text to determine whether the cut-out uncertain recognized text meets the punctuation and sentence breaking condition.

[0036] In the embodiment of the present application, whether the punctuation and sentence breaking condition is met is specifically that, when detecting the uncertain recognized text which is ready for punctuation and sentence breaking detection, if the uncertain recognized text which is ready for punctuation and sentence breaking detection contains the first type of sentence breaking punctuation symbol, it is determined that the sentence breaking is successful. In the embodiment of the present application, the number of the first type of punctuation symbol sentence breaking success can be one or more, but all are determined as sentence breaking success. When the uncertain recognized text which is ready for punctuation and sentence breaking detection contains one first type of sentence breaking punctuation symbol, the first type of sentence breaking punctuation symbol at this time is the punctuation and sentence breaking point, and the part of the uncertain recognized text before the punctuation and sentence breaking point is determined as the determined recognized text this time. When the uncertain recognized text which is ready for punctuation and sentence breaking detection contains two or more first type of sentence breaking punctuation symbols, the last first type of sentence breaking punctuation symbol is the punctuation and sentence breaking point at this time, and the part of the uncertain recognized text before the punctuation and sentence breaking point is determined as the determined recognized text this time. In the embodiment of the present application, if the uncertain recognized text which is ready for punctuation and sentence breaking detection does not contain the first type of sentence breaking punctuation symbol, it is detected whether it contains the second type of sentence breaking punctuation symbol. When the uncertain recognized text which is ready for punctuation and sentence breaking detection contains two or more second type of sentence breaking punctuation symbols, it is determined that the sentence breaking is successful at this time, the last second type of sentence breaking punctuation symbol is the punctuation and sentence breaking point at this time, and the part of the uncertain recognized text before the punctuation and sentence breaking point is determined as the determined recognized text this time. Otherwise, it is determined that the sentence breaking fails, that is, when the uncertain recognized text which is ready for punctuation and sentence breaking detection does not contain the first type of sentence breaking punctuation symbol and does not contain two or more second type of sentence breaking punctuation symbols, it is determined that the sentence breaking fails. If the sentence breaking fails, no new determined recognized text is generated.

[0037] In the embodiment of the present application, the first type of sentence breaking punctuation symbol is period, exclamation mark, question mark or semicolon, and the second type of sentence breaking punctuation symbol is comma.

[0038] In the embodiment of the present application, the determined recognized text obtained by punctuation and sentence breaking each time is sent to the translation engine for translation, and the translated text is synthesized into speech. In order to prevent too many translated and synthesized TTS speeches from being accumulated, and to make the user wait for a long time for the translated speech, the method of the present application can also adjust the playing speed of the TTS speech of the user terminal, that is, when the string length of the text which has been completed speech synthesis but has not been played by the user terminal is greater than or equal to 40, the user terminal adjusts the speech playing speed to 1.5 times of the default value; when the string length of the text which has been completed speech synthesis but has not been played by the user terminal is less than 40, the user terminal adjusts the speech playing speed to the default value, which is determined according to the speed of normal speaking.

[0039] In order to describe the purpose of the present application in more detail, taking Fig. 2 as an example, "isLast":false,"rText":indicates uncertain recognized text, Arabic numerals "123456789" (including punctuation marks in subsequent lines) represent recognized text, and each line in Fig. 2 is an uncertain recognition result responded by the recognition engine to the user terminal. After the user terminal receives the uncertain recognized text, the user terminal first performs character number detection on the uncertain recognized text. At this time, taking n as 20 as an example for description. Since the user terminal performs character number detection on the characters of the uncertain recognized text of the first four lines, and since the number of characters is less than 20, the user terminal does not trigger punctuation and sentence detection. However, since the number of characters of the fifth line exceeds 20, the user terminal triggers the first punctuation and sentence detection on the uncertain recognized text "123456789, 123456789, 123" of the fifth line, and performs punctuation and sentence detection on the uncertain recognized text of the fifth line. During the detection process, it is found that the uncertain recognized text of the fifth line does not contain a period, i.e., it does not contain a first type of sentence punctuation mark, so successful sentence breaking cannot be performed. Then, the user terminal judges again whether the uncertain recognized text of the fifth line contains two or more commas, i.e., whether it contains two or more second type of sentence punctuation marks. It is found that the uncertain recognized text of the fifth line contains two commas, so the second comma is taken as the punctuation and sentence breaking point of this time, and the recognized text "123456789, 123456789," before the second comma is taken as the determined recognized text of this time. Since the uncertain recognized text of the sixth to ninth lines does not increase 20 characters relative to the uncertain recognized text of the fifth line, the user terminal does not trigger punctuation and sentence detection on each of the sixth to ninth lines.The uncertain recognition text of the 10th row is increased by 20 characters compared with the uncertain recognition text of the 5th row, at this time, the punctuation and sentence breaking detection is triggered for the uncertain recognition text of the 10th row, when the user terminal performs the punctuation and sentence breaking detection for the uncertain recognition text of the 10th row, since the uncertain recognition text of the 10th row contains the determined recognition text "123456789, 123456789," of the 5th row which is successfully broken by punctuation and sentence, at this time, the similarity comparison is performed between the uncertain recognition text of the 10th row and the determined recognition text "123456789, 123456789,", through the similarity comparison, the determined recognition text "123456789, 123456789," obtained previously is ignored in the uncertain recognition text "123456789, 123456789, 123456789. 123456789, 123" of the 10th row, the remaining uncertain recognition text "123456789. 123456789, 123" is intercepted, and the punctuation and sentence breaking detection is performed on the intercepted uncertain recognition text "123456789. 123456789, 123", at this time, the intercepted uncertain recognition text "123456789. 123456789, 123" contains a period, it is judged that the intercepted uncertain recognition text reaches the punctuation and sentence breaking condition, and the part "123456789." before the period is taken as the determined recognition text this time. In the embodiment of the application, the determined recognition text "123456789, 123456789," or "123456789." is sent to the translation engine for translation and voice synthesis, and finally the synthesized voice is played by the user terminal.

[0040] As shown in Fig. 3, the application further provides another voice processing method for realizing streaming TTS, which comprises the following steps:

[0041] The user starts to speak and performs voice wake-up on the user terminal, the user terminal sends a voice recognition request to the recognition engine, after the voice recognition request is successful, the user terminal sends a streaming voice data packet to the recognition engine for voice recognition, the recognition engine continuously responds to the user terminal with uncertain recognition text, the user terminal detects and recognizes the uncertain recognition text, which specifically shows that after the user terminal continuously receives the uncertain recognition text responded by the recognition engine for multiple times, the user terminal comprehensively compares and judges the N times of uncertain recognition text responded recently, N is in the range of (2, 6), when the recognition content of the first M sentences of the N times of uncertain recognition text does not change, M is in the range of (1, 5), it is determined that the recognition content of the first M sentences is the determined recognition text this time;

[0042] The user terminal sends each determined identified text to the translation engine for translation and speech synthesis, and finally plays the synthesized speech. In the embodiment, when the N unstable identified texts recently responded by the user terminal all contain one or more previously determined identified texts, the N unstable identified texts are compared with the one or more previously determined identified texts in similarity, the one or more previously determined identified texts are ignored, the remaining part of the N unstable identified texts is cut out, and then the remaining part of the N unstable identified texts is compared and judged. When the first M sentences of the remaining part of the N unstable identified texts do not change, the first M sentences are determined as the determined identified text. Similarly, the user terminal sends the determined identified text to the translation engine for translation and speech synthesis. When the length of the text string that has been synthesized but not played by the user terminal is greater than or equal to 40, the user terminal adjusts the speech playing speed to 1.5 times of the default value; when the length of the text string that has been synthesized but not played by the user terminal is less than 40, the user terminal adjusts the speech playing speed to the default value. Taking FIG. 2 as an example, N is 5 and M is 2. Since the first four rows of unstable identified texts do not have two sentences, the first two sentences are ignored. From the fifth row of unstable identified text to the tenth row of unstable identified text, there are five unstable identified texts, and each row of unstable identified text contains the same first two sentences "123456789, 123456789,". Therefore, the first two sentences are determined as the determined identified text. Since the unstable identified texts from the sixth row to the ninth row do not have two sentences after the determined identified text "123456789, 123456789," is ignored, they are not compared and judged. However, the unstable identified texts from the tenth row to the fourteenth row have two or more sentences after the determined identified text "123456789, 123456789," is ignored, the unstable identified texts from the tenth row to the fourteenth row are compared with the determined identified text "123456789, 123456789," in similarity, the determined identified text "123456789, 123456789," is ignored, and the remaining unstable identified texts are compared and judged. It is found that the remaining unstable identified texts have the same part "123456789.123456789". Therefore, the remaining unstable identified texts are determined as the determined identified text "123456789.123456789".In the embodiment of the present application, the determined recognition texts "123456789,123456789," and "123456789.123456789" obtained by the method are all sent to the translation engine for translation and voice synthesis, and finally the synthesized voice is played by the user terminal.

[0043] In summary, the present application has the following advantages,

[0044] 1. The present application detects the punctuation and sentence division of the uncertain recognition text recently responded by the recognition engine through the user terminal, determines the partial determined recognition result in advance, and sends the determined recognition result to the translation engine for translation and voice synthesis in a streaming manner. This way, the user terminal can obtain or determine the recognition text before the recognition engine responds the determined recognition text or the final recognition text to the user terminal, which reduces the waiting time for the user to hear the translated voice and achieves the effect of artificial simultaneous interpretation.

[0045] 2. The present application also compares and judges the multiple uncertain recognition texts recently responded by the recognition engine through the user terminal, determines the partial determined recognition result in advance, and sends the determined recognition result to the translation engine for translation and voice synthesis in a streaming manner, which also reduces the waiting time for the user to hear the translated voice and achieves the effect of artificial simultaneous interpretation.

[0046] 3. The present application dynamically adjusts the TTS voice text playing speed by judging the string length of the text that has completed voice synthesis but has not been played by the user terminal, which reduces the playing time of the translated voice.

[0047] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited by the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without deviating from the spirit and principles of the present application are equivalent replacement methods and are included in the protection scope of the present application.

Claims

1. A speech processing method for realizing streaming TTS, the method comprising: a user starts to speak and voice wakes up a user terminal, the user terminal sends a speech recognition request to a recognition engine, after the speech recognition request is successful, the user terminal sends a streaming speech data packet to the recognition engine for speech recognition, and the recognition engine continuously responds to the user terminal with uncertain recognition text, characterized in that the user terminal detects and recognizes the uncertain recognition text, specifically: the user terminal detects characters of the uncertain recognition text responded by the recognition engine each time, when the number of detected characters is greater than or equal to n for the first time, n is (20, 50), at this time, the uncertain recognition text is recorded as B1, the user terminal triggers a first punctuation and sentence breaking detection for the uncertain recognition text B1, and each time n new characters are added to the uncertain recognition text B1, a punctuation and sentence breaking detection is triggered. Each time a punctuation and sentence breaking detection is successfully triggered, the user terminal performs a punctuation and sentence breaking detection action, if the punctuation and sentence breaking detection of the user terminal meets a punctuation and sentence breaking condition, it is determined that sentence breaking is successful, and the part of the uncertain recognition text before the punctuation and sentence breaking point is determined as the determined recognition text of this time. After the determined recognition text is determined each time, the user terminal sends the determined recognition text determined this time to a translation engine for translation and completes speech synthesis, and finally the user terminal plays out the synthesized speech.

2. The speech processing method of implementing streaming TTS according to claim 1, wherein, Each time a punctuation and sentence breaking detection is successfully triggered, it is determined whether the latest uncertain recognition text responded from the recognition engine contains one or more determined recognition texts obtained through successful sentence breaking before. If the latest uncertain recognition text does not contain the determined recognition text obtained through successful sentence breaking before, it is directly determined whether the latest uncertain recognition text meets the punctuation and sentence breaking condition.

3. The speech processing method for implementing streaming TTS according to claim 2, wherein, If the latest uncertain recognition text contains one or more determined recognition texts obtained through successful sentence breaking before, the latest uncertain recognition text is first compared with the one or more determined recognition texts obtained through successful sentence breaking before in terms of similarity, the one or more determined recognition texts obtained before are ignored, the remaining uncertain recognition text is cut out, and punctuation and sentence breaking detection is performed on the cut-out uncertain recognition text to determine whether the cut-out uncertain recognition text meets the punctuation and sentence breaking condition.

4. The speech processing method for implementing streaming TTS according to claim 3, wherein, The punctuation and sentence breaking condition specifically refers to, when detecting the uncertain recognition text to be subjected to punctuation and sentence breaking detection, if the uncertain recognition text to be subjected to punctuation and sentence breaking detection contains a first type of sentence breaking punctuation symbol, it is determined that sentence breaking is successful. If the uncertain recognition text to be subjected to punctuation and sentence breaking detection contains a first type of sentence breaking punctuation symbol, the first type of sentence breaking punctuation symbol at this time is the punctuation and sentence breaking point of this time, and the part of the uncertain recognition text before the punctuation and sentence breaking point is determined as the determined recognition text of this time.

5. The speech processing method for implementing streaming TTS according to claim 3, wherein, If the uncertain recognized text to be punctuated and segmented contains two or more than two first-type punctuation marks, the last first-type punctuation mark is the punctuation segmentation point at this time, and the part of the uncertain recognized text before the punctuation segmentation point is determined as the determined recognized text at this time.

6. The speech processing method for implementing streaming TTS according to claim 3, wherein, If the uncertain recognized text to be punctuated and segmented contains no first-type punctuation mark, it is detected whether it contains a second-type punctuation mark, and if the uncertain recognized text to be punctuated and segmented contains two or more than two second-type punctuation marks, it is determined that the segmentation is successful at this time, the last second-type punctuation mark is taken as the punctuation segmentation point at this time, and the part of the uncertain recognized text before the punctuation segmentation point is determined as the determined recognized text at this time; otherwise, it is determined that the segmentation fails.

7. The speech processing method for implementing streaming TTS according to any one of claims 3 to 6, characterized in that, The first-type punctuation mark includes a period, an exclamation mark, a question mark, and a semicolon.

8. The speech processing method for implementing streaming TTS according to claim 6, wherein, The second-type punctuation mark includes a comma.

9. The speech processing method for implementing streaming TTS according to claim 1, wherein, Each time the determined recognized text obtained by punctuating and segmenting is sent to the translation engine for translation, and the translated text is synthesized into speech, when the string length of the text that has been synthesized into speech but not played by the user terminal is greater than or equal to 40, the user terminal adjusts the speech playing speed to 1.5 times of the default value; when the string length of the text that has been synthesized into speech but not played by the user terminal is less than 40, the user terminal adjusts the speech playing speed to the default value.

10. A speech processing method for implementing streaming TTS, the method comprising: A user starts to speak and performs voice wake-up on a user terminal, the user terminal sends a speech recognition request to an identification engine, after the speech recognition request succeeds, the user terminal sends a streaming speech data packet to the identification engine for speech recognition, the identification engine continuously responds to the user terminal with uncertain recognized texts, The user terminal detects and identifies the uncertain recognized texts, specifically, after the user terminal continuously receives the uncertain recognized texts responded by the identification engine for multiple times, the user terminal comprehensively compares and judges the N times of recently responded uncertain recognized texts, N is in the range of (2, 6), when the first M sentences of the N times of uncertain recognized texts have no change, M is in the range of (1, 5), the first M sentences of the uncertain recognized texts are determined as the determined recognized texts at this time; The user terminal sends each time of the determined recognized texts to the translation engine for translation and completes speech synthesis, and finally the user terminal plays out the synthesized speech.

11. The speech processing method for implementing streaming TTS according to claim 10, wherein, The method further comprises: when the N times of recently responded unstable recognition texts of the user terminal all contain one or more determined recognition texts that have been determined previously, then comparing the N times of recently responded unstable recognition texts respectively with the one or more determined recognition texts that have been determined previously, ignoring the one or more determined recognition texts that have been determined previously, cutting out the remaining part of the N times of unstable recognition texts, and then comprehensively comparing and judging the remaining part of the N times of unstable recognition texts, when the first M sentences of the remaining part of the N times of unstable recognition texts do not change, then the first M sentences of the remaining part of the N times of unstable recognition texts are determined as the determined recognition text.

12. The speech processing method for implementing streaming TTS according to claim 10, wherein, The user terminal sends the determined recognition text to the translation engine for translation each time, and performs speech synthesis on the translated text, when the string length of the text that has completed speech synthesis but has not been played by the user terminal is greater than or equal to 40, the user terminal adjusts the speech playing speed to 1.5 times of the default value, and when the string length of the text that has completed speech synthesis but has not been played by the user terminal is less than 40, the user terminal adjusts the speech playing speed to the default value.

Citation Information

Patent Citations

  • Speech translation method and device, and device for speech translation

    CN107632980A

  • Information processing method, system and device, electronic equipment and storage medium

    CN111711853A

  • Speech translation method, electronic equipment and computer readable storage medium

    CN112735417A

  • Appartus and method for providing simultaneous interpretation service and translation

    KR1020180047789A

  • Translation system, translation method, translation device, and speech input / output device

    WO2019186639A1