Method for synthesizing speech

By using the template text and slot structure in the preset template library, voice waveforms of slot content are generated in real time and spliced, the problem of response delay and waste of computing resources in intelligent voice synthesis technology is solved, achieving more efficient voice synthesis and reducing costs.

CN114049874BActive Publication Date: 2025-07-29KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111326772.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-07-29
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

The existing intelligent voice synthesis technology has a large response delay in smart device terminals, severe waste of computing resources, affects user experience and increases costs.

Method used

The template text and slot structure in the preset template library are used to generate the voice waveform corresponding to the slot content in real time, and splice it with the pre-acquisitioned template text voice waveform to reduce the calculation amount and processing time.

Benefits of technology

It reduces the response delay of smart device terminals, improves user experience, reduces waste of computing resources, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114049874B_ABST
    Figure CN114049874B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method for synthesizing speech, belonging to the field of artificial intelligence. The method includes: finding a matching preset template that matches the received text in a preset template library, wherein in the preset template library, each preset template includes template text and slots and the speech waveform corresponding to the template text has been pre-obtained; obtaining the slot content corresponding to the slots included in the matching preset template in the received text; generating the slot acoustic features corresponding to the slot content; generating the slot speech waveforms corresponding to the slot acoustic features; and splicing the slot speech waveforms with the speech waveforms corresponding to the template text included in the matching preset template to obtain the speech waveform corresponding to the received text, thereby synthesizing speech. Thereby, the actual computational amount during speech synthesis is reduced, the response delay of the intelligent device terminal is reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] An embodiment of the present invention relates to a method for synthesizing speech. Background Art

[0002] With the rapid iteration of computer performance and the wide application of deep learning technology, intelligent voice interaction technology has achieved unprecedented development. Nowadays, the human-computer interaction interface is undergoing a transformation from touch to voice. The purpose of text-to-speech (TTS) technology is to enable computers and various intelligent devices to speak like humans. Driven by emerging machine learning algorithms, TTS technology can already generate speech that is very close to that of real humans and even indistinguishable from the real thing.

[0003] Since mainstream deep learning algorithms rely on powerful computing power, nowadays, each intelligent device manufacturer and technology service provider deploy this part in a computing center to provide intelligent voice services in the form of cloud computing. However, for TTS, this method has two obvious disadvantages: 1) Due to factors such as the complexity of the algorithm itself and network latency, the response latency of the intelligent device terminal is relatively large, affecting the user experience; 2) Since user requests are unevenly distributed in time (such as peak requests in the morning and evening and low requests in the early morning), in order to cope with peak requests, manufacturers need to prepare oversaturated computing devices, which will cause waste of computing resources and thus increase costs. Summary of the Invention

[0004] To at least partially solve the above problems, an aspect of an embodiment of the present invention provides a method for synthesizing speech, the method including: finding a matching preset template that matches the received text in a preset template library, wherein in the preset template library, each preset template includes template text and slots and the speech waveform corresponding to the template text is pre-obtained; obtaining the slot content corresponding to the slots included in the matching preset template in the received text; generating the slot acoustic features corresponding to the slot content; generating the slot speech waveforms corresponding to the slot acoustic features; and splicing the slot speech waveforms with the speech waveforms corresponding to the template text included in the matching preset template together to obtain the speech waveforms corresponding to the received text, thereby synthesizing speech.

[0005] In addition, another aspect of an embodiment of the present invention further provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to cause a machine to execute the above method.

[0006] In addition, another aspect of an embodiment of the present invention further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.

[0007] Through the above technical solution, a preset template is used to synthesize speech. The speech waveforms corresponding to the template texts in the preset template have been pre-obtained. When synthesizing speech, it is only necessary to generate in real time the speech waveforms corresponding to the slot contents of the slot parts filled in the preset template, and splice the speech waveforms corresponding to the template texts and the speech waveforms corresponding to the slot contents together to synthesize speech, without having to generate in real time the speech waveforms corresponding to the entire received text. In this way, the actual computational load during speech synthesis is reduced, thereby reducing the response latency of the intelligent device terminal and enhancing the user experience. In addition, when synthesizing speech, by means of the preset template to synthesize speech, the processing time of a single task is reduced, and it is not necessary to prepare an oversaturated computing device to withstand the computing intensity at the peak of user requests. In this way, the waste of computing resources is reduced and the cost is lowered.

[0008] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementation, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:

[0010] Figure 1 is a flowchart of a method for synthesizing speech provided by one aspect of the embodiments of the present invention;

[0011] Figure 2 is a schematic diagram of matching a preset template provided by another embodiment of the present invention;

[0012] Figure 3 is a schematic diagram of compensation signal illustration provided by another embodiment of the present invention;

[0013] Figure 4 is a schematic diagram of the logic of training a compensation model provided by another embodiment of the present invention;

[0014] Figure 5 is a schematic diagram of compensation logic provided by another embodiment of the present invention;

[0015] Figure 6 is a schematic diagram of compensating adjacent acoustic features provided by another embodiment of the present invention; and

[0016] Figure 7 is a schematic diagram of generating slot acoustic features and compensating adjacent acoustic features provided by another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following will describe in detail the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.

[0018] In an intelligent voice interaction system, there is a natural language generation module (NLG) that generates the text content for the system to respond to the user. The TTS module will receive this text and convert it into speech for output to the user. A mainstream design method of NLG is to design dialogue templates, which contain various slots. Different response texts can be generated by filling different values in the slots. For example, a dialogue template example is: The current time is [hour] o'clock [minute] minutes in the morning. In this dialogue template, there are 2 slots. By filling different contents in the 2 slots, a type of response text for answering the current time can be generated. Response text examples: The current time is 6:58 in the morning; The current time is 7:45 in the morning; The current time is 8:20 in the morning. Through statistics, it is found that in a real intelligent voice interaction system, this type of structured response text generated through dialogue templates appears very frequently, especially during the peak request periods in the morning and evening. The design idea of the technical solution provided by the embodiments of the present invention is to pre-make preset templates for these high-frequency dialogue templates, and only generate the speech corresponding to the content filled in the slots when synthesizing speech online in real time. Taking the previous example, pre-make the speech content of "The current time is [hour] o'clock [minute] minutes in the morning." as a preset template, where [] represents a slot. During real-time speech synthesis, only the speech segments of numbers such as "6" and "58" need to be generated and spliced into the template as the final output of TTS. Among them, the TTS algorithm solution includes 3 functional modules: a linguistic feature generation module, an acoustic feature generation module, and a speech waveform generation module. Linguistic features are generated from the text, then acoustic features are generated from the linguistic features, and then a speech waveform is generated from the acoustic features, thereby obtaining a speech signal.

[0019] One aspect of the embodiments of the present invention provides a method for synthesizing speech. Figure 1 is a flowchart of a method for synthesizing speech provided by an embodiment of the present invention. As Figure 1 shown, the method includes the following content.

[0020] In step S10, a matching preset template that matches the received text is found in the preset template library. Among them, in the preset template library, each preset template includes template text and slots, and the speech waveform corresponding to the template text has been obtained in advance. The received text is the text for which speech synthesis is to be performed. In the preset template, the slots are not filled with text. In a preset template, one or more slots may be included. The preset template in the preset template library that matches the received text is the matching preset template. For finding a matching preset template that matches the received text in the preset template library, for example, each preset template in the preset template library has a template ID. For the case where template information can be obtained from the NLG (Natural Language Generation Module), the template ID of the preset template that matches the received text can be directly obtained, and the matching preset template that matches the received text is found in the preset template library according to the template ID. In addition, in the embodiments of the present invention, in the preset template library, the acoustic features and / or acoustic feature hidden states and / or speech waveform hidden states corresponding to the template text may also be obtained in advance, where the acoustic feature hidden state is the hidden state of the acoustic feature generation network when generating acoustic features, and the speech waveform hidden state is the hidden state of the speech waveform generation network when generating the speech waveform corresponding to the acoustic features. Speech is a typical time-series signal. When processing time-series signals with a neural network, the network usually needs to "remember" historical information, such as the cell state c (cell state) and the hidden state h (hidden state) in the LSTM. The hidden state is not equal to the model parameters of the neural network. The model parameters are generally some fixed-value weights, while the hidden state refers to the state parameters and / or intermediate values that change with time during the inference process. Still taking the LSTM as an example, the hidden state at time t can be expressed as [c t 、h t . In a real model, there may be multiple LSTM layers, and each layer may have multiple LSTM units. Then, in the real model, the hidden state is a matrix sequence. In the process of using the acoustic feature generation network to generate the acoustic features of the speech template, assuming that the last frame of the acoustic features of the template text is generated at time t, and the first frame of the acoustic features of the slot text is generated at time t + 1, then the hidden state [c t 、h t is exported and stored as part of the speech template before the end of time t; similarly, the hidden state corresponding to the speech waveform generation network can also be exported and stored as part of the speech template.

[0021] In step S11, the slot content corresponding to the slots included in the matching preset template in the received text is obtained. For example, the received text may be compared with the matching preset template to obtain the slot content corresponding to the slots.

[0022] In step S12, slot acoustic features corresponding to the slot content are generated. For example, the slot acoustic features can be generated according to a streaming speech synthesis algorithm.

[0023] In step S13, slot speech waveforms corresponding to the slot acoustic features are generated. For example, the slot speech waveforms can be generated according to a streaming speech synthesis algorithm.

[0024] In step S14, the slot speech waveforms are spliced together with the speech waveforms corresponding to the template text included in the matching preset template to obtain the speech waveform corresponding to the received text, thereby synthesizing speech.

[0025] Through the above technical solution, a preset template is used to synthesize speech. The speech waveforms corresponding to the template text in the preset template have been pre-obtained. When synthesizing speech, only the speech waveforms corresponding to the slot content of the slot part filled in the preset template need to be generated in real time, and the speech waveforms corresponding to the template text are spliced together with the speech waveforms corresponding to the slot content to synthesize speech. There is no need to generate the speech waveform corresponding to the entire received text in real time. In this way, the actual computational amount during speech synthesis is reduced, thereby reducing the response delay of the intelligent device terminal and improving the user experience. In addition, when synthesizing speech, the preset template is used to synthesize speech, reducing the processing time of a single task. There is no need to prepare an over-saturated computing device to withstand the computing intensity at the peak of user requests. In this way, the waste of computing resources is reduced and the cost is lowered.

[0026] Optionally, in the embodiments of the present invention, establishing a preset template library includes two aspects. One is the screening of preset template texts, and the other is the production of preset templates. The screening of preset template texts can adopt the following screening methods: 1) Obtain all designed dialogue templates from the NLG module; 2) Delete dialogue templates with sparse text content. For example, a dialogue template with sparse content can be a dialogue template whose average segment length is less than 2 characters after segmenting the template text by slots; 3) Delete fixed phrases, that is, templates with zero slot quantity; 4) Use the remaining dialogue templates as the preset templates of this proposal. In the production of preset templates, a simple method is to pre-record the preset templates by recording audio. However, this method is restricted in many ways. For example, it is only suitable for a few specific voices and cannot be extended to more voices; it is only suitable for a small number of pre-made dialogue templates and cannot be extended to later added dialogue templates and high-frequency texts. In the embodiments of the present invention, the production of preset templates includes the following content. 1) Fill in variable instances in the slots of the dialogue template to complete the template into a complete and smooth text. 2) Use the Streaming Speech Synthesis algorithm to synthesize the text after completion. 3) Record some intermediate values during the synthesis process, such as the acoustic features of the template text, the hidden states during the prediction of the acoustic process (the hidden states of the acoustic feature generation network when generating acoustic features), the hidden states of the vocoder when generating the waveform (the hidden states of the speech waveform generation network when generating the speech waveform from the acoustic features), and the speech waveform of the template text, etc., so as to quickly restore the inference context when generating slot voices online (that is, the input data, hidden states, and other variables and necessary computing resources corresponding to predicting this frame). 4) For the results output in steps 2) to 3) (i.e., acoustic features, speech waveforms, hidden states), delete the parts corresponding to the slot content. 5) Record the template number, template text, slot position description information, and the results of step 4) as a preset template. The slot position description information includes the quantity and order of the slots, and the offset positions of each slot in the text, acoustic features, and speech waveform data of the template. For example, the current time is [hour] o'clock [minute] in the morning, the offset of [hour] is 7, and the offset of [minute] is 8. Through the above content, a preset template library is established.

[0027] Optionally, in the embodiments of the present invention, the following content can be adopted to find a matching preset template in the preset template library that matches the received text. Among them, the following method of finding a matching preset template in the preset template library is applicable to the situation where template information cannot be obtained from the NLG. In most voice interaction systems, it is inconvenient to interact information between TTS and NLG. At this time, it is necessary to match the most suitable template in the preset template library by passing in the text of TTS.

[0028] Sort all the characters included in the received text in the order from the head to the tail of the received text. For example, if the received text includes m characters, sort the m characters, and use c1, c2, …, c m to represent, where c i represents the i-th character. For any preset template in the preset template library, perform the following matching operations to match all the preset templates in the preset template library with the received text and determine the matching preset template. Segment the template text corresponding to the preset template based on the slots to obtain segmented template text. According to the positions of the segmented template text in the preset template, sort the segmented template text in the order from the head to the tail of the preset template, that is, sort all the segmented template text. For example, the template text corresponding to a preset template includes n segmented template texts, which are represented by t1, t2, …, t n to represent, and t i represents the i-th segmented template text. In addition, the slots included in the preset template can also be sorted in the order from the head to the tail of the preset template. For example, there are n - 1 slots in total, which are represented by s1, s2, …, s n-1 to represent, and s i represents the i-th slot. Specifically, it can be understood with reference to Table 1. According to the sorting of the template text, sequentially search for the segmented template text in the received text. Among them, the characters in the received text that are the same as the characters included in a segmented template text are no longer used for subsequent searches. Here, the search means finding a string in the received text that is the same as the segmented template text. When searching for the segmented template text in the received text, search in the order of the segmented template text. For example, taking the above segmented template text t1, t2, …, t n as an example, first find t1, then find t2, and so on. When some characters in the received text are the same as a segmented template text, those characters are no longer used for subsequent search work. For example, taking the above example, the characters in the received text are represented by c1, c2, …, c m to represent, and the segmented template text corresponding to the preset template is represented by segmented template text t1, t2, …, t n to represent. When the characters c1, c2, c3, c4 are the same as the characters in the segmented template text t1, the characters c1, c2, c3, c4 are no longer used for the segmented template text t2, …, t nFor the search work, the search for the segmented template text t2 starts from the character c5. For a segmented template text, finding the segmented template text in the received text means finding an exactly identical string in the received text. For example, as shown in Table 1, for the segmented template text t1 "The current time is morning", when it is found in the received text, it must be 7 characters and must be exactly "The current time is morning". When all segmented template texts are found in the received text, according to the matching preset conditions, it is determined whether the preset template matches the received text successfully; when the preset template matches the received text successfully, according to the optimal template setting rule, it is determined whether to set the preset template as the optimal template, where, when all preset templates in the preset template library have completed the matching operation, the optimal template is the matching preset template. If all preset templates in the preset template library have completed the matching operation but there is no optimal template, it means that no matching preset template is found in the preset template library. In addition, in the received text, the characters that are not included in the template text of the matching preset template are the characters corresponding to the slots. Additionally, the slots and segmented template texts included in the matching preset template can be sorted simultaneously, and by comparing with the received text, the slot content corresponding to each slot included in the matching preset template is determined.

[0029] Table 1

[0030]

[0031] Optionally, in the embodiments of the present invention, sequentially searching for segmented template texts in the received text according to the template text sorting includes the following content. Search for the first segmented template text in the received text according to the following content, where the first segmented template text is the segmented template text that ranks first in the preset template according to the template text sorting. For example, for the segmented template texts t1, t2,..., t n , t1 is the first segmented template text. When there is no slot at the head of the first segmented template text, compare the first segmented template text with the string in the received text whose character length is the same as that of the first segmented template text starting from the first character according to the character sorting, where the first character is the character that ranks first in the received text according to the character sorting. For example, there is no slot at the head of t1. If among the characters c1, c2,..., c m Search for the segmented template texts t1, t2,..., t n, if t1 consists of 7 characters, then only t1 can be compared with the strings c1 - c7. c1 is the first character. When there is a slot at the head of the first segmented template text, the first segmented template text is compared with any string in the received text that has the same character length as the first segmented template text and is consecutive according to the character sorting order. That is, when there is a slot at the head of the first segmented template text, there is no restriction on the string in the received text for comparison. When the first segmented template text is found in the received text, the matching position for the second segmented template text is set according to the character length of the first segmented template text and the sequence number of the starting character in the string that matches the first segmented template text in the received text according to the character sorting order. Here, the second segmented template text is the next segmented template text after the first segmented template text sorted according to the template text, and the matching position is the sequence number of the starting character for searching the segmented template text in the received text according to the character sorting order. For example, for segmented template texts t1, t2, …, t n , t2 is the second segmented template text. Taking the above example, the string that matches t1 is c1 - c7, the length of the string of t1 is 7, and the sequence number of the starting character c1 for comparison with t1 is 1, then the matching position for t2 is 8, that is, starting from c8, t2 is compared. Starting from the matching position, search for the second segmented template text in the received text; when the second segmented template text is found, update the matching position according to the character length of the second segmented template text and the sequence number of the starting character in the string that matches the second segmented template text in the received text according to the character sorting order. It can be understood by referring to the content of setting the matching position for the second segmented template text above, and search for the segmented template texts other than the first segmented template text and the second segmented template text in the preset template in the received text in sequence according to the content of searching for the second segmented template text until all the segmented template texts in the preset template are searched. Here, during the process of searching for all the segmented template texts including the first segmented template text, the second segmented template text and those after them in the preset template, if any segmented template text is not found, the preset template fails to match the received text, and the operation of sequentially searching for the segmented template texts in the received text according to the template text sorting ends. For example, taking the above example, when t2 is found and no matching string is found, the search operation for this preset template ends, and this preset template fails to match the received text.

[0032] Optionally, in the embodiments of the present invention, matching the preset conditions includes the following. If all the segmented template texts are found in the received text and the maximum serial number of the characters sorted by characters in the received text is greater than or equal to the last matching position corresponding to the last segmented template text and there is no slot at the tail of the last segmented template text, then the preset template fails to match successfully with the received text, where the last segmented template text is the segmented template text ranked last according to the sorting of the template texts in the preset template, and the last matching position is the matching position set according to the character length of the last segmented template text and the serial number of the starting character sorted by characters in the string matching the last segmented template text in the received text. For example, when searching for segmented template texts t1, t2,..., t5 in characters c1, c2,..., c9, the maximum serial number is 9, the last segmented template text is t5, there is no slot at the tail of t5, t5 includes 2 characters, and the serial number of the starting character sorted by characters in the string matching t5 among characters c1, c2,..., c9 is 6, then the last matching position is 8, and the maximum serial number 9 is greater than the last matching position 8, indicating that there are extra characters, so the matching fails. If all the segmented template texts are found in the received text and the maximum serial number is less than the last matching position corresponding to the last segmented template text and / or there is a slot at the tail of the last segmented template text, then the preset template matches successfully with the received text.

[0033] Optionally, in the embodiments of the present invention, the optimal template setting rule includes the following. If the optimal template is not set, then set the preset template as the optimal template, that is, set the currently successfully matched preset template as the optimal template. If the optimal template has been set, but the number of characters included in the preset template is more than the number of characters included in the optimal template, then set the preset template as the optimal template. If the optimal template has been set and the number of characters included in the preset template is the same as the number of characters included in the optimal template, but the number of slots included in the preset template is less than the number of slots included in the optimal template, then set the preset template as the optimal template. If the optimal template has been set, but the number of characters included in the preset template is less than or equal to the number of characters included in the optimal template, then do not set the preset template as the optimal template. If the optimal template has been set and the number of characters included in the preset template is the same as the number of characters included in the optimal template, but the number of slots included in the preset template is greater than or equal to the number of slots included in the optimal template, then do not set the preset template as the optimal template.

[0034] Searching for segmented template texts in the received text means matching the segmented template texts with the characters in the received text. A string that is matched is considered found, and a string that is not matched is considered not found. For example, for the above segmented template texts t1, t2,..., t n and characters c1, c2,..., c mFor example, refer to Figure 2 to give an exemplary introduction to the search process. During the search, each segmented template text in the t sequence of the segmented template text is sequentially searched in the received text. If the entire t sequence can be found and there are no extra characters in the received text, it is considered a successful match; otherwise, it is a failed match. As Figure 2 shown, in Figure A, t1 has been matched to c1 and c2, so t2 should start searching from c3; Figure B shows that t2 has been matched to c4 and c5. Figure C shows that t n has been matched to c m-1 , c m and there are no more characters in the received text, then the preset template matches the received text successfully. For the same received text segment, multiple preset templates may match successfully. In this case, the preset template with the most characters is selected. If the number of characters is the same, the preset template with the fewest slots is selected.

[0035] Specifically, the matching algorithm may include the following content, and the following operations are performed on all preset templates in the preset template library.

[0036] Step 1: Segment the template text by slots. The segmented template text is denoted as t1, t2, …, t n , and the slots are denoted as s1, s2, …, s n-1 (when the number of slots is less than the number of text segments), or s1, s2, …, s n (when the number of slots is equal to the number of text segments), or s1, s2, …, s n+1 (when the number of slots is more than the number of text segments). Sort the segments in the order they appear in the template. The results may be: (1) t1, s1, t2, s2, …, s n-1 , t n ; (2) t1, s1, t2, s2, …, s n-1 , t n , s n ; (3) s1, t1, s2, t2, …, s n , t n ; (4) s1, t1, s2, t2, …, s n , t n , s n+1 .

[0037] Step 2: Sequentially search for each segmented template text in the received text c1, c2, …, c m (since the slots (i.e., the s sequence) can match 0 to multiple characters, only the text segments (i.e., the t sequence) are searched).

[0038] Step 2.1: If it is a preset template of type (1) or (2), compare t1 with c1, c2, …, c L1, L1 is the length of t1. If they are the same, update the matching position search_beg = L1 + 1 and go to step 2.2; otherwise, the preset template matching fails. If it is a type (3) or (4) template, search for t1 in c1, c2, …, c m If a string exactly the same as t1 is found at c P1 , and c P1 is the starting character of this string, then update the matching position search_beg = P1 + L1 and go to step 2.2; otherwise, the preset template matching fails. In addition, the content corresponding to the slot s1 is between c1 and c P1-1 .

[0039] Step 2.2: If the last segmented template text t n has been successfully matched, update the matching position. According to t n and the serial number of the starting character that matches t m in c1, c2, …, c n , then go to step 2.3; otherwise, start searching for the next template text segment t i from the search_beg of the input text. If a string exactly the same as t Pi is found at c i , then update the matching position search_beg = P i +L i and re-enter step 2.2. L i is the length of t i , otherwise the preset template matching fails.

[0040] Step 2.3: If m ≥ search_beg and the preset template is of type (1) or (3), it means there are extra characters in the received text relative to this preset template, then the preset template matching fails; otherwise, the template matching is successful and go to step 3.

[0041] Step 3: If the optimal template is not set, set the current preset template as the optimal template; otherwise, compare the number of characters of the current preset template and the optimal template. If the number of characters of the current preset template is more, set the current template as the optimal template; if the number of characters is the same, compare the number of slots of the current preset template and the optimal template. If the number of slots of the current preset template is less, set the current preset template as the optimal template. After all preset templates are matched, if the optimal template is set, select the optimal template; otherwise, the received text fails to match a template in the preset template library and the matching fails.

[0042] Since the voice signal is a time-series signal, it is greatly affected by the preceding and following voices. If the voice of the slot content is directly generated and directly spliced into the preset template, only a very mechanical voice will be generated, that is, the voice has poor naturalness due to inconsistencies in fundamental frequency, energy, speech rate, and prosody, and the glitch noise caused by discontinuous phases, etc. In the embodiments of the present invention, in order to ensure the continuity of the synthesized voice, certain adjustments are also required for the acoustic features of the partial frames adjacent to the slot in the preset template itself. Specifically, in the preset template library, the acoustic features corresponding to the template text included in each preset template are pre-obtained, and the method further includes the following content. Compensate the adjacent acoustic features in the acoustic features corresponding to the template text included in the matching preset template to obtain compensated acoustic features, where the adjacent acoustic features include all acoustic features with serial numbers less than or equal to a preset value starting from the acoustic feature closest to the slot acoustic feature among the acoustic features corresponding to the template text included in the matching preset template. For example, the acoustic features corresponding to the template text include i1, i2, i3, i4,..., i 20 , the slot is before the acoustic feature i1, and the preset value is set to 5. Then, starting from the acoustic feature i1, all acoustic features with serial numbers less than or equal to 5 are adjacent acoustic features. Therefore, the adjacent acoustic features include acoustic features i1, i2, i3, i4, i5. Still taking the acoustic features corresponding to the template text including i1, i2, i3, i4,..., i 20 as an example, the slot is after the acoustic feature i 20 , the preset value is set to 5. Then, starting from the acoustic feature i 20 , all acoustic features with serial numbers less than or equal to 5 are adjacent acoustic features. However, at this time, the serial numbers need to be counted in reverse order. Therefore, the adjacent acoustic features include acoustic features i 20 , i 19 , i 18 , i 17 , i 16 . Still taking the acoustic features corresponding to the template text including i1, i2, i3, i4,..., i 20 as an example, the slot is between the acoustic features i 11 and i 12 , the preset value is set to 5. Then, starting from the acoustic feature i 11 in reverse order and starting from the acoustic feature i 12 in sequential order, all acoustic features with serial numbers less than or equal to 5 are adjacent acoustic features. Therefore, the adjacent acoustic features include acoustic features i 11 , i 10 , i9, i8, i7 and acoustic features i 12 , i 13 , i 14 , i 15 , i 16It should be noted that no matter how many slots there are in a preset template, the above method can be used to obtain adjacent acoustic features. Based on the compensated acoustic features, the speech waveform corresponding to the adjacent acoustic features is regenerated to generate an updated adjacent template text speech waveform. For example, an updated adjacent template text speech waveform is generated according to a streaming speech synthesis algorithm. The slot speech waveform is spliced together with the speech waveform corresponding to the template text included in the matching preset template, that is, the slot speech waveform, the updated adjacent template text speech waveform, and the speech waveform corresponding to the acoustic features remaining after removing the adjacent acoustic features among the acoustic features included in the matching preset template are spliced together.

[0043] Optionally, in the embodiments of the present invention, a compensation signal may be generated, and the adjacent acoustic features are compensated based on the compensation signal.

[0044] The compensation signal is the difference between the acoustic features generated by the model and the ideal acoustic features. For example, in the embodiments of the present invention, when compensating the adjacent acoustic features, the ideal acoustic features are the compensated acoustic features. Therefore, the compensation signal is the difference between the adjacent acoustic features before compensation and the adjacent acoustic features after compensation. The pronunciation of a speech segment is affected by the content before and after it, which helps the continuity and consistency of the entire speech. For the same preset template, the content filled in its slots is different, and the ideal pronunciation of the template is also different, especially for some speech frames adjacent to the slots. The fundamental reason for this difference is that the encoder part in the model for generating acoustic features uses the linguistic features of the complete sentence. When the slot content is replaced, these linguistic features have changed. The goal of the compensation signal is to repair the impact of this change on pronunciation. To put it more straightforwardly: When generating the preset template, it is based on a certain condition (i.e., a piece of text), and when using the preset template, this condition has changed a little (a small part of the text has changed), and we have to make a little change to the generated template so that it can better adapt to the new condition, and the amplitude of the change is the compensation signal.

[0045] For a trained acoustic model, if the preset template in the technical solution provided in the embodiments of the present invention is not used, that is, the acoustic features generated from scratch can be considered as ideal acoustic features, as Figure 3 shown in the acoustic feature A2 and acoustic feature B2 in the figure. In Figure 3Among them, the acoustic features of two sentences (the italic underlined part is the slot content) are respectively generated by the trained acoustic model. The acoustic feature A2 and the acoustic feature B2 in the figure both represent the acoustic features adjacent to the slot on the preset template, that is, the adjacent acoustic features described in the above embodiments. The difference between the two acoustic features is the compensation signal (denoted by C), where the compensation signal C is the compensation signal for compensating the acoustic feature B2 to obtain the acoustic feature A2. In order to predict the compensation signal C, it is necessary to train a compensation model with the acoustic feature B2 and the linguistic feature A1 as input data, as Figure 4 shown. In the embodiment of the present invention, the following content can be used to train the compensation model. The input data is [acoustic feature B2, linguistic feature A1], and the linguistic feature A1 refers to the linguistic feature corresponding to the entire sentence "No problem, let's watch Boonie Bears together". The output data is [compensation signal C′]. The loss function is Mel-SD(acoustic feature A2, acoustic feature A′), that is, the Mel spectrum distortion measure between the target acoustic feature A2 and the acoustic feature A′ obtained by compensating the acoustic feature B2. Among them, acoustic feature A′ = acoustic feature B2 + compensation signal C′. The optimization method uses Adam. In order to make full use of data resources, the input data can also be [acoustic feature A2, linguistic feature B1], that is, the acoustic feature B2 is obtained by compensating the acoustic feature A2. At this time, the loss function calculates Mel-SD(acoustic feature B2, acoustic feature B′), that is, two sets of training data can be obtained from any two filled sentences of the same template.

[0046] After the compensation model is trained, the actually predicted compensation signal is C′, that is, the compensation signal used during online real-time speech synthesis. Before online real-time speech synthesis, the adjacent acoustic feature is used as the acoustic feature B2; during online real-time speech synthesis, the linguistic feature of the received text is generated as the linguistic feature A1. After obtaining the compensation signal C′ through the compensation model, it is added to the template acoustic feature B2 to obtain the acoustic feature A′, which is the compensated output feature, as Figure 5 shown. The generation of linguistic features is to convert text into pronunciation marks (such as phonemes and tones) and prosody marks (such as stress and pause). When generating linguistic features online in real time, different contents may be filled in the slots. To ensure the effectiveness of the slot linguistic features, the complete input text needs to be used to generate the complete linguistic features.

[0047] Optionally, in the embodiments of the present invention, different methods may be adopted to generate the slot voice waveform and / or update the adjacent template text voice waveform according to the position of the slot in the preset template. The generation of the voice waveform is to convert the acoustic features into voice signals for output. Specifically, the following content may be referred to. Among them, in the preset template library, the hidden state of the voice waveform corresponding to the template text included in each preset template is pre-obtained, and the voice waveform hidden state is the hidden state of the voice waveform generation network when generating the voice waveform corresponding to the acoustic features. The acoustic features corresponding to the template text included in the matching preset template and the slot acoustic features are placed together according to the relationship between the matching preset template and the slot to obtain the overall acoustic features, and the acoustic features included in the overall acoustic features are sorted in the order from the head to the tail. For example, the acoustic features corresponding to the template text include i1, i2, i3, i4, …, i 20 , the slot acoustic features include j1, j2, j3, j4, and the slot is located at the head of the matching preset template. Therefore, the overall acoustic features are the acoustic features j1, j2, j3, j4, i1, i2, i3, i4, …, i 20 , and then all the acoustic features included in the overall acoustic features are sorted in the order from the head to the tail. When the slot is located at the tail and in the middle of the matching preset template, the sorting is performed according to the above content.

[0048] When the slot is located at the tail of the matching preset template, generating the slot voice waveform and / or updating the adjacent template text voice waveform includes: generating the voice samples corresponding to the slot acoustic features and / or the voice samples corresponding to the adjacent acoustic features according to the following content: generating the i + 1th frame acoustic feature y i+1 corresponding voice sample z j+L , z j+1+L , …, z j+2L-1 is based on the streaming voice synthesis algorithm, with y i+1 and z j , z j+1 , …, z j+L-1 as the input, and s i ’ as the condition for generation, where z j , z j+1 , …, z j+L-1 are the voice samples corresponding to the i-th frame acoustic feature y i , L is the frame length, and s i ’ is the voice waveform hidden state corresponding to the i + 1th frame acoustic feature; synthesizing the slot voice waveform and / or updating the adjacent template text voice waveform according to the voice samples corresponding to the slot acoustic features and / or the voice samples corresponding to the adjacent acoustic features. When the slot is located at the tail of the preset template, the voice samples are generated frame by frame in the order of the acoustic features. For example, in the order of the acoustic features y1, y2, …, y i , …, yN-1 and y N in sequence and use them as inputs to generate speech samples z1, z2, …, z j , z j+1 , …, z j+L , …z M , where z j , z j+1 , …, z j+L-1 corresponds to the acoustic feature y i , where L is the frame length and M is the length of the complete speech. During the generation process, a set of hidden states s0’, s1’, s2’, s3’, …, s i ’, …, s N-1 ’ and s N ’ are maintained, where s0’ is the initial state and s i ’ is the hidden state of the speech waveform corresponding to the (i + 1)-th frame acoustic feature, that is, s i ’ is the hidden state of the speech waveform generation network when generating the speech waveform corresponding to the (i + 1)-th frame acoustic feature. When generating the speech waveform corresponding to the (i + 1)-th frame acoustic feature, use y i+1 and z j , z j+1 , …, z j+L-1 as inputs and s i ’ as the condition to generate z j+L , z j+1+L , …, z j+2L-1 . Specifically, in the embodiments of the present invention, it may be from the adjacent acoustic feature to the end of the slot acoustic feature, and the speech samples are generated frame by frame according to the above content. In order to maintain the natural continuity of the speech, the adjacent acoustic features are adjusted. When generating the speech waveform, it is necessary to regenerate the speech waveform corresponding to the adjacent acoustic feature. Specifically, it may be to generate speech samples frame by frame from the adjacent acoustic feature to the slot acoustic feature.

[0049] When the slot is located at the head of the matching preset template, generating the slot speech waveform and / or updating the adjacent template text speech waveform includes: generating the speech samples corresponding to the slot acoustic feature and / or the speech samples corresponding to the adjacent acoustic feature according to the following content: generating the speech samples z i-1 corresponding to the (i - 1)-th frame acoustic feature y j-1 , z j-2 , …, z j-L is based on the streaming speech synthesis algorithm, using y i-1 and z j+L-1 , z j+L-2 , …, z j as inputs and s i ” as the condition for generation, where z j+L-1 , z j+L-2 , …, z jis the acoustic feature y of the i-th frame i The corresponding speech sample points, L is the frame length, s i ” is the hidden state of the speech waveform corresponding to the acoustic feature of the (i - 1)-th frame, that is, s i ” is the hidden state of the speech waveform generation network when generating the speech waveform corresponding to the acoustic feature of the (i - 1)-th frame; according to the speech sample points corresponding to the slot acoustic features and / or the speech sample points corresponding to the adjacent acoustic features, synthesize the slot speech waveform and / or update the adjacent template text speech waveform. When the slot is at the head of the preset template, generate the speech sample points frame by frame in the reverse order of the order of the acoustic features. For example, it takes the acoustic features y N 、y N-1 、…、y i 、…、y2、y1 as inputs, and generate the speech sample points z M 、z M-1 、…、z j+L 、z j+L-1 、…、z j 、…、z1 frame by frame or point by point, where z j+L-1 、…、z j+1 、z j corresponds to the acoustic feature y i , L is the frame length, and M is the complete speech length. A group of hidden states s N+1 ”、s N ”、s N-1 ”、…、s i ”、…、s3”、s2”、s1” will be maintained during the generation process, where s N+1 ” is the initial state, s i ” is the hidden state of the speech waveform corresponding to the (i - 1)-th frame acoustic feature, that is, s i ” is the hidden state of the speech waveform generation network when generating the speech waveform corresponding to the (i - 1)-th frame acoustic feature. When generating the speech waveform corresponding to the (i - 1)-th frame acoustic feature, use y i-1 and z j+L-1 、…、z j as inputs, and s i ” is used as the condition to generate z j-1 、z j-2 、…、z j-L . Specifically, in the embodiments of the present invention, it can be from the adjacent acoustic features to the end of the slot acoustic features, and generate the speech sample points frame by frame according to the above content in sequence. In order to maintain the natural continuity of the speech, the adjacent acoustic features are adjusted. When generating the speech waveform, it is necessary to regenerate the speech waveform corresponding to the adjacent acoustic features, specifically, it can be to generate the speech sample points frame by frame from the adjacent acoustic features to the slot acoustic features.

[0050] When the slot is in the middle of the preset template that is matched, when generating the slot voice waveform and updating the adjacent template text voice waveform, a spliced voice waveform obtained by splicing the slot voice waveform and the updated adjacent template text voice waveform together. When the slot is in the middle of the preset template that is matched, the acoustic feature distributions included in the adjacent acoustic features are on both sides of the slot acoustic feature. The adjacent acoustic features and the slot acoustic feature are put together to form a spliced acoustic feature, and the spliced voice waveform corresponding to the spliced acoustic feature is generated. Specifically, the following contents may be included. Generate the voice samples corresponding to the slot acoustic feature and the acoustic features in the adjacent acoustic features whose serial numbers are less than the serial number of the slot acoustic feature to obtain the first spliced voice samples: Generate the acoustic feature y of the (i + 1)-th frame i+1 The corresponding voice sample z j+L , z j+1+L , …, z j+2L-1 is based on the streaming speech synthesis algorithm, with y i+1 and z j , z j+1 , …, z j+L-1 as the input, and s i ’ as the condition for generation, where z j , z j+1 , …, z j+L-1 are the voice samples corresponding to the acoustic feature y of the i-th frame i , L is the frame length, and s i ’ is the voice waveform hidden state corresponding to generating the acoustic feature of the (i + 1)-th frame. For example, the acoustic features corresponding to the template text include i1, i2, i3, i4, ……, i 20 , the slot acoustic feature includes j1, j2, j3, j4, the slot is between the acoustic features i 11 and i 12 , and the adjacent acoustic features include the acoustic features i 11 , i 10 , i9, i8, i7 and the acoustic features i 12 , i 13 , i 14 , i 15 , i 16 . The voice samples corresponding to the acoustic features i7, i8, i9, i 10 , i 11 , j1, j2, j3, j4 are the first spliced voice samples. Generate the voice samples corresponding to the slot acoustic feature and the acoustic features in the adjacent acoustic features whose serial numbers are greater than the serial number of the slot acoustic feature to obtain the second spliced voice samples: Generate the voice sample z i-1 corresponding to the acoustic feature y of the (i - 1)-th frame j-1 , z j-2 , …, zj-L Based on the streaming speech synthesis algorithm, with y i-1 and z j+L-1 、z j+L-2 、…、z j as inputs, and generating with s i ” as the condition, where z j+L-1 、z j+L-2 、…、z j are the speech sample points corresponding to the i-th frame acoustic feature y i , L is the frame length, and s i ” is the speech waveform hidden state corresponding to the generation of the (i - 1)-th frame acoustic feature. For example, the acoustic features corresponding to the template text include i1, i2, i3, i4, ……, i 20 , the slot acoustic features include j1, j2, j3, j4, and the slot is located between acoustic features i 11 and i 12 . The adjacent acoustic features include acoustic features i 11 、i 10 、i9, i8, i7 and acoustic features i 12 、i 13 、i 14 、i 15 、i 16 . The acoustic features j1, j2, j3, j4, i 12 、i 13 、i 14 、i 15 、i 16The corresponding speech sample point is the second spliced speech sample point. Based on the first spliced speech sample point, the second spliced speech sample point, the weight corresponding to the first spliced speech sample point, and the weight corresponding to the second spliced speech sample point, the slot acoustic feature and the spliced speech sample point corresponding to the adjacent acoustic feature are generated. Specifically, the weight is a weight sequence. Multiply the first spliced speech sample point by the elements included in its corresponding weight sequence, multiply the second spliced speech sample point by the elements included in its corresponding weight sequence, and add the results of the two multiplications to obtain the spliced speech sample point. Among them, the number of elements included in the weight sequence corresponding to the first spliced speech sample point is the same as the number of elements included in the weight sequence corresponding to the second spliced speech sample point, and the number of elements included in the weight sequence is the number of speech sample points included in the first spliced speech sample point or the second spliced speech sample point. According to the spliced speech sample point, the spliced speech waveform is synthesized. When the slot is located in the middle of the preset template, it is equivalent to dividing it into two cases: the slot is located at the end of the preset template and the slot is located at the end of the preset template. According to the sorting of the acoustic features, for the acoustic feature located at the head of the slot acoustic feature among the adjacent acoustic features, the first spliced speech sample point is generated by the method of generating speech sample points in the above order; for the acoustic feature located at the tail of the slot acoustic feature among the adjacent acoustic features, the second spliced speech sample point is generated by the method of generating speech sample points in the reverse order of the above. The final spliced speech sample point with the adjacent acoustic feature and the slot acoustic feature spliced together is obtained by weighted summation. For example, the length of the first spliced speech sample point or the second spliced speech sample point is N, and the weight sequence corresponding to the first spliced speech sample point can be The weight sequence corresponding to the second spliced speech sample point can be

[0051] The generation of acoustic features is to convert linguistic features into acoustic features (such as Mel spectrum coefficients). In the embodiments of the present invention, the acoustic features corresponding to the template text in the preset template are pre-generated, and the generated acoustic features and the hidden states during the generation of the acoustic features are saved for later use. The acoustic features and hidden states of the template text are required when generating the slot acoustic features. Optionally, in the embodiments of the present invention, different methods can be used to generate the slot acoustic features according to the different positions of the slot in the preset template. Among them, in the preset template library, the acoustic features and acoustic feature hidden states corresponding to the template text included in each preset template are pre-obtained, and the acoustic feature hidden state is the hidden state corresponding to the generation of the acoustic features corresponding to the template text. The acoustic features corresponding to the template text included in the matching preset template and the slot acoustic features are placed together according to the relationship between the matching preset template and the slot to obtain the overall acoustic features, and the acoustic features included in the overall acoustic features are sorted in the order from the head to the tail. For the relevant introduction about the sorting, reference can be made to the relevant content described in the above embodiments.

[0052] When the slot is at the end of the matching preset template, for example, in the example "Today is the [April 17th] of the Xin Chou year in the lunar calendar", where "Today is the Xin Chou year in the lunar calendar" is the template text and "April 17th" is the slot text. When the slot is at the end of the preset template, based on the streaming speech synthesis algorithm, the acoustic features are generated frame by frame according to the sorting order for acoustic features. Among them, generating the acoustic features sequentially takes the linguistic features X of a complete text as the input, and generates the acoustic features y1, y2,..., y i 、…、y N-1 、y N 。During the generation process, a set of hidden states s0, s1, s2,..., s i 、…、s N-1 、s N will be maintained, where s0 is the initial state, and this hidden state is the hidden state of the acoustic feature generation network when generating acoustic features. Specifically, generating the slot acoustic features includes generating according to the following: The acoustic feature y i+1 of the (i + 1)-th frame is generated based on the streaming speech synthesis algorithm, with Enc(X) and y i as the input and s i as the condition. Among them, Enc() is the linguistic feature encoding network (such as the CBHG-based Encoder of Tacotron and the Transformer-based Encoder of Fastspeech), s i is the acoustic feature hidden state corresponding to the acoustic feature of the i-th frame, that is, s i is the hidden state of the acoustic feature generation network when generating the acoustic feature of the i-th frame, y i is the acoustic feature of the i-th frame, and X is the linguistic feature corresponding to the received text. In addition, in the embodiments of the present invention, adjacent acoustic features can be compensated. When compensating adjacent acoustic features, the obtained acoustic features corresponding to the received text can be referred to as shown in Figure 6 ,where the acoustic feature compensation signal generated by the adjacent part of the template and the slot shown in Figure 6 is the compensation signal for compensating adjacent acoustic features.

[0053] When the slot is at the head of the matching preset template, for example, in the sample: "Jiangnan" one-minute audition version is for you. You can purchase a membership on the Xiaohai Butler app to enjoy the full version of the song. Here, "Jiangnan" is the slot text, and "one-minute audition version is for you. You can purchase a membership on the Xiaohai Butler app to enjoy the full version of the song" is the template text. The difference from the mainstream streaming speech synthesis algorithm solution is that in the present invention, for this situation, the acoustic features are generated in reverse order, that is, the acoustic features of the character "Nan" are generated first, and then the acoustic features of the character "Jiang" are generated. When the slot is at the head of the preset template, based on the streaming speech synthesis algorithm, the acoustic features are generated frame by frame in reverse order according to the sorting of the acoustic features. Among them, generating the acoustic features frame by frame in reverse order takes the linguistic feature X of a complete text as the input and generates the acoustic feature y frame by frame in reverse order from the back to the front N 、y N-1 、…、y j 、…、y2、y1. A group of hidden states s N+1 、s N 、s N-1 、…、s j 、…、s2、s1 will be maintained during the generation process, where s N+1 is the initial state, and this hidden state is the hidden state corresponding to the generation of the acoustic features. When generating the (j - 1)-th frame of acoustic features, Enc(X) and y j are used as the input, and s j is used as the condition to generate y j-1 , where Enc() is the linguistic feature encoding network. Specifically, when the slot is at the head of the preset template, generating the slot acoustic features includes generating according to the following: The (j - 1)-th frame of acoustic feature y j-1 is generated based on the streaming speech synthesis algorithm with Enc(X) and y j as the input and s j as the condition, where Enc() is the linguistic feature encoding network, s j is the acoustic feature hidden state corresponding to the (j - 1)-th frame of acoustic feature, that is, s j is the hidden state of the acoustic feature generation network when generating the (j - 1)-th frame of acoustic feature, y j is the j-th frame of acoustic feature, and X is the linguistic feature corresponding to the received text. In addition, in the embodiments of the present invention, adjacent acoustic features can be compensated. When compensating adjacent acoustic features (for example, it can be compensated by using the method for compensating adjacent acoustic features provided in the embodiments of the present invention), the slot acoustic features and adjacent acoustic features can be understood with reference to Figure 7 .

[0054] When the slot is in the middle of the matching preset template, which is different from the cases where the slot is at the head and tail of the preset template, in this case, the slot part will be affected by the preset template parts at both of its ends simultaneously. Specifically, in the embodiments of the present invention, it may be based on a streaming speech synthesis algorithm to generate acoustic features frame by frame according to the sorting order for acoustic features. Specifically, generating the slot acoustic features includes generating according to the following: the acoustic feature y of the (i + 1)-th frame i+1 is based on the streaming speech synthesis algorithm, with Enc(X), y i 、y n and Enc d (n - i - 1) as the input and s i as the condition for generation, where Enc() is the linguistic feature encoding network, Enc d () is the distance encoding function, s i is the acoustic feature hidden state corresponding to the acoustic feature of the (i + 1)-th frame, y i is the acoustic feature of the i-th frame, y n is the acoustic feature of the n-th frame and is the acoustic feature with the serial number greater than the serial number of the slot acoustic feature and closest to the slot acoustic feature among adjacent acoustic features, X is the linguistic feature corresponding to the received text. For example, the adjacent acoustic features include y a-l+1 、…、y a 、y b 、y b+1 、…、y b+l-1 , l is the adjusted window length (that is, the preset value described in the embodiments of the present invention). The one between y a and y b is the slot acoustic feature, and y b is exactly the acoustic feature with the serial number greater than the serial number of the slot acoustic feature and closest to the slot acoustic feature in the sorting of acoustic features.

[0055] In summary, in the embodiments of the present invention, the computational complexity of the TTS algorithm is reduced from the algorithm structure, thereby reducing the response delay of the intelligent device terminal; from the characteristics of service data and algorithm strategies, the task processing time is reduced, thereby reducing computing resources and lowering the service cost.

[0056] Another aspect of the embodiments of the present invention further provides a machine-readable storage medium, on which a program is stored, and when the program is executed by a processor, the method described in the above embodiments is implemented.

[0057] Another aspect of the embodiments of the present invention further provides a processor, and the processor is used to run a program, wherein when the program runs, the method described in the above embodiments is executed.

[0058] Another aspect of the embodiments of the present invention further provides a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the methods described in the above embodiments are implemented. The device herein can be a server, a PC, a PAD, a mobile phone, etc.

[0059] Another aspect of the embodiments of the present invention further provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the methods described in the above embodiments are implemented.

[0060] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0061] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0062] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0063] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1Steps for the functions specified in one or more boxes.

[0064] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0065] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0066] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0067] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0068] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for synthesizing speech, characterized in that, The method includes: finding a matching preset template that matches the received text in a preset template library, wherein in the preset template library, each preset template includes template text and slots, and the speech waveform corresponding to the template text has been pre-acquired, and the acoustic features corresponding to the template text included in each preset template have been pre-acquired; acquiring slot contents corresponding to the slots included in the matching preset template in the received text, including comparing the received text with the matching preset template to obtain the slot contents corresponding to the slots; generating slot acoustic features corresponding to the slot contents; generating slot speech waveforms corresponding to the slot acoustic features; and stitching together the slot speech waveforms and the speech waveforms corresponding to the template text included in the matching preset template to obtain a speech waveform corresponding to the received text, thereby synthesizing speech, The method further includes: compensating adjacent acoustic features in the acoustic features corresponding to the template text included in the matching preset template to obtain compensated acoustic features, wherein the adjacent acoustic features include all the acoustic features whose serial numbers are less than or equal to a preset value starting from the acoustic feature closest to the slot acoustic feature among the acoustic features corresponding to the template text included in the matching preset template; and regenerating the speech waveforms corresponding to the adjacent acoustic features based on the compensated acoustic features to generate updated adjacent template text speech waveforms; wherein the stitching together of the slot speech waveforms and the speech waveforms corresponding to the template text included in the matching preset template is to stitch together the slot speech waveforms, the updated adjacent template text speech waveforms, and the speech waveforms corresponding to the remaining acoustic features except the adjacent acoustic features in the acoustic features corresponding to the template text included in the matching preset template.

2. The method according to claim 1, characterized in that, The finding a matching preset template that matches the received text in the preset template library includes: sorting all the characters included in the received text in the order from the head to the tail of the received text; and performing the following matching operations on any one of the preset templates in the preset template library: segmenting the template text based on the slots to obtain segmented template text; sorting the segmented template text in the order from the head to the tail of the preset template according to the position of the segmented template text in the preset template; sequentially searching for the segmented template text in the received text according to the sorting of the template text, wherein the characters in the received text that are the same as the characters included in a segmented template text are no longer used for subsequent searches; when all the segmented template text is found in the received text, determining whether the preset template matches the received text successfully according to a matching preset condition; and When the preset template successfully matches the received text, determine whether to set the preset template as the optimal template according to the optimal template setting rules. Among them, when all the preset templates in the preset template library have completed the matching operation, the optimal template is the matching preset template. Among them, the optimal template setting rules include: If the optimal template has not been set, set the preset template as the optimal template; If the optimal template has been set, but the number of characters included in the preset template is more than the number of characters included in the optimal template, set the preset template as the optimal template; and If the optimal template has been set and the number of characters included in the preset template is the same as the number of characters included in the optimal template, but the number of slots included in the preset template is less than the number of slots included in the optimal template, set the preset template as the optimal template.

3. The method according to claim 2, wherein Search for the segmented template text in the received text in sequence according to the template text sorting, including: Search for the first segmented template text in the received text according to the following content. Among them, the first segmented template text is the segmented template text that is ranked first according to the template text sorting in the preset template: When there is no slot at the head of the first segmented template text, compare the first segmented template text with the string in the received text that has the same character length as the first segmented template text starting from the first character according to the character sorting. Among them, the first character is the character that is ranked first according to the character sorting in the received text; and When there is a slot at the head of the first segmented template text, compare the first segmented template text with any string that has the same character length as the first segmented template text and has consecutive character sorting serial numbers in the received text; When the first segmented template text is found in the received text, set the matching position for the second segmented template text according to the character length of the first segmented template text and the serial number of the starting character of the string that matches the first segmented template text in the received text according to the character sorting. Among them, the second segmented template text is the next segmented template text of the first segmented template text according to the template text sorting, and the matching position is the serial number of the starting character of the segmented template text searched in the received text according to the character sorting; Start from the matching position and search for the second segmented template text in the received text; and In the case where the second segmented template text is found, update the matching position according to the character length of the second segmented template text and the serial number of the starting character in the string matching the second segmented template text in the received text in the character sorting order, and sequentially search for the segmented template text other than the first segmented template text and the second segmented template text in the received text according to the content of the second segmented template text until all the segmented template texts in the preset template are searched. Among them, in the process of searching for all the segmented template texts including the first segmented template text, the second segmented template text and those after them in the preset template, if any of the segmented template texts is not found, the preset template fails to match the received text, and the operation of sequentially searching for the segmented template texts in the received text according to the template text sorting ends.

4. The method according to claim 2, wherein The matching preset conditions include: If all the segmented template texts are found in the received text and the maximum serial number of the characters in the received text in the character sorting order is greater than or equal to the last matching position corresponding to the last segmented template text and there is no such slot at the tail of the last segmented template text, the preset template fails to match the received text successfully. Among them, the last segmented template text is the segmented template text ranked last in the preset template in the template text sorting order, and the last matching position is the matching position set according to the character length of the last segmented template text and the serial number of the starting character in the string matching the last segmented template text in the received text in the character sorting order; and If all the segmented template texts are found in the received text and the maximum serial number is less than the last matching position corresponding to the last segmented template text and / or there is such a slot at the tail of the last segmented template text, the preset template matches the received text successfully.

5. The method according to claim 1, characterized in that, Generating the slot voice waveform and / or updating the adjacent template text voice waveform is generated according to the streaming speech synthesis algorithm.

6. The method according to claim 5, characterized in that, In the preset template library, the voice waveform hidden state corresponding to the template text included in each preset template is pre-acquired. The voice waveform hidden state is the hidden state of the voice waveform generation network when generating the voice waveform corresponding to the acoustic feature. When the slot is located at the tail of the matching preset template, the acoustic feature corresponding to the template text included in the matching preset template and the slot acoustic feature are put together according to the relationship between the matching preset template and the slot to obtain the overall acoustic feature, and the acoustic features included in the overall acoustic feature are sorted in the order from the head to the tail. Generating the slot voice waveform and / or updating the adjacent template text voice waveform includes: Generate the speech sample corresponding to the slot acoustic feature and / or the speech sample corresponding to the adjacent acoustic feature according to the following content: Generate the acoustic feature y of the i+1th frame i+1 The corresponding speech sample z j+L 、z j+1+L ,…,z j+2L-1 Based on the streaming speech synthesis algorithm, y i+1 and z j 、z j+1 ,…,z j+L-1 As input, s i 'Generate conditions, where z j 、z j+1 ,…,z j+L-1 is the acoustic feature y of the i-th frame i The corresponding speech sample point, L is the frame length, s i ' is the speech waveform hidden state corresponding to the acoustic feature of the i+1th frame; and Synthesize the slot voice waveform and / or the updated adjacent template text voice waveform according to the voice sample points corresponding to the slot acoustic features and / or the voice sample points corresponding to the adjacent acoustic features.

7. The method according to claim 5, characterized in that, In the preset template library, the voice waveform hidden state corresponding to the template text included in each preset template is pre-obtained. The voice waveform hidden state is the hidden state of the voice waveform generation network when generating the voice waveform corresponding to the acoustic feature. When the slot is located at the head of the matching preset template, the acoustic feature corresponding to the template text included in the matching preset template and the slot acoustic feature are put together according to the relationship between the matching preset template and the slot to obtain an overall acoustic feature, and the acoustic features included in the overall acoustic feature are sorted in the order from the head to the tail. Generating the slot voice waveform and / or the updated adjacent template text voice waveform includes: Generate the speech sample points corresponding to the slot acoustic features and / or the speech sample points corresponding to the adjacent acoustic features according to the following content: Generate the acoustic feature y of the (i-1)th frame i-1 The corresponding speech sample point z j-1 , z j-2 , …, z j-L is based on the streaming speech synthesis algorithm, with y i-1 and z j+L-1 , z j+L-2 , …, z j as the input, and s i ’’ as the condition for generation, where z j+L-1 , z j+L-2 , …, z j are the speech sample points corresponding to the acoustic feature y of the ith frame i L is the frame length, and s i ’’ is the hidden state of the speech waveform corresponding to the acoustic feature of the (i-1)th frame; and Synthesize the slot voice waveform and / or the updated adjacent template text voice waveform according to the voice sample points corresponding to the slot acoustic features and / or the voice sample points corresponding to the adjacent acoustic features.

8. The method according to claim 5, wherein In the preset template library, the voice waveform hidden state corresponding to the template text included in each preset template is pre-obtained. The voice waveform hidden state is the hidden state of the voice waveform generation network when generating the voice waveform corresponding to the acoustic feature. When the slot is located in the middle of the matching preset template, the acoustic feature corresponding to the template text included in the matching preset template and the slot acoustic feature are put together according to the relationship between the matching preset template and the slot to obtain an overall acoustic feature, and the acoustic features included in the overall acoustic feature are sorted in the order from the head to the tail. Generating the slot voice waveform and the updated adjacent template text voice waveform is a spliced voice waveform obtained by splicing the generated slot voice waveform and the updated adjacent template text voice waveform. Generating the spliced voice waveform includes: Generate the speech sample corresponding to the slot acoustic feature and the speech sample corresponding to the acoustic feature whose sequence number is smaller than the sequence number of the slot acoustic feature in the adjacent acoustic features to obtain the first spliced speech sample according to the following content: Generate the acoustic feature y of the i+1th frame i+1 The corresponding speech sample z j+L 、z j+1+L ,…,z j+2L-1 Based on the streaming speech synthesis algorithm, y i+1 and z j 、z j+1 ,…,z j+L-1 As input, s i 'Generate conditions, where z j 、z j+1 ,…,z j+L-1 is the acoustic feature y of the i-th frame i The corresponding speech sample point, L is the frame length, s i 'To generate the speech waveform hidden state corresponding to the acoustic feature of the i+1th frame; Generate the speech samples corresponding to the slot acoustic features and the speech samples corresponding to the acoustic features in the adjacent acoustic features whose sequence numbers are greater than the sequence number of the slot acoustic features to obtain the second spliced speech samples: Generate the acoustic feature y of the (i-1)th frame i-1 The corresponding speech sample z j-1 , z j-2 , …, z j-L Based on the streaming speech synthesis algorithm, with y i-1 and z j+L-1 , z j+L-2 , …, z j as the input, and with s i ’’ as the condition for generation, where z j+L-1 , z j+L-2 , …, z j are the speech samples corresponding to the acoustic feature y of the ith frame i , L is the frame length, and s i ’’ is the hidden state of the speech waveform corresponding to the acoustic feature of the (i-1)th frame; Generating the spliced voice sample points corresponding to the slot acoustic feature and the adjacent acoustic feature based on the first spliced voice sample points, the second spliced voice sample points, the weight corresponding to the first spliced voice sample points, and the weight corresponding to the second spliced voice sample points; and Synthesize the spliced voice waveform according to the spliced voice sample points.

9. The method according to any one of claims 1-4, characterized in that, Generating the slot acoustic feature is generated according to the streaming voice synthesis algorithm.

10. The method according to claim 9, characterized in that, In the preset template library, the acoustic features and the acoustic feature hidden states corresponding to the template texts included in each preset template are pre-acquired. The acoustic feature hidden state is the hidden state of the acoustic feature generation network when generating the acoustic features. When the slot is located at the tail of the matching preset template, the acoustic features corresponding to the template texts included in the matching preset template and the slot acoustic features are put together according to the relationship between the matching preset template and the slot to obtain the overall acoustic features, and the acoustic features included in the overall acoustic features are sorted in the order from the head to the tail. Generating the slot acoustic features includes generating according to the following content: The acoustic feature y of the (i + 1)-th frame i+1 is generated based on the streaming speech synthesis algorithm with Enc(X) and y i as inputs and s i as a condition, where Enc() is the linguistic feature encoding network, and s i is the acoustic feature hidden state corresponding to the acoustic feature of the (i + 1)-th frame, y i is the acoustic feature of the i-th frame, and X is the linguistic feature corresponding to the received text.

11. The method according to claim 9, wherein In the preset template library, the acoustic features and the acoustic feature hidden states corresponding to the template texts included in each preset template are pre-acquired. The acoustic feature hidden state is the hidden state of the acoustic feature generation network when generating the acoustic features. When the slot is located at the head of the matching preset template, the acoustic features corresponding to the template texts included in the matching preset template and the slot acoustic features are put together according to the relationship between the matching preset template and the slot to obtain the overall acoustic features, and the acoustic features included in the overall acoustic features are sorted in the order from the head to the tail. Generating the slot acoustic features includes generating according to the following content: The acoustic feature y of the (j-1)-th frame j-1 Based on the streaming speech synthesis algorithm, using Enc(X) and y j as inputs, and s j as a condition for generation, where Enc() is the linguistic feature encoding network, and s j is the acoustic feature hidden state corresponding to the acoustic feature of the (j-1)-th frame, and y j is the acoustic feature of the j-th frame, and X is the linguistic feature corresponding to the received text.

12. The method according to claim 9, wherein In the preset template library, the acoustic features and the acoustic feature hidden states corresponding to the template texts included in each preset template are pre-acquired. The acoustic feature hidden state is the hidden state of the acoustic feature generation network when generating the acoustic features. When the slot is located in the middle of the matching preset template, the acoustic features corresponding to the template texts included in the matching preset template and the slot acoustic features are put together according to the relationship between the matching preset template and the slot to obtain the overall acoustic features, and the acoustic features included in the overall acoustic features are sorted in the order from the head to the tail. Generating the slot acoustic features includes generating according to the following content: The acoustic feature y of the (i + 1)-th frame i+1 Based on the streaming speech synthesis algorithm, with Enc(X), y i , y n and Enc d (n - i - 1) as the input, and s i as the condition for generation, where Enc() is the linguistic feature encoding network, Enc d () is the distance encoding function, s i is the acoustic feature hidden state corresponding to the acoustic feature of the (i + 1)-th frame, y i is the acoustic feature of the i-th frame, y n is the acoustic feature of the n-th frame and is the acoustic feature with the serial number greater than that of the slot acoustic feature among the adjacent acoustic features and closest to the slot acoustic feature, and X is the linguistic feature corresponding to the received text.

Citation Information

Patent Citations

  • Voice role switching method, device and equipment and storage medium

    CN107340991A

  • Voice information generation method and device, electronic equipment and storage medium

    CN111128121A

  • Method for improving template sentence synthetic effect in voice synthetic system

    CN1945691A