Voice processing method, device, equipment and storage medium

By adding acoustic processing padding blocks to streaming speech synthesis and performing streaming acoustic processing, the problem of long first-packet response time in streaming speech synthesis is solved, and speech processing efficiency and user experience are improved.

CN114882866BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210524046.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-09-23
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

Existing streaming speech synthesis solutions have the problem of long first-packet response time, and some solutions fail to effectively support streaming processing, resulting in a reduced user experience.

Method used

By adding an acoustic processing padding block to the acoustic processing data block and processing it using the streaming acoustic processing module, a streaming acoustic processing result is generated, and then streaming or non-streaming speech synthesis is performed to ensure that the data block size is consistent to support subsequent audio output.

Benefits of technology

It improves voice processing efficiency, reduces audio response time, especially the first packet response time, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882866B_ABST
    Figure CN114882866B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speech processing method, apparatus, device and storage medium, which relate to the field of artificial intelligence, and in particular to the field of speech technology. The specific implementation scheme is: based on the parameter characteristics of the streaming acoustic processing module in the acoustic model, an acoustic processing filling block is obtained; the acoustic processing filling block is added to the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; wherein the i-th data block to be acoustically processed is obtained by slicing the data to be processed into n data blocks to be acoustically processed; the i is a natural number not greater than n, and the n is a natural number not less than 2; the i-th target acoustic processing data block is input into the streaming acoustic processing module in the acoustic model for streaming acoustic processing to obtain the i-th streaming acoustic processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to the field of artificial intelligence and voice technology. Background Art

[0002] Streaming text-to-speech (TTS), also known as online speech synthesis, is primarily used to optimize the latency of the first packet of speech synthesis. Specifically, streaming speech synthesis generates synthesized audio in segments and plays it back as it is being synthesized. This significantly reduces the first-packet response time compared to generating the entire audio track all at once (non-streaming speech synthesis). However, existing streaming speech synthesis solutions only partially support streaming processing, and therefore, the problem of long first-packet response time still exists. Summary of the Invention

[0003] The present disclosure provides a speech processing method, apparatus, device, and storage medium.

[0004] According to one aspect of the present disclosure, there is provided a speech processing method, comprising:

[0005] Based on the parameter characteristics of the streaming acoustic processing module in the acoustic model, an acoustic processing filling block is obtained;

[0006] Adding the acoustic processing filling block to the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; wherein the i-th data block to be acoustically processed is obtained by slicing the data to be processed into n data blocks to be acoustically processed; i is a natural number not greater than n, and n is a natural number not less than 2;

[0007] The i-th target acoustic processing data block is input into the streaming acoustic processing module in the acoustic model for streaming acoustic processing to obtain the i-th streaming acoustic processing result.

[0008] According to another aspect of the present disclosure, there is provided a speech processing apparatus, comprising:

[0009] A parameter feature processing unit, configured to obtain an acoustic processing filling block based on parameter features of a streaming acoustic processing module in an acoustic model;

[0010] a data preprocessing unit, configured to add the acoustic processing filling block to the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; wherein the i-th data block to be acoustically processed is obtained by slicing the data to be processed into n data blocks to be acoustically processed; wherein i is a natural number not greater than n, and n is a natural number not less than 2;

[0011] The speech synthesis unit is used to input the i-th target acoustic processing data block into the streaming acoustic processing module in the acoustic model for streaming acoustic processing to obtain the i-th streaming acoustic processing result.

[0012] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method of any embodiment of the present disclosure.

[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to any embodiment of the present disclosure.

[0017] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method according to any embodiment of the present disclosure when executed by a processor.

[0018] In this way, streaming acoustic processing can be performed, which improves voice processing efficiency and reduces audio response time.

[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0021] Figure 1 This is a schematic flow chart of the speech processing method according to the embodiment of the present application. Figure 1 ;

[0022] Figure 2 This is an example diagram of adding an acoustic processing filling block in a specific example of the speech processing method according to an embodiment of the present application;

[0023] Figure 3 This is a schematic flow chart of the speech processing method according to the embodiment of the present application. Figure 2 ;

[0024] Figure 4 This is a schematic flow chart of the speech processing method according to the embodiment of the present application. Figure 3 ;

[0025] Figure 5 This is a schematic flow chart of the speech processing method according to the embodiment of the present application. Figure 4 ;

[0026] Figure 6 This is a schematic flow chart of the speech processing method according to the embodiment of the present application. Figure 5 ;

[0027] Figure 7 This is a schematic flow chart of the speech processing method according to the embodiment of the present application. Figure 6 ;

[0028] Figure 8 This is a schematic flow chart of the speech processing method according to the embodiment of the present application. Figure 7 ;

[0029] Figure 9 This is a schematic flow chart of the speech processing method according to the embodiment of the present application. Figure 8 ;

[0030] Figure 10 This is an example diagram of data block processing in a specific example according to the speech processing method of an embodiment of the present application;

[0031] Figure 11 This is a schematic flow chart of the speech processing method according to the embodiment of the present application. Figure 9 ;

[0032] Figure 12 is a structural diagram of a speech processing device according to an embodiment of the present disclosure;

[0033] Figure 13 is a flowchart of a specific example of a speech processing method according to an embodiment of the present application;

[0034] Figure 14 It is a block diagram of an electronic device used to implement the voice processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0035] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0036] In voice interaction scenarios, the overall response time for the entire interaction chain is required to be less than 1500ms. As the final step in the interaction, speech synthesis generally needs to return data and start voice broadcasting within 300ms. Otherwise, users will experience delays, which will reduce the user experience.

[0037] Here, speech synthesis mainly includes three steps:

[0038] The first step is text processing, which maps text to phonemes. This process can use a text processing model.

[0039] Step 2: Perform acoustic processing on the phonemes obtained in the first step and map them to audio features (ie, obtain audio features). This step may use an acoustic model, such as a streaming acoustic model AM (Acoustic Model, AM).

[0040] Step 3: Map the audio features obtained in step 2 to audio sample points to obtain audio sample points. Here, the audio sample points are determined based on the audio sampling rate, thus providing support for the subsequent output of audio data. This step can use a speech synthesis model, such as a streaming vocoder (Voc).

[0041] Based on this, the present disclosure provides a streaming speech synthesis solution, specifically, Figure 1 This is a schematic flow chart of a speech processing method according to an embodiment of the present application. Figure 1 The method may optionally be applied to electronic devices, such as personal computers or servers, but is not limited thereto. The method includes at least part of the following contents. Figure 1 As shown, including:

[0042] Step S101: obtaining an acoustic processing filling block based on parameter characteristics of the streaming acoustic processing module in the acoustic model.

[0043] Here, the streaming acoustic processing module refers to a module in the acoustic model that supports streaming processing. Correspondingly, the non-streaming acoustic processing module described below refers to a module in the acoustic model that does not support streaming processing.

[0044] It is understandable that the disclosed solution does not limit the specific acoustic model. As long as the model can perform acoustic processing, it can be applied to the disclosed solution and can improve the efficiency of speech processing. Furthermore, the disclosed solution does not limit the number and type of specific modules in the acoustic model used to support streaming acoustic processing. For example, all modules in the acoustic model can support streaming acoustic processing, or only some modules can support streaming acoustic processing, etc., all of which are within the scope of protection of the disclosed solution.

[0045] Step S102: Add the acoustic processing filling block to the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; wherein the i-th data block to be acoustically processed is obtained by slicing the data to be processed into n data blocks to be acoustically processed; i is a natural number not greater than n, and n is a natural number not less than 2.

[0046] That is, the i-th data block to be acoustically processed is any one of the n data blocks to be acoustically processed, and the n data blocks to be acoustically processed are obtained by slicing the data to be processed.

[0047] Step S103: inputting the i-th target acoustic processing data block into the streaming acoustic processing module in the acoustic model for streaming acoustic processing to obtain the i-th streaming acoustic processing result.

[0048] Here, the outputted i-th streaming acoustic processing result may specifically be an audio feature, that is, an audio feature obtained after streaming acoustic processing is performed on the text data, thereby providing data support for subsequent speech synthesis.

[0049] It is understood that the streaming method described in the disclosed solution refers to the real-time return of partial processing results during the data processing process, that is, processing the data to be processed piece by piece and outputting the processing results piece by piece. In contrast, the non-streaming method requires that the entire data to be processed be processed before outputting the overall processing results.

[0050] In this way, the disclosed solution can perform streaming acoustic processing, thereby further improving the efficiency of voice processing and further reducing the audio response time, such as the first packet response time; effectively improving the user experience.

[0051] In a specific example of the disclosed solution, at least one of the following methods is used to add the acoustic processing filling block to the i-th data block to be acoustically processed to obtain the i-th target acoustically processed data block:

[0052] When the value of i is 1, the acoustic processing filling block is added after the first data block to be acoustically processed to obtain the first target acoustic processing data block;

[0053] When i is any value from 2 to n-1, the acoustic processing filling block is added before and after the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block;

[0054] When the value of i is n, the acoustic processing filling block is added in front of the n-th data block to be acoustically processed to obtain the n-th target acoustic processing data block.

[0055] For example, if Figure 2 As shown, when i is 1, the acoustic processing filling block is added after the first data block to be acoustically processed F1 to obtain the first target acoustic processing data block C1; when i is any value from 2 to n-1, the acoustic processing filling block is added before and after the i-th data block to be acoustically processed Fi to obtain the i-th target acoustic processing data block Ci; for example, when i is 2, the acoustic processing filling block is added before and after the second data block to be acoustically processed F2 to obtain the second target acoustic processing data block C2. When i is n, the acoustic processing filling block is added before the n-th data block to be acoustically processed Fn to obtain the n-th target acoustic processing data block Cn. In this way, after splicing the segment results after streaming acoustic processing (such as the i-th streaming acoustic processing result), the total result corresponding to the data to be processed is obtained. At this time, since the data block to be processed is processed based on the acoustic processing filling block, it is possible to ensure that the data block size of the total result is consistent with the size of the total result of the data to be processed output by the non-streaming acoustic processing. This provides support for the subsequent effective speech synthesis and audio data output.

[0056] It can be understood that the above processing method is only a specific example. In actual applications, other processing methods can also be used to ensure that the splicing results obtained after streaming acoustic processing are consistent with the results of non-streaming acoustic processing. The present disclosure does not impose specific restrictions on this.

[0057] In this way, since the acoustic processing filling block is obtained based on the parameter characteristics of the streaming acoustic processing module in the acoustic model, it effectively ensures that the total result obtained after splicing the segment results after streaming acoustic processing (that is, the i-th streaming acoustic processing result) has the same data block size as the total result output by the non-streaming acoustic processing, laying the foundation for subsequent effective speech synthesis and audio data output.

[0058] In a specific example of the present disclosure, Figure 3 As shown, after step S103 is executed, the following steps are further included:

[0059] Step S104: Based on the i-th streaming acoustic processing result, perform streaming speech synthesis to obtain a streaming speech output result; or, based on the i-th streaming acoustic processing result, perform non-streaming speech synthesis to obtain a non-streaming speech output result.

[0060] That is to say, in this example, after performing streaming acoustic processing, you can use streaming speech synthesis to perform streaming speech synthesis on the processing results after streaming acoustic processing. In this way, due to the streaming acoustic processing and streaming speech synthesis processing, the speech processing efficiency is further improved, and the audio response time can be further reduced, such as the first packet response time. Alternatively, you can also use non-streaming speech processing to perform non-streaming processing on the processing results after streaming acoustic processing. In this case, due to the streaming acoustic processing, the speech processing efficiency can also be improved and the audio response time can be reduced; in this way, the user experience is effectively improved.

[0061] Figure 4 This is a schematic flow chart of a speech processing method according to an embodiment of the present application. Figure 2 The method may optionally be applied to electronic devices, such as personal computers or servers, but is not limited thereto. The method includes at least part of the following contents. Figure 2 As shown, including:

[0062] Step S401: Processing text data to obtain a target phoneme sequence corresponding to the text data; the text data is sub-text data obtained by segmenting the total text data.

[0063] That is to say, the text data is obtained after segmenting the total text data. For example, the total text data is segmented based on semantics to obtain the text data. In this way, on the one hand, the naturalness of the subsequently output audio is guaranteed, and on the other hand, the quality of the subsequently obtained audio features is guaranteed, thereby ensuring the quality of the generated audio data and effectively improving the user experience.

[0064] Step S402: obtaining the data to be processed based on the target phoneme sequence.

[0065] Furthermore, in a specific example, after obtaining the data to be processed, the data to be processed is sliced ​​to obtain n data blocks to be acoustically processed, where n is a natural number not less than 2. In this way, the data blocks to be acoustically processed are processed one by one to achieve streaming acoustic processing.

[0066] It is worth noting that the size of the data block to be acoustically processed can be set based on actual needs, and the present disclosure does not impose any restrictions on this.

[0067] Step S403: Obtain an acoustic processing filling block based on the parameter characteristics of the streaming acoustic processing module in the acoustic model.

[0068] Step S404: Add the acoustic processing filling block to the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; wherein the i-th data block to be acoustically processed is obtained by slicing the data to be processed into n data blocks to be acoustically processed; and i is a natural number not greater than n.

[0069] It can be understood that the execution order of step S201 and step S202 and step S203 is not specifically limited. For example, step S203 can be executed first, and then step S201 and step S202 are executed. As long as the processing result can be obtained before step S204, it is within the protection scope of the present disclosure.

[0070] Step S405: inputting the i-th target acoustic processing data block into the streaming acoustic processing module in the acoustic model for streaming acoustic processing to obtain the i-th streaming acoustic processing result.

[0071] In this way, since the total text data can be segmented in advance to obtain text data, and then the text data can be processed to obtain the target phoneme sequence, this lays the foundation for improving the efficiency of speech processing.

[0072] In a specific example of the disclosed solution, the target phoneme sequence can be obtained in the following manner. Specifically, the above-mentioned processing of text data to obtain the target phoneme sequence corresponding to the text data specifically includes: inputting the text data into a text processing model to obtain an initial phoneme sequence; processing the initial phoneme sequence based on preset phoneme rules to obtain the target phoneme sequence. For example, a customized phoneme rule can be made based on actual needs, and then after obtaining the initial phoneme sequence, it can be modified based on the customized phoneme rule. In this way, the quality of the subsequently obtained audio features (such as the i-th streaming acoustic processing result) is guaranteed, and thus the quality of the subsequently generated audio data is also guaranteed.

[0073] In a specific example of the disclosed solution, the data to be processed may be obtained in the following manner, specifically including:

[0074] Method 1: The data to be processed is a subphoneme sequence; the specific process is as follows Figure 5 As shown, including:

[0075] Step S501: Processing text data to obtain a target phoneme sequence corresponding to the text data; the text data is sub-text data obtained by segmenting the total text data.

[0076] That is to say, the text data is obtained after segmenting the total text data. For example, the total text data is segmented based on semantics to obtain the text data. In this way, on the one hand, the naturalness of the subsequently output audio is guaranteed, and on the other hand, the quality of the subsequently obtained audio features is guaranteed, thereby ensuring the quality of the generated audio data and effectively improving the user experience.

[0077] Step S502: Processing the text data to obtain a target prosody identification sequence; the prosody identifications in the target prosody identification sequence are used to mark the prosody corresponding to the phonemes in the target phoneme sequence.

[0078] In one example, steps S501 and S502 can be performed simultaneously, for example, by inputting text data into a text processing model to obtain a target phoneme sequence and a target prosody signature sequence corresponding to the text data. Alternatively, in another example, steps S501 and S502 can be performed separately, for example, by using two models to process the text data separately and obtain a target phoneme sequence and a target prosody signature sequence, respectively. It is understood that the disclosed solution is not limited to this, as long as the target phoneme sequence and the target prosody signature sequence can be obtained.

[0079] In a specific example of the present disclosure, the target prosody identification sequence can be obtained in the following manner, specifically including: inputting the text data into a text processing model to obtain an initial prosody identification sequence; processing the initial prosody identification sequence based on preset prosody rules to obtain the target prosody identification sequence.

[0080] For example, prosody correction rules can be pre-set based on actual needs, and then after obtaining the initial prosody identification sequence, correction can be performed based on the prosody correction rules. In this way, the quality of the subsequently obtained audio features (such as the i-th streaming acoustic processing result) is guaranteed, and the quality of the subsequently generated audio data is also guaranteed.

[0081] Step S503: Based on the target prosody identifier sequence, the target phoneme sequence is segmented to obtain a plurality of subphoneme sequences.

[0082] That is to say, in this example, the text data obtained after segmentation can be segmented again based on the target prosody identification sequence, that is, secondary segmentation, to obtain multiple subphoneme sequences, and then each obtained subphoneme sequence can be processed.

[0083] Step S504: When there is no non-streaming acoustic processing module in the acoustic model, the subphoneme sequence is used as the data to be processed.

[0084] That is to say, in this example, there is no non-streaming acoustic processing module in the acoustic model. In other words, each module in the acoustic model supports streaming acoustic processing. At this time, the sub-phoneme sequence can be directly used as the data to be processed.

[0085] It is understandable that the execution order of steps S501 to S504 and steps S101 and S102 is not specifically limited, as long as steps S501 to S504 can be completed before step S103.

[0086] Furthermore, in one specific example, after obtaining the data to be processed, i.e., the subphoneme sequence, the subphoneme sequence is sliced ​​to obtain n data blocks to be acoustically processed. These data blocks are then processed one by one, achieving streaming acoustic processing. It is worth noting that the size of the data blocks to be acoustically processed can be set based on actual needs and is not limited by the present disclosure.

[0087] Thus, because the disclosed solution can further segment the target phoneme sequence based on the target prosody identifier sequence, it effectively ensures the smooth quality and naturalness of the synthesized audio data. Moreover, due to the streaming acoustic processing, the speech processing efficiency is further improved, and the audio response time, such as the first packet response time, can be further reduced, thus further effectively improving the user experience.

[0088] Method 2: The data to be processed is the output result of the non-streaming acoustic processing module in the acoustic model processing the subphone sequence, that is, the first non-streaming acoustic output result. The specific process is as follows: Figure 6 As shown, including:

[0089] Step S601: Processing text data to obtain a target phoneme sequence corresponding to the text data; the text data is sub-text data obtained by segmenting the total text data.

[0090] That is to say, the text data is obtained after segmenting the total text data. For example, the total text data is segmented based on semantics to obtain the text data. In this way, on the one hand, the naturalness of the subsequently output audio is guaranteed, and on the other hand, the quality of the subsequently obtained audio features is guaranteed, thereby ensuring the quality of the generated audio data and effectively improving the user experience.

[0091] Step S602: Process the text data to obtain a target prosody identification sequence; the prosody identifications in the target prosody identification sequence are used to mark the prosody corresponding to the phonemes in the target phoneme sequence.

[0092] In one example, steps S601 and S602 can be performed simultaneously, for example, by inputting text data into a text processing model to obtain a target phoneme sequence and a target prosody signature sequence corresponding to the text data. Alternatively, in another example, steps S601 and S602 can be performed separately, for example, by using two models to process the text data separately and obtain a target phoneme sequence and a target prosody signature sequence, respectively. It is understood that the disclosed solution is not limited to this, as long as the target phoneme sequence and the target prosody signature sequence can be obtained.

[0093] In a specific example of the present disclosure, the target prosody identification sequence can be obtained in the following manner, specifically including: inputting the text data into a text processing model to obtain an initial prosody identification sequence; processing the initial prosody identification sequence based on preset prosody rules to obtain the target prosody identification sequence.

[0094] For example, prosody correction rules can be pre-set based on actual needs, and then after obtaining the initial prosody identification sequence, correction can be performed based on the initial prosody identification sequence. In this way, the quality of the subsequently obtained audio features (such as the i-th streaming acoustic processing result) is guaranteed, and the quality of the subsequently generated audio data is also guaranteed.

[0095] Step S603: Based on the target prosody identifier sequence, the target phoneme sequence is segmented to obtain a plurality of subphoneme sequences.

[0096] That is to say, in this example, the text data obtained after segmentation can be segmented again based on the target prosody identification sequence, that is, secondary segmentation, to obtain multiple subphoneme sequences, and then each obtained subphoneme sequence can be processed.

[0097] Step S604: When a non-streaming acoustic processing module exists in the acoustic model, the subphoneme sequence is input into the non-streaming acoustic processing module in the acoustic model for processing to obtain a first non-streaming acoustic output result corresponding to the subphoneme sequence.

[0098] That is to say, in this example, the acoustic model not only includes a module that supports streaming acoustic processing, i.e., a streaming acoustic processing module, but also includes a module that does not support streaming acoustic processing, i.e., a non-streaming acoustic processing module. At this time, the sub-phoneme sequence needs to be input into the non-streaming acoustic processing module for processing first, and after the non-streaming acoustic processing module completes the processing, the processing result, i.e., the first non-streaming acoustic output result is processed by the streaming acoustic processing module. In this way, the problem of being unable to perform streaming acoustic processing normally due to the existence of a non-streaming acoustic processing module in the acoustic model is effectively avoided. While effectively improving the efficiency of speech processing, it takes into account adaptability (for example, the adapted acoustic models are more diverse) and practicality, and enriches the usage scenarios while improving the user experience.

[0099] Step S605: taking the first non-streaming acoustic output result corresponding to the subphoneme sequence as the data to be processed.

[0100] It is understandable that the execution order of steps S601 to S605 and steps S101 and S102 is not specifically limited, as long as steps S601 to S605 can be completed before step S103.

[0101] Furthermore, in one specific example, after obtaining the data to be processed, i.e., the first non-streaming acoustic output result, the first non-streaming acoustic output result is sliced ​​to obtain n data blocks to be acoustically processed. In this manner, each of these data blocks to be acoustically processed is processed one by one, thereby achieving streaming acoustic processing. It is worth noting that the size of the data blocks to be acoustically processed can be set based on actual needs and is not limited by the disclosed solution.

[0102] Thus, because the disclosed solution can further segment the target phoneme sequence based on the target prosody identifier sequence, it effectively ensures the smooth quality and naturalness of the synthesized audio data. Moreover, due to the streaming acoustic processing, the speech processing efficiency is further improved, and the audio response time, such as the first packet response time, can be further reduced, thus further effectively improving the user experience.

[0103] Method 3: The data to be processed is the target phoneme sequence. The specific process is as follows: Figure 7 As shown, including:

[0104] Step S701: Processing text data to obtain a target phoneme sequence corresponding to the text data; the text data is sub-text data obtained by segmenting the total text data.

[0105] That is to say, the text data is obtained after segmenting the total text data. For example, the total text data is segmented based on semantics to obtain the text data. In this way, on the one hand, the naturalness of the subsequently output audio is guaranteed, and on the other hand, the quality of the subsequently obtained audio features is guaranteed, thereby ensuring the quality of the generated audio data and effectively improving the user experience.

[0106] Step S702: When there is no non-streaming acoustic processing module in the acoustic model, the target phoneme sequence is used as the data to be processed.

[0107] That is to say, in this example, there is no non-streaming acoustic processing module in the acoustic model. In other words, each module in the acoustic model supports streaming acoustic processing. At this time, the target phoneme sequence can be directly used as the data to be processed.

[0108] It is understandable that there is no specific limitation on the execution order of step S701 and step S702 and step S101 and step S102, as long as step 701 and step S702 can be completed before step S103.

[0109] Furthermore, in a specific example, after obtaining the data to be processed, that is, the target phoneme sequence, the target phoneme sequence is sliced ​​to obtain n data blocks to be acoustically processed. In this way, the data blocks to be acoustically processed are processed one by one to achieve streaming acoustic processing.

[0110] It is worth noting that the size of the data block to be acoustically processed can be set based on actual needs, and the present disclosure does not impose any restrictions on this.

[0111] In this way, due to the streaming acoustic processing, the voice processing efficiency is further improved. At the same time, the audio response time can be further reduced, such as the first packet response time, thus further effectively improving the user experience.

[0112] Method 4: The data to be processed is the output result after processing by the non-streaming acoustic processing module in the acoustic model, that is, the second non-streaming acoustic output result. The specific process is as follows: Figure 8 As shown, including:

[0113] Step S801: Processing text data to obtain a target phoneme sequence corresponding to the text data; the text data is sub-text data obtained by segmenting the total text data.

[0114] That is to say, the text data is obtained after segmenting the total text data. For example, the total text data is segmented based on semantics to obtain the text data. In this way, on the one hand, the naturalness of the subsequently output audio is guaranteed, and on the other hand, the quality of the subsequently obtained audio features is guaranteed, thereby ensuring the quality of the generated audio data and effectively improving the user experience.

[0115] Step S802: When there is a non-streaming acoustic processing module in the acoustic model, the target phoneme sequence is input into the non-streaming acoustic processing module in the acoustic model for processing to obtain a second non-streaming acoustic output result corresponding to the target phoneme sequence.

[0116] That is to say, in this example, the acoustic model not only includes a module that supports streaming acoustic processing, i.e., a streaming acoustic processing module, but also includes a module that does not support streaming acoustic processing, i.e., a non-streaming acoustic processing module. At this time, the target phoneme sequence needs to be input into the non-streaming acoustic processing module for processing first, and after the non-streaming acoustic processing module completes the processing, the processing result, i.e., the second non-streaming acoustic output result is processed by the streaming acoustic processing module. In this way, the problem of being unable to perform streaming acoustic processing normally due to the existence of a streaming acoustic processing module that does not support the acoustic model is effectively avoided. While effectively improving the efficiency of speech processing, it takes into account adaptability (for example, the adapted acoustic models are more diverse) and practicality, and enriches the usage scenarios while improving the user experience.

[0117] Step S803: taking the second non-streaming acoustic output result corresponding to the target phoneme sequence as the data to be processed.

[0118] It is understandable that there is no specific limitation on the execution order of steps S801 to S803 and steps S101 and S102, as long as steps S801 to S803 can be completed before step S103.

[0119] Furthermore, in one specific example, after obtaining the data to be processed, i.e., the second non-streaming acoustic output result, the second non-streaming acoustic output result is sliced ​​to obtain n data blocks to be acoustically processed. In this manner, each of these data blocks to be acoustically processed is processed one by one, thereby achieving streaming acoustic processing. It is worth noting that the size of the data blocks to be acoustically processed can be set based on actual needs and is not limited by the disclosed solution.

[0120] In this way, due to the streaming acoustic processing, the voice processing efficiency is further improved. At the same time, the audio response time can be further reduced, such as the first packet response time, thus further effectively improving the user experience.

[0121] In a specific example of the disclosed solution, streaming speech synthesis may be performed in the following manner. Specifically, the above-described streaming speech synthesis based on the i-th streaming acoustic processing result specifically includes:

[0122] Step 1: Based on the parameter features of the streaming speech synthesis module in the speech synthesis model, a speech synthesis filling block is obtained.

[0123] Here, the streaming speech synthesis module refers to the module in the speech synthesis model that supports streaming speech synthesis. Correspondingly, the non-streaming speech synthesis module described below refers to the module in the speech synthesis model that does not support streaming speech synthesis.

[0124] It is understood that the present disclosure does not restrict specific speech synthesis models. Any model capable of speech synthesis can be applied to the present disclosure and can improve speech processing efficiency. Furthermore, the present disclosure does not restrict the number and type of specific modules in the speech synthesis model used to support streaming speech synthesis. For example, all modules in the speech synthesis model can support streaming speech synthesis, or only some modules can support streaming speech synthesis, etc., all of which are within the scope of protection of the present disclosure.

[0125] Step 2: Add the speech synthesis filling block to the jth data block to be speech synthesized to obtain the jth target speech processing data block; wherein the jth data block to be speech processed is obtained based on the i-th streaming acoustic processing result.

[0126] It can be understood that the j-th target speech processing data block is the input of the speech synthesis model, and the j-th target speech processing data block is obtained based on the j-th data block to be speech synthesized, and the j-th data block to be speech synthesized is based on the output of the acoustic model, that is, the i-th streaming acoustic processing result. In other words, the input of the speech synthesis model is obtained based on the output of the acoustic model.

[0127] In a specific example of the disclosed solution, at least one of the following methods is used to add the speech synthesis filling block to the jth data block to be speech synthesized, thereby obtaining the jth target speech processing data block:

[0128] When the value of j is 1, the speech synthesis filling block is added after the first data block to be speech synthesized to obtain the first target speech processing data block;

[0129] When the value of j is any value between 2 and m-1, the speech synthesis filling block is added before and after the j-th data block to be speech synthesized to obtain the j-th target speech processing data block;

[0130] When the value of j is m, the speech synthesis filling block is added in front of the mth data block to be speech synthesized to obtain the mth target speech processing data block.

[0131] For example, the addition method in this example is similar to the acoustic processing filling block. In this case, when j takes a value of 1, the speech synthesis filling block is added after the first data block to be speech synthesized to obtain the first target speech processing data block. When j takes a value between 2 and m-1, the speech synthesis filling block is added before and after the i-th data block to be speech synthesized to obtain the i-th target speech processing data block. For example, when j takes a value of 2, the speech synthesis filling block is added before and after the second data block to be speech synthesized to obtain the second target speech processing data block. When j takes a value of m, the speech synthesis filling block is added before the m-th data block to be speech synthesized to obtain the m-th target speech processing data block. In this way, after splicing the segment results after speech synthesis, the total result corresponding to the i-th streaming acoustic processing result is obtained. At this time, since the speech synthesis data block is processed based on the speech synthesis filling block, it is possible to ensure that the data block size of the total result is consistent with the size of the total result for the i-th streaming acoustic processing result output by the non-streaming speech synthesis. This provides support for the subsequent effective speech synthesis and audio data output.

[0132] It is understandable that the above processing method is only a specific example. In actual applications, other processing methods can also be used to ensure that the splicing result obtained after streaming voice processing is consistent with the result of non-streaming voice processing. The present disclosure does not impose specific restrictions on this.

[0133] In this way, since the speech synthesis filling block is obtained based on the parameter characteristics of the streaming speech synthesis module in the speech synthesis model, it is effectively ensured that the total result obtained after splicing the fragment results after streaming speech synthesis has the same data block size as the total result output by non-streaming speech synthesis, laying the foundation for subsequent effective speech synthesis and audio data output.

[0134] Step 3: Input the j-th target speech processing data block into the streaming speech synthesis module in the speech synthesis model for streaming speech synthesis.

[0135] Here, the output of the acoustic model, such as the i-th streaming acoustic processing result, can be specifically an audio feature, and the output of the speech synthesis model (for example, when the j-th target speech processing data block is input into the streaming speech synthesis module in the speech synthesis model, the j-th streaming speech synthesis result is obtained) can be specifically an audio sample point corresponding to the audio feature. The audio sampling point is related to the audio sampling rate, thereby providing data support for subsequent audio post-processing.

[0136] In this way, the disclosed solution can perform streaming acoustic processing and streaming speech synthesis, thereby further improving the speech processing efficiency and further reducing the audio response time, such as the first packet response time, effectively improving the user experience.

[0137] In a specific example of the disclosed solution, the jth data block to be synthesized can be obtained in the following manner, specifically including:

[0138] Method 1: If Figure 9 As shown, after obtaining the i-th streaming acoustic processing result, as Figure 10 As shown, perform the following steps:

[0139] Step S901: When there is no non-streaming speech synthesis module in the speech synthesis model, remove the acoustic processing filling block from the i-th streaming acoustic processing result and perform slicing processing to obtain m data blocks to be speech synthesized.

[0140] That is to say, in this example, there is no non-streaming speech synthesis module in the speech synthesis model. In other words, each module in the speech synthesis model supports streaming speech synthesis. At this time, the streaming acoustic processing result after removing the acoustic processing filling block, such as the i-th streaming acoustic processing result, can be directly sliced ​​to obtain m data blocks to be speech synthesized, where m is a natural number not less than 2.

[0141] For example, if Figure 10As shown, taking the acoustic processing filling block as an example, after filling the acoustic processing filling block to obtain the first target acoustic processing data block C1, the second target acoustic processing data block C2 to the nth target acoustic processing data block Cn respectively, each target acoustic processing data block obtained will be input into the acoustic model (for example, the streaming acoustic processing module in the acoustic model) for inference to obtain corresponding outputs, such as the first streaming acoustic processing result O1, the second streaming acoustic processing result O2 to the nth streaming acoustic processing result On. At this time, each streaming acoustic processing result contains the acoustic processing filling block after model processing. Therefore, in order to ensure the validity of the data, it is necessary to remove the acoustic processing filling block from each streaming acoustic processing result obtained to obtain a valid streaming acoustic processing result, that is, after removing the acoustic processing filling block, the first target streaming acoustic processing result V1, the second target streaming acoustic processing result V2, to the nth target streaming acoustic processing result Vn are obtained.

[0142] It should be noted that the disclosed solution does not restrict the order in which each target acoustic processing data block is input into the model for inference. For example, they can be input into the model in sequence, or sorted in a queue, and the data input into the model is determined based on the queue sorting results, etc. As long as streaming acoustic processing can be achieved, it falls within the scope of protection of the disclosed solution.

[0143] Furthermore, after obtaining the i-th target streaming acoustic processing result, the i-th target streaming acoustic processing result is sliced ​​to obtain m data blocks to be speech synthesized. In this way, it is convenient to treat the speech synthesis data blocks one by one to realize streaming speech synthesis.

[0144] It is worth noting that the size of the data block to be speech synthesized can be set based on actual needs, and the present disclosure does not impose any restrictions on this.

[0145] Step S902: obtaining a speech synthesis filling block based on parameter features of the streaming speech synthesis module in the speech synthesis model.

[0146] It is understandable that step S901 and step S902 can be swapped. For example, step S902 can be performed first, and then step S901, as long as the data block to be speech synthesized can be obtained before step S903.

[0147] Step S903: adding the speech synthesis filling block to the jth data block to be speech synthesized to obtain the jth target speech processing data block.

[0148] Here, the j-th data block to be speech synthesized is one of the m data blocks to be speech synthesized; and j is a natural number not greater than m.

[0149] Step S904: inputting the j-th target speech processing data block into the streaming speech synthesis module in the speech synthesis model for streaming speech synthesis.

[0150] In this way, since the disclosed solution can effectively perform streaming acoustic processing and streaming speech synthesis, the speech processing efficiency is further improved. At the same time, it can also further reduce the audio response time, such as the first packet response time, thereby further effectively improving the user experience.

[0151] Method 2: If Figure 11 As shown, after obtaining the i-th streaming acoustic processing result, as Figure 10 As shown, perform the following steps:

[0152] Step S1101: When there is a non-streaming speech synthesis module in the speech synthesis model, after removing the acoustic processing filling block from the i-th streaming acoustic processing result, it is input into the non-streaming speech synthesis module in the speech synthesis model to obtain the i-th non-streaming speech synthesis result.

[0153] That is to say, in this example, the speech synthesis model includes not only a module that supports streaming speech synthesis, i.e., a streaming speech synthesis module, but also a module that does not support streaming speech synthesis, i.e., a non-streaming speech synthesis module. At this time, the data output by the acoustic model needs to be input into the non-streaming speech synthesis module for processing first, and after the processing of the non-streaming speech synthesis module is completed, the processing result, i.e., the i-th non-streaming speech synthesis result is processed by the streaming speech synthesis result block. In this way, the problem of being unable to perform streaming speech synthesis normally due to the existence of a module that does not support streaming speech synthesis in the speech synthesis model is effectively avoided. While effectively improving the speech processing efficiency, it takes into account adaptability (for example, the adapted speech synthesis models are more diverse) and practicality, and enriches the usage scenarios while improving the user experience.

[0154] It is understood that, in this example, a specific example of removing the acoustic processing filling block from the i-th streaming acoustic processing result can be found in Figure 10 The relevant description will not be repeated here.

[0155] Step S1102: Slice the i-th non-streaming speech synthesis result to obtain m data blocks to be speech synthesized, where m is a natural number not less than 2.

[0156] It can be understood that there is a non-streaming speech synthesis module in this method. At this time, the i-th streaming acoustic processing result after removing the acoustic processing filling block needs to be input into the non-streaming speech synthesis module for processing, and the corresponding output result, that is, the i-th non-streaming speech synthesis result, is obtained. Subsequently, the i-th non-streaming speech synthesis result is sliced ​​to obtain m data blocks to be speech synthesized. The data blocks to be speech synthesized are the input data to be input into the streaming speech synthesis module.

[0157] Step S1103: obtaining a speech synthesis filling block based on parameter features of the streaming speech synthesis module in the speech synthesis model.

[0158] Step S1104: adding the speech synthesis filling block to the jth data block to be speech synthesized to obtain the jth target speech processing data block.

[0159] Here, the j-th data block to be speech synthesized is one of the m data blocks to be speech synthesized; and j is a natural number not greater than m.

[0160] It is understandable that there is no specific limitation on the execution order of step S1101, step S1102, and step S1103, as long as step S1001 and step S1002 can be completed before step S1104.

[0161] Step S1105: inputting the j-th target speech processing data block into the streaming speech synthesis module in the speech synthesis model for streaming speech synthesis.

[0162] In this way, due to the streaming acoustic processing and streaming speech synthesis, the speech processing efficiency is further improved. At the same time, the audio response time can be further reduced, such as the first packet response time, thereby further effectively improving the user experience.

[0163] The present disclosure provides a speech processing device, such as Figure 12 As shown, including:

[0164] A parameter feature processing unit 1201 is configured to obtain an acoustic processing filling block based on parameter features of the streaming acoustic processing module in the acoustic model;

[0165] The data preprocessing unit 1202 is configured to add the acoustic processing filling block to the i-th data block to be acoustically processed to obtain the i-th target acoustically processed data block; wherein the i-th data block to be acoustically processed is obtained by slicing the data to be processed into n data blocks to be acoustically processed; wherein i is a natural number not greater than n, and n is a natural number not less than 2;

[0166] The speech synthesis unit 1203 is configured to input the i-th target acoustic processing data block into the streaming acoustic processing module in the acoustic model for streaming acoustic processing to obtain an i-th streaming acoustic processing result.

[0167] In a specific example of the presently disclosed scheme, the data preprocessing unit is also used to process text data to obtain a target phoneme sequence corresponding to the text data; the text data is sub-text data obtained after segmentation processing of the total text data; based on the target phoneme sequence, the data to be processed is obtained.

[0168] In a specific example of the present disclosure, the data preprocessing unit is specifically used to input the text data into a text processing model to obtain an initial phoneme sequence; based on preset phoneme rules, the initial phoneme sequence is processed to obtain the target phoneme sequence.

[0169] In a specific example of the disclosed solution, the data preprocessing unit is further configured to:

[0170] Processing the text data to obtain a target prosody identification sequence; the prosody identifications in the target prosody identification sequence are used to mark the prosody corresponding to the phonemes in the target phoneme sequence;

[0171] Based on the target prosody identifier sequence, segmenting the target phoneme sequence to obtain a plurality of subphoneme sequences;

[0172] In the case that there is no non-streaming acoustic processing module in the acoustic model, the subphoneme sequence is used as the data to be processed.

[0173] In a specific example of the disclosed solution, the data preprocessing unit is further configured to:

[0174] Processing the text data to obtain a target prosody identification sequence; the prosody identifications in the target prosody identification sequence are used to mark the prosody corresponding to the phonemes in the target phoneme sequence;

[0175] Based on the target prosody identifier sequence, segmenting the target phoneme sequence to obtain a plurality of subphoneme sequences;

[0176] In a case where a non-streaming acoustic processing module exists in the acoustic model, inputting the subphoneme sequence into the non-streaming acoustic processing module in the acoustic model for processing to obtain a first non-streaming acoustic output result corresponding to the subphoneme sequence;

[0177] The first non-streaming acoustic output result corresponding to the subphoneme sequence is used as the data to be processed.

[0178] In a specific example of the disclosed solution, the data preprocessing unit is further configured to:

[0179] Inputting the text data into a text processing model to obtain an initial prosody identification sequence;

[0180] The initial prosody identification sequence is processed based on a preset prosody rule to obtain the target prosody identification sequence.

[0181] In a specific example of the disclosed solution, the data preprocessing unit is specifically configured to use the target phoneme sequence as the data to be processed when there is no non-streaming acoustic processing module in the acoustic model.

[0182] In a specific example of the present disclosure, the data preprocessing unit is specifically configured to:

[0183] In a case where a non-streaming acoustic processing module exists in the acoustic model, inputting the target phoneme sequence into the non-streaming acoustic processing module in the acoustic model for processing to obtain a second non-streaming acoustic output result corresponding to the target phoneme sequence;

[0184] The second non-streaming acoustic output result corresponding to the target phoneme sequence is used as the data to be processed.

[0185] In a specific example of the present disclosure, the data preprocessing unit is specifically configured to:

[0186] When the value of i is 1, the acoustic processing filling block is added after the first data block to be acoustically processed to obtain the first target acoustic processing data block;

[0187] When i is any value from 2 to n-1, the acoustic processing filling block is added before and after the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block;

[0188] When the value of i is n, the acoustic processing filling block is added in front of the n-th data block to be acoustically processed to obtain the n-th target acoustic processing data block.

[0189] In a specific example of the disclosed solution, the speech synthesis unit is further used to perform streaming speech synthesis based on the i-th streaming acoustic processing result to obtain a streaming speech output result; or, to perform non-streaming speech synthesis based on the i-th streaming acoustic processing result to obtain a non-streaming speech output result.

[0190] In a specific example of the disclosed solution, the parameter feature processing unit is further configured to obtain a speech synthesis filling block based on parameter features of a streaming speech synthesis module in the speech synthesis model;

[0191] The data preprocessing unit is further configured to add the speech synthesis filling block to the jth data block to be speech synthesized to obtain a jth target speech processing data block; wherein the jth data block to be speech processed is obtained based on the i-th streaming acoustic processing result;

[0192] The speech synthesis unit is specifically used to input the j-th target speech processing data block into the streaming speech synthesis module in the speech synthesis model for streaming speech synthesis.

[0193] In a specific example of the disclosed solution, the data preprocessing unit is further configured to:

[0194] When there is no non-streaming speech synthesis module in the speech synthesis model, remove the acoustic processing filling block from the i-th streaming acoustic processing result and perform slicing processing to obtain m data blocks to be speech synthesized;

[0195] The j-th data block to be speech synthesized is one of the m data blocks to be speech synthesized; j is a natural number not greater than m, and m is a natural number not less than 2.

[0196] In a specific example of the disclosed solution, the data preprocessing unit is further configured to:

[0197] In the case where a non-streaming speech synthesis module exists in the speech synthesis model, after removing the acoustic processing filling block from the i-th streaming acoustic processing result, the result is input into the non-streaming speech synthesis module in the speech synthesis model to obtain the i-th non-streaming speech synthesis result;

[0198] Slicing the i-th non-streaming speech synthesis result to obtain m data blocks to be speech synthesized;

[0199] The j-th data block to be speech synthesized is one of the m data blocks to be speech synthesized; j is a natural number not greater than m, and m is a natural number not less than 2.

[0200] In a specific example of the present disclosure, the data preprocessing unit is specifically configured to:

[0201] When the value of j is 1, the speech synthesis filling block is added after the first data block to be speech synthesized to obtain the first target speech processing data block;

[0202] When the value of j is any value between 2 and m-1, the speech synthesis filling block is added before and after the j-th data block to be speech synthesized to obtain the j-th target speech processing data block;

[0203] When the value of j is m, the speech synthesis filling block is added in front of the mth data block to be speech synthesized to obtain the mth target speech processing data block.

[0204] For the description of specific functions and examples of each unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0205] The following is a further detailed description of the disclosed solution with reference to specific examples; specifically, Figure 13 The streaming speech synthesis method of the disclosed solution specifically comprises the following steps:

[0206] Step 1: Create four thread pools; the threads in thread pool 1 are used to map the data to be synthesized to a phoneme sequence, the threads in thread pool 2 are used for streaming AM inference, the threads in thread pool 3 are used for streaming Voc inference, and the threads in thread pool 4 are used for audio post-processing.

[0207] Step 2: Extract a single thread from each of the four thread pools to obtain thread 1, which is used to map the text to be synthesized into a phoneme sequence; thread 2, which is used for streaming AM inference; thread 3, which is used for streaming Voc inference; and thread 4, which is used for audio post-processing. Configure threads 1 to 4 so that they share data.

[0208] Step 3: Start thread 1:

[0209] First, the text to be synthesized is segmented based on semantics to obtain text 1 to text p; this ensures the naturalness of the audio on the one hand, and the quality of the audio features on the other hand, thereby ensuring the quality of the generated audio.

[0210] Subsequently, each text segment is processed by the text front-end module in turn to obtain a phoneme sequence and a rhythm identification sequence.

[0211] Here, the phoneme sequence and prosody mark sequence processed by the text front-end module can be corrected based on the input custom phonemes and prosody correction rules. Furthermore, the corrected phoneme sequence is re-segmented based on the corrected prosody mark sequence, such as prosody markers (e.g., #3 (intonation phrase) or #4 (sentence ending)). The re-segmented phoneme sequence is then sent to thread 2.

[0212] Start threads 2 to 4 in sequence, and threads 2 to 4 can be executed in parallel.

[0213] Step 4: Start thread 2, use the secondary segmented phoneme sequence as the input of streaming AM, i.e., thread 2, obtain audio features, and store the obtained audio features in the audio feature receiving queue for thread 3.

[0214] Step Five: Start Thread 3, use the audio features in the audio feature queue as the input of the streaming Voc, obtain audio sample points, and store the obtained audio sample points into the audio receiving queue for Thread 4.

[0215] Step Six: Start Thread 4, post-process the audio sample points in the audio queue, including changing the sound speed, changing the volume, changing the sampling rate, audio encoding and decoding (pcm, adpcm, opus), etc., and store the processed audio into the output audio queue.

[0216] Step Seven: The audio in the output audio queue can be taken for other operations by the upper-layer application.

[0217] It should be noted that the disclosed solution can be applied to multiple applications, and only the output audio queue needs to be operated later.

[0218] The processing flows of each thread are further described in detail below;

[0219] First, for Thread 1, mapping text to phonemes, the main processes include:

[0220] a) Segment the text to be synthesized based on semantics to obtain each text;

[0221] b) Regularize the text after semantic segmentation in turn. Here, the regularization process includes at least one of the following: converting numbers to text, converting special symbols to text, etc.;

[0222] c) Input the regularized text into the phoneme mapping model to obtain the phoneme sequence and prosody mark sequence corresponding to the text; then perform phonetic conversion according to the positions of special characters (such as polyphonic characters, the character "儿", the character "一", etc.) to obtain a suitable phoneme sequence and prosody mark sequence;

[0223] d) Correct the phoneme sequence and prosody mark sequence obtained in c) according to the input custom phonemes and prosody correction rules;

[0224] e) Based on the corrected prosody mark sequence, perform secondary segmentation on the corrected phoneme sequence to obtain the phoneme sequence after secondary segmentation, and this phoneme sequence after secondary segmentation is used as the input of Thread 2, that is, the streaming AM.

[0225] Second, for Thread 2, streaming AM inference, the main processes include:

[0226] a) Use the phoneme sequence obtained by thread 1 as the input of the streaming AM. According to the situation of the streaming AM, for example, based on the input of the streaming reasoning part of the streaming AM (that is, the streaming acoustic processing module mentioned above) being recorded as Input1, set the valid block1 supported by each AM reasoning, that is, the valid data block.

[0227] It should be noted that if all modules in the streaming AM support streaming acoustic processing, the phoneme sequence obtained by thread 1 can be marked as Input1; if there is a module in the streaming AM that does not support streaming acoustic processing, that is, there is a non-streaming acoustic processing module, then the phoneme sequence obtained by thread 1 is used as the input of the non-streaming acoustic processing module, and the output of the non-streaming acoustic processing module can be marked as Input1.

[0228] b) Based on the size of the effective block1, Input1 is sliced ​​without overlap to obtain n data blocks to be acoustically processed, denoted as [F1, F2, F3, ..., Fn], where the maximum size of Fi (i = 1, 2, 3, ..., n) is block.

[0229] c) According to the configuration parameters of the module supporting streaming acoustic processing in the streaming AM (such as kernel Kermel, size, stride), the acoustic processing filling block, such as the value of pad1, is calculated. The pad1 can ensure that the total result obtained by splicing the fragment results obtained by the streaming acoustic processing module in the streaming AM (such as the total result corresponding to the phoneme sequence input by thread 1) is consistent in size with the output result of the non-streaming acoustic processing module in the streaming AM.

[0230] Here, the size of the acoustic processing padding block is obtained based on the value of pad1.

[0231] d) If Figure 10 As shown, an acoustic processing padding block is added after F1, such as adding pad1 input values, adding pad1 input values ​​before and after Fi (i = 1, 2, 3, ..., n-1), and adding pad1 input values ​​before Fn to generate a target acoustic processing data block, denoted as [C1, C2, C3, ..., Cn]; and performing streaming AM inference in sequence.

[0232] e) The output of streaming AM inference is audio features. To ensure data validity, the final audio features are obtained by removing the values ​​of pad1 from Ci (i = 1, 2, 3, ..., n), and then added to the audio feature queue.

[0233] Third, thread 3, streaming Voc, the main process includes:

[0234] a) Add the audio features in the audio feature queue of thread 2 as the input of the streaming Voc. According to the situation of the streaming Voc, for example, based on the input of the streaming reasoning part of the streaming Voc (that is, the streaming speech synthesis module described above) as Input2, set the valid block2 supported by each Voc reasoning, that is, the valid data block.

[0235] It should be noted that if all modules in the streaming Voc support streaming speech synthesis, the audio features obtained by thread 2 can be marked as Input2; if there is a module in the streaming Voc that does not support streaming speech synthesis, that is, there is a non-streaming speech synthesis module, then the audio features obtained by thread 2 are used as the input of the non-streaming speech synthesis module, and the output of the non-streaming speech synthesis module can be marked as Input2.

[0236] b) Based on the effective size of block2, Input2 is sliced ​​without overlap to obtain m blocks of data to be synthesized, denoted as [F1, F2, F3, ..., Fm], where the maximum size of Fj (j = 1, 2, 3, ..., m) is block.

[0237] c) Based on the configuration parameters of the module supporting streaming speech synthesis in the streaming Voc (such as kernel Kermel, size, and stride), calculate the speech synthesis padding block, such as the value of pad2. The pad2 can ensure that the total result obtained by splicing the fragment results obtained by the streaming speech synthesis module in the streaming Voc (such as the total result corresponding to the audio features input by thread 2) is consistent with the output result of the non-streaming speech synthesis module in the streaming Voc.

[0238] Here, the size of the speech synthesis padding block is obtained based on the value of pad2.

[0239] d) Similar to Figure 10 As shown, a speech synthesis padding block is added after F1, for example, pad inputs are added to, pad2 input values ​​are added before and after Fj (j = 1, 2, 3, ..., m-1), and pad2 input values ​​are added before Fm to obtain the target speech processing data block, which is recorded as [C1, C2, C3, ..., Cm]; and streaming Voc inference is performed in sequence.

[0240] e) The output of streaming Voc inference is an audio sample point. Here, to ensure the validity of the data, the value of pad2 size is removed from Ci (i = 1, 2, 3, ..., m) to obtain the final audio sample point, and then the final audio sample point is added to the audio receiving queue.

[0241] Fourth, thread 4, audio post-processing, the main process includes:

[0242] a) The audio sample points in the audio receiving queue are used as the input of the audio post-processing module, and the audio sample points are processed in sequence by changing the sound speed, changing the volume, changing the sampling rate, and audio codec conversion.

[0243] b) Store the processed audio sample points in the output audio queue for upper-layer application operations (e.g., http, websocket).

[0244] In this way, the disclosed solution can implement streaming AM, thus reducing the first packet response time while ensuring audio quality. Furthermore, the disclosed solution uses a two-level segmentation approach and supports customizable prosody and phoneme correction rules, thus ensuring the smoothness and naturalness of the synthesized audio quality. Furthermore, the disclosed solution supports multiple audio post-processing functions. In other words, the disclosed solution does not restrict audio post-processing functions, thus meeting different user needs and adapting to different application scenarios.

[0245] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0246] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0247] Figure 14 A schematic block diagram of an example electronic device 1400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0248] like Figure 14As shown, device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1402 or a computer program loaded from a storage unit 1408 into a random access memory (RAM) 1403. Various programs and data required for the operation of device 1400 can also be stored in RAM 1403. Computing unit 1401, ROM 1402, and RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to bus 1404.

[0249] Various components in device 1400 are connected to I / O interface 1405, including: an input unit 1406, such as a keyboard, mouse, etc.; an output unit 1407, such as various types of displays, speakers, etc.; a storage unit 1408, such as a magnetic disk, optical disk, etc.; and a communication unit 1409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1409 allows device 1400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0250] The computing unit 1401 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1401 performs the various methods and processes described above, such as the speech processing method. For example, in some embodiments, the speech processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1408. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1400 via the ROM 1402 and / or the communication unit 1409. When the computer program is loaded into the RAM 1403 and executed by the computing unit 1401, one or more steps of the speech processing method described above can be performed. Alternatively, in other embodiments, the computing unit 1401 can be configured to perform the speech processing method by any other appropriate means (e.g., by means of firmware).

[0251] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0252] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0253] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0254] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0255] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0256] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0257] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0258] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A speech processing method, comprising: Based on the parameter characteristics of the streaming acoustic processing module in the acoustic model, an acoustic processing filling block is obtained; Adding the acoustic processing filling block to the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; wherein the i-th data block to be acoustically processed is obtained by slicing the data to be processed into n data blocks to be acoustically processed; i is a natural number not greater than n, and n is a natural number not less than 2; Inputting the i-th target acoustic processing data block into the streaming acoustic processing module in the acoustic model for streaming acoustic processing to obtain an i-th streaming acoustic processing result; The acoustic processing filling block is used to ensure that the total result obtained after splicing the streaming acoustic processing results output by the streaming acoustic processing matches the size of the total result corresponding to the to-be-processed data output by the non-streaming acoustic processing.

2. The method according to claim 1, further comprising: Processing the text data to obtain a target phoneme sequence corresponding to the text data; The text data is sub-text data obtained by segmenting the total text data; The data to be processed is obtained based on the target phoneme sequence.

3. The method according to claim 2, wherein: The processing of the text data to obtain a target phoneme sequence corresponding to the text data includes: Inputting the text data into a text processing model to obtain an initial phoneme sequence; The initial phoneme sequence is processed based on a preset phoneme rule to obtain the target phoneme sequence.

4. The method according to claim 3, further comprising: Processing the text data to obtain a target prosodic identification sequence; The rhythm identifiers in the target rhythm identifier sequence are used to mark the rhythms corresponding to the phonemes in the target phoneme sequence; The step of obtaining the data to be processed based on the target phoneme sequence includes: Based on the target prosody identifier sequence, segmenting the target phoneme sequence to obtain a plurality of subphoneme sequences; In the case that there is no non-streaming acoustic processing module in the acoustic model, the subphoneme sequence is used as the data to be processed.

5. The method according to claim 3, further comprising: Processing the text data to obtain a target prosodic identification sequence; The rhythm identifiers in the target rhythm identifier sequence are used to mark the rhythms corresponding to the phonemes in the target phoneme sequence; The step of obtaining the data to be processed based on the target phoneme sequence includes: Based on the target prosody identifier sequence, segmenting the target phoneme sequence to obtain a plurality of subphoneme sequences; In a case where a non-streaming acoustic processing module exists in the acoustic model, inputting the subphoneme sequence into the non-streaming acoustic processing module in the acoustic model for processing to obtain a first non-streaming acoustic output result corresponding to the subphoneme sequence; The first non-streaming acoustic output result corresponding to the subphoneme sequence is used as the data to be processed.

6. The method according to claim 4 or 5, wherein: The processing of the text data to obtain a target prosody identification sequence includes: Inputting the text data into a text processing model to obtain an initial prosody identification sequence; The initial prosody identification sequence is processed based on a preset prosody rule to obtain the target prosody identification sequence.

7. The method according to claim 2 or 3, wherein: The step of obtaining the data to be processed based on the target phoneme sequence includes: When there is no non-streaming acoustic processing module in the acoustic model, the target phoneme sequence is used as the data to be processed.

8. The method according to claim 2 or 3, wherein: The step of obtaining the data to be processed based on the target phoneme sequence includes: In a case where a non-streaming acoustic processing module exists in the acoustic model, inputting the target phoneme sequence into the non-streaming acoustic processing module in the acoustic model for processing to obtain a second non-streaming acoustic output result corresponding to the target phoneme sequence; The second non-streaming acoustic output result corresponding to the target phoneme sequence is used as the data to be processed.

9. The method according to any one of claims 1 to 5, wherein: At least one of the following methods is used to add the acoustic processing filling block to the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block: When the value of i is 1, the acoustic processing filling block is added after the first data block to be acoustically processed to obtain the first target acoustic processing data block; When i is any value from 2 to n-1, the acoustic processing filling block is added before and after the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; When the value of i is n, the acoustic processing filling block is added in front of the n-th data block to be acoustically processed to obtain the n-th target acoustic processing data block.

10. The method according to any one of claims 1 to 5, further comprising: Performing streaming speech synthesis based on the i-th streaming acoustic processing result to obtain a streaming speech output result; or, Based on the i-th streaming acoustic processing result, non-streaming speech synthesis is performed to obtain a non-streaming speech output result.

11. The method according to claim 10, wherein: The performing streaming speech synthesis based on the i-th streaming acoustic processing result includes: Based on parameter features of a streaming speech synthesis module in a speech synthesis model, a speech synthesis filling block is obtained; Adding the speech synthesis filling block to the jth data block to be speech synthesized to obtain the jth target speech processing data block; wherein the jth data block to be speech synthesized is obtained based on the i-th streaming acoustic processing result; The j-th target speech processing data block is input into the streaming speech synthesis module in the speech synthesis model for streaming speech synthesis.

12. The method according to claim 11, further comprising: When there is no non-streaming speech synthesis module in the speech synthesis model, remove the acoustic processing filling block from the i-th streaming acoustic processing result and perform slicing processing to obtain m data blocks to be speech synthesized; The j-th data block to be speech synthesized is one of the m data blocks to be speech synthesized; j is a natural number not greater than m, and m is a natural number not less than 2.

13. The method according to claim 11, further comprising: In the case where a non-streaming speech synthesis module exists in the speech synthesis model, after removing the acoustic processing filling block from the i-th streaming acoustic processing result, the result is input into the non-streaming speech synthesis module in the speech synthesis model to obtain the i-th non-streaming speech synthesis result; Slicing the i-th non-streaming speech synthesis result to obtain m data blocks to be speech synthesized; The j-th data block to be speech synthesized is one of the m data blocks to be speech synthesized; j is a natural number not greater than m, and m is a natural number not less than 2.

14. The method according to claim 11, wherein At least one of the following methods is used to add the speech synthesis filling block to the j-th data block to be speech synthesized, to obtain the j-th target speech processing data block: When the value of j is 1, the speech synthesis filling block is added after the first data block to be speech synthesized to obtain the first target speech processing data block; When the value of j is any value between 2 and m-1, the speech synthesis filling block is added before and after the j-th data block to be speech synthesized to obtain the j-th target speech processing data block; When the value of j is m, the speech synthesis filling block is added in front of the mth data block to be speech synthesized to obtain the mth target speech processing data block.

15. A speech processing device, comprising: A parameter feature processing unit, configured to obtain an acoustic processing filling block based on parameter features of a streaming acoustic processing module in an acoustic model; a data preprocessing unit, configured to add the acoustic processing filling block to the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; wherein the i-th data block to be acoustically processed is obtained by slicing the data to be processed into n data blocks to be acoustically processed; wherein i is a natural number not greater than n, and n is a natural number not less than 2; a speech synthesis unit, configured to input the i-th target acoustic processing data block into the streaming acoustic processing module in the acoustic model for streaming acoustic processing, thereby obtaining an i-th streaming acoustic processing result; The acoustic processing filling block is used to ensure that the total result obtained after splicing the streaming acoustic processing results output by the streaming acoustic processing matches the size of the total result corresponding to the to-be-processed data output by the non-streaming acoustic processing.

16. The device according to claim 15, wherein The data preprocessing unit is further used to process text data to obtain a target phoneme sequence corresponding to the text data; the text data is sub-text data obtained after segmenting the total text data; and the data to be processed is obtained based on the target phoneme sequence.

17. The device according to claim 16, wherein The data preprocessing unit is specifically configured to input the text data into a text processing model to obtain an initial phoneme sequence; and process the initial phoneme sequence based on a preset phoneme rule to obtain the target phoneme sequence.

18. The device according to claim 17, wherein The data preprocessing unit is further used for: Processing the text data to obtain a target prosody identification sequence; the prosody identifications in the target prosody identification sequence are used to mark the prosody corresponding to the phonemes in the target phoneme sequence; Based on the target prosody identifier sequence, segmenting the target phoneme sequence to obtain a plurality of subphoneme sequences; In the case that there is no non-streaming acoustic processing module in the acoustic model, the subphoneme sequence is used as the data to be processed.

19. The device according to claim 17, wherein The data preprocessing unit is further used for: Processing the text data to obtain a target prosody identification sequence; the prosody identifications in the target prosody identification sequence are used to mark the prosody corresponding to the phonemes in the target phoneme sequence; Based on the target prosody identifier sequence, segmenting the target phoneme sequence to obtain a plurality of subphoneme sequences; In a case where a non-streaming acoustic processing module exists in the acoustic model, inputting the subphoneme sequence into the non-streaming acoustic processing module in the acoustic model for processing to obtain a first non-streaming acoustic output result corresponding to the subphoneme sequence; The first non-streaming acoustic output result corresponding to the subphoneme sequence is used as the data to be processed.

20. The device according to claim 18 or 19, wherein The data preprocessing unit is further used for: Inputting the text data into a text processing model to obtain an initial prosody identification sequence; The initial prosody identification sequence is processed based on a preset prosody rule to obtain the target prosody identification sequence.

21. The device according to claim 16 or 17, wherein The data preprocessing unit is specifically configured to use the target phoneme sequence as the data to be processed when there is no non-streaming acoustic processing module in the acoustic model.

22. The device according to claim 16 or 17, wherein The data preprocessing unit is specifically used to: In a case where a non-streaming acoustic processing module exists in the acoustic model, inputting the target phoneme sequence into the non-streaming acoustic processing module in the acoustic model for processing to obtain a second non-streaming acoustic output result corresponding to the target phoneme sequence; The second non-streaming acoustic output result corresponding to the target phoneme sequence is used as the data to be processed.

23. The device according to any one of claims 15 to 19, wherein The data preprocessing unit is specifically used to: When the value of i is 1, the acoustic processing filling block is added after the first data block to be acoustically processed to obtain the first target acoustic processing data block; When i is any value from 2 to n-1, the acoustic processing filling block is added before and after the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; When the value of i is n, the acoustic processing filling block is added in front of the n-th data block to be acoustically processed to obtain the n-th target acoustic processing data block.

24. The device according to any one of claims 15 to 19, wherein The speech synthesis unit is further configured to perform streaming speech synthesis based on the i-th streaming acoustic processing result to obtain a streaming speech output result; Alternatively, non-streaming speech synthesis is performed based on the i-th streaming acoustic processing result to obtain a non-streaming speech output result.

25. The apparatus according to claim 24, wherein The parameter feature processing unit is further configured to obtain a speech synthesis filling block based on the parameter features of the streaming speech synthesis module in the speech synthesis model; The data preprocessing unit is further configured to add the speech synthesis filling block to the jth data block to be speech synthesized to obtain a jth target speech processing data block; wherein the jth data block to be speech synthesized is obtained based on the i-th streaming acoustic processing result; The speech synthesis unit is specifically used to input the j-th target speech processing data block into the streaming speech synthesis module in the speech synthesis model for streaming speech synthesis.

26. The device according to claim 25, wherein The data preprocessing unit is further used for: When there is no non-streaming speech synthesis module in the speech synthesis model, remove the acoustic processing filling block from the i-th streaming acoustic processing result and perform slicing processing to obtain m data blocks to be speech synthesized; The j-th data block to be speech synthesized is one of the m data blocks to be speech synthesized; j is a natural number not greater than m, and m is a natural number not less than 2.

27. The apparatus according to claim 25, wherein The data preprocessing unit is further used for: In the case where a non-streaming speech synthesis module exists in the speech synthesis model, after removing the acoustic processing filling block from the i-th streaming acoustic processing result, the result is input into the non-streaming speech synthesis module in the speech synthesis model to obtain the i-th non-streaming speech synthesis result; Slicing the i-th non-streaming speech synthesis result to obtain m data blocks to be speech synthesized; The j-th data block to be speech synthesized is one of the m data blocks to be speech synthesized; j is a natural number not greater than m, and m is a natural number not less than 2.

28. The apparatus according to claim 25, wherein The data preprocessing unit is specifically used to: When the value of j is 1, the speech synthesis filling block is added after the first data block to be speech synthesized to obtain the first target speech processing data block; When the value of j is any value between 2 and m-1, the speech synthesis filling block is added before and after the j-th data block to be speech synthesized to obtain the j-th target speech processing data block; When the value of j is m, the speech synthesis filling block is added in front of the mth data block to be speech synthesized to obtain the mth target speech processing data block.

29. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.

30. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.

31. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Speech synthesis method and device, computer equipment and storage medium

    CN114360490A