Speech synthesis method and device based on multi-scale style, equipment and medium

By employing a multi-scale style speech synthesis method that combines style analysis of audio and text and integrates style information from different scales, the problem of a single speech style in existing speech synthesis technologies is solved, and more emotional speech synthesis is achieved.

CN116597807BActive Publication Date: 2026-04-17PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2023-06-15
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing speech synthesis technologies are relatively simple in terms of emotion and expression, have a strong machine-like feel, and lack rich differences in speech style.

Method used

A multi-scale style speech synthesis method is adopted, which extracts style embedding vectors from the original speech and text, and fuses style information at different scales to synthesize the target speech.

Benefits of technology

It enhances the emotional richness of speech synthesis, reduces the robotic feel, and improves the quality of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116597807B_ABST
    Figure CN116597807B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology and discloses a speech synthesis method, apparatus, computer device, and storage medium based on multi-scale style, which solves the problems of traditional speech synthesis schemes having a strong machine feel and insufficient emotional richness. The method includes: extracting target audio and target text corresponding to the original speech; performing style analysis on the target audio to obtain a first style embedding vector; performing style prediction on the target text to obtain a second style embedding vector; fusing the first style embedding vector and the second style embedding vector to obtain a target style embedding vector; and synthesizing target speech based on the target style embedding vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a speech synthesis method, apparatus, computer device and storage medium based on multi-scale style. Background Technology

[0002] While existing speech synthesis technology has made great strides, in real-life production and daily life, people can easily distinguish whether the other end of a conversation is a robot or a real person. This is because synthesized speech data generally aims for stability, and therefore does not have much richness in terms of emotion and expression.

[0003] In recent years, as people have shown increasing interest and demand for emotional and personalized synthesis, the current focus of emotional speech synthesis work is mainly on obtaining contextual information from sentences to build a single-scale model, while ignoring the differences in speech style at different scales. This results in synthesized speech styles being relatively monotonous, not rich enough, and with a noticeable machine feel. Summary of the Invention

[0004] This application provides a speech synthesis method, apparatus, computer device, and storage medium based on multi-scale style to solve the problem that the style of synthesized speech in traditional solutions is relatively simple, not rich enough, and has a noticeable machine feel.

[0005] A speech synthesis method based on multi-scale style includes:

[0006] Extract the target audio and target text corresponding to the original speech;

[0007] Perform style analysis on the target audio to obtain a first style embedding vector;

[0008] Perform style prediction on the target text to obtain a second style embedding vector;

[0009] The first style embedding vector and the second style embedding vector are fused to obtain the target style embedding vector;

[0010] The target speech is synthesized based on the target style embedding vector.

[0011] A speech synthesis device based on multi-scale style includes:

[0012] The extraction module is used to extract the target audio and target text corresponding to the original speech;

[0013] The style analysis module is used to perform style analysis on the target audio to obtain a first style embedding vector;

[0014] The style prediction module is used to perform style prediction on the target text to obtain a second style embedding vector;

[0015] A fusion module is used to fuse the first style embedding vector and the second style embedding vector to obtain a target style embedding vector;

[0016] A synthesis module is used to synthesize target speech based on the target style embedding vector.

[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the multi-scale style-based speech synthesis method described above.

[0018] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described multi-scale style-based speech synthesis method.

[0019] The above-mentioned multi-scale style-based speech synthesis method, device, computer equipment, and storage medium propose a multi-scale style extraction and embedding method compared with traditional methods. It fully extracts speech styles from different scales, highlights the style and emotion of synthesized speech data, and introduces speech style analysis and prediction at different scales to help express the emotions of synthesized speech, improve the synthesis quality of emotional speech, and obtain the final emotional synthesized speech, thus solving the problem that traditional speech synthesis schemes have a strong machine feel and lack rich emotion. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of an application environment for a speech synthesis method based on multi-scale style in one embodiment of this application;

[0022] Figure 2 This is a flowchart of a speech synthesis method based on multi-scale style in one embodiment of this application;

[0023] Figure 3 yes Figure 2 A flowchart illustrating a specific implementation of step S20;

[0024] Figure 4This is another flowchart of a speech synthesis method based on multi-scale style in one embodiment of this application;

[0025] Figure 5 yes Figure 3 A flowchart illustrating a specific implementation of step S25;

[0026] Figure 6 yes Figure 2 A flowchart of a specific implementation method for step S30;

[0027] Figure 7 This is a schematic diagram of a speech synthesis device based on multi-scale style in one embodiment of this application;

[0028] Figure 8 This is a schematic diagram of the structure of a computer device according to one embodiment of this application. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] The speech synthesis method based on multi-scale style provided in this application can be applied to, for example... Figure 1 In this application environment, the client can communicate with the server via a network. The client can provide various raw speech samples. After receiving the raw speech, the server can extract the corresponding target audio and target text; perform style analysis on the target audio to obtain a first style embedding vector; perform style prediction on the target text to obtain a second style embedding vector; fuse the first and second style embedding vectors to obtain a target style embedding vector, thus obtaining multi-scale style information; finally, synthesize the target speech based on the target style embedding vector. Compared with traditional methods, this approach proposes a multi-scale style extraction and embedding method, fully extracting speech styles from different scales to highlight the style and emotion of the synthesized speech data. It also proposes a multi-scale style prediction module that incorporates context, introducing speech style analysis and prediction at different scales to help express the emotion of the synthesized speech, improving the synthesis quality of emotional speech, and ultimately obtaining emotionally rich synthesized speech. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0031] First, before describing the embodiments of this application, it is necessary to clarify several terms involved in this application:

[0032] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0033] Natural Language Processing (NLP): NLP is a type of artificial intelligence that specializes in analyzing human language. Its working principle is roughly as follows: it receives natural language, which has evolved from human natural use and is used by humans to communicate every day. It translates natural language and analyzes it through probabilistic algorithms and outputs the results.

[0034] Embedding: Embedding is a vector representation that refers to representing an object with a low-dimensional vector. The object can be a word, a product, a movie, etc.; or an emotional style mentioned in this application. It is often used in machine learning. In the process of building machine learning models, objects are encoded into low-dimensional dense vectors and then passed to the network or encoder / decoder for processing to improve processing efficiency.

[0035] In one embodiment, such as Figure 2 As shown, a speech synthesis method based on multi-scale style is provided, which is then applied to... Figure 1 Taking the server in the example, the following steps are included:

[0036] S10: Extract the target audio and target text corresponding to the original speech;

[0037] In this embodiment, a speech segment is first acquired, which includes multiple sentences. Any sentence within this speech segment is recorded as the original speech. That is, the original speech is any sentence within the acquired speech segment. After obtaining the original speech, the corresponding audio and text are extracted from the original speech and recorded as the target audio and target text.

[0038] It should be noted that, in specific implementation, the original speech can be input into the trained audio encoder and text encoder respectively to obtain the target audio and target text corresponding to the original speech. The above extraction process can be obtained based on commonly used audio extraction algorithms and text extraction algorithms, which will not be described in detail here.

[0039] Furthermore, the same processing method is applied to any other original speech. For ease of description, this application uses one sentence of original speech as an example for illustration.

[0040] S20: Perform style analysis on the target audio to obtain a first style embedding vector;

[0041] S30: Perform style prediction on the target text to obtain a second style embedding vector;

[0042] After obtaining the target audio and target text corresponding to the original speech, in order to obtain rich speech synthesis effects rather than speech that sounds robotic and lacks emotional expression, this embodiment of the application performs two different style scales of processing in two separate ways. The first processing path is to perform style analysis on the target audio to obtain a first style embedding vector, denoted as the first style embedding vector. This first style embedding vector represents the style information of the audio dimension. It can be understood that the audio contains the most direct emotional style expression of the original speech, so it is necessary to perform style analysis on the target audio to obtain the style corresponding to the audio and convert it into a style embedding vector. The second processing path is to perform style prediction on the target text to obtain another style embedding vector, denoted as the second style embedding vector. This second style embedding vector represents the style information of the character dimension. It can be understood that the text in the original speech often also reflects emotional style. For example, "Why did you do this to me? I'm so sad." From the above text, it can be seen that the text carries a sad style. Another example is "Why did you do this to me? I'm so sad!!", from which it can be seen that in addition to carrying a sad style, the text also carries an angry or resentful style. It is evident that the text contains another layer of emotional style expression of the original speech. Therefore, it is necessary to perform style prediction on the target text corresponding to the original speech, predict the style behind the text, and convert it into a second style embedding vector to represent it.

[0043] S40: Merge the first style embedding vector and the second style embedding vector to obtain the target style embedding vector;

[0044] S50: Synthesize the target speech based on the target style embedding vector.

[0045] After obtaining the emotional style at both the audio and text scales, the corresponding style embedding vectors are acquired. Finally, multi-scale style fusion is performed, fusing the first and second style embedding vectors to obtain the target style embedding vector. Finally, the target speech is synthesized based on the target style embedding vector and the corresponding speech data. In one embodiment, the above fusion refers to directly superimposing the first and second style embedding vectors to obtain the target style embedding vector. No specific limitation is imposed, and other fusion methods are also possible; this application does not limit the specific embodiments.

[0046] As can be seen, the embodiments of this application provide a speech synthesis method based on multi-scale style. Compared with traditional schemes, it proposes a multi-scale style extraction and embedding method, which fully extracts speech styles from different scales to highlight the style and emotion of synthesized speech data. It also proposes a multi-scale style prediction module that combines context, introduces speech style analysis and prediction at different scales, helps to express the emotion of synthesized speech, improves the synthesis quality of emotional speech, and can obtain the final emotional synthesized speech, solving the problem that traditional speech synthesis schemes have a strong machine feel and lack rich emotion.

[0047] In conjunction with the above embodiments, in this application embodiment, in order to further improve the emotional expression of synthesized speech and enhance speech richness, and also considering the correlation between information at different levels, further optimizations have been made in the construction of style embedding vectors. Specifically, contextual information between different levels is combined to help the synthesized speech express emotions and improve the synthesis quality of emotional speech.

[0048] In one embodiment, such as Figure 3 As shown, step S20, namely, performing style analysis on the target audio to obtain the first style embedding vector, specifically includes the following steps: (Write the corresponding implementation examples for the weights).

[0049] S21: Extract the Mel spectrum of the target audio as a local Mel spectrum;

[0050] S22: Obtain the Mel spectrum of the context speech of the target audio, and concatenate the Mel spectrum of the context audio and the Mel spectrum of the target audio to obtain the global Mel spectrum;

[0051] S23: Extract the Mel spectrum of the sub-audio in the target audio according to the sub-word phoneme boundary as the segment Mel spectrum;

[0052] S24: Style encoding is performed on the global Mel spectrum, local Mel spectrum and segment Mel spectrum respectively, and the encoded style information is input into the corresponding style tag layer to obtain the global audio style vector, local audio style vector and segment audio style vector;

[0053] S25: Based on the global audio style vector, local audio style vector, and segment audio style vector, the total emotional style variable of the target audio is obtained as the first style embedding vector.

[0054] For steps S21-S25, when constructing the first style embedding vector corresponding to the target audio, the corresponding local Mel spectrum, global Mel spectrum, and segment Mel spectrum are first analyzed. The local Mel spectrum is the Mel spectrum obtained by Mel spectrum transformation of the current target audio. The global Mel spectrum is obtained by concatenating the local Mel spectrum of the target audio with the Mel spectra of its context audio. Specifically, it involves concatenating the Mel spectra of 2n+1 sentences (where n refers to the number of sentences involving the context of the target audio). n can be set empirically and is not specifically limited. The segment Mel spectrum is the Mel spectrum obtained by Mel spectrum transformation of the audio segments corresponding to the words divided based on the phoneme boundaries of the target audio.

[0055] In this specific implementation, a reference encoder is introduced to process global, local, and fragment-scale Mel spectra, namely a global reference encoder, a local reference encoder, and a fragment reference encoder. The Mel spectra corresponding to 2n+1 sentences of the target audio (n refers to the number of sentences involving the context) are concatenated as the input of the global reference encoder. The current target audio is used as the input of the local reference encoder. The word boundaries between phonemes of the target audio are used as the segmentation criteria to divide the audio segments corresponding to the sub-words as the input of the fragment reference encoder. Thus, the global reference encoder, the local reference encoder, and the fragment reference encoder output the global Mel spectra, the local Mel spectra, and the fragment Mel spectra, respectively.

[0056] It should be noted that there are multiple ways to segment words, such as various methods of spectral and phoneme or character alignment. For example, the MFA (Forced Alignment Tool) method used in Fastspeech2 maps speech and text together, identifying which segment of the spectrum each character corresponds to, thus obtaining the corresponding set of words.

[0057] After obtaining the global Mel spectrum, local Mel spectrum, and segment Mel spectrum, style encoding is performed on each of them. The encoded style information is then input into the corresponding style tag layers—that is, the global style tag layer, local style tag layer, and segment style tag layer—to obtain the global audio style vector, local audio style vector, and segment audio style vector. Finally, based on these vectors, the total emotional style variable of the target audio is obtained, and this total emotional style variable serves as the first style embedding vector. Specifically, the global audio style vector, local audio style vector, and segment audio style vector are superimposed to obtain the first style embedding vector.

[0058] In this embodiment, based on the hierarchical relationship of the target audio and considering the correlation between information at different levels, contextual information between different levels is combined to construct the style vector corresponding to the audio, so as to help the synthesized speech express emotions, improve the synthesis quality of emotional speech, and fully consider the information correlation between extracted speech styles to obtain better emotional synthesized speech.

[0059] It should be noted that, in one embodiment, considering the issue of information redundancy, and in order to further reduce the processing workload and improve voice processing efficiency, in one embodiment, such as... Figure 5 As shown, step S25, namely, performing style encoding on the global Mel spectrum, local Mel spectrum, and segment Mel spectrum respectively, and inputting the encoded style information into the corresponding style tag layer to obtain the global audio style vector, local audio style vector, and segment audio style vector, includes:

[0060] S251: Style-encode the global Mel spectrum to obtain a global audio style as the first residual style, and style-encode the local Mel spectrum and the segment Mel spectrum respectively to obtain a local audio style and a segment audio style;

[0061] S252: Subtract the global audio style from the local audio style to obtain the second residual style;

[0062] S253: Subtract the local audio style from the segment audio style to obtain the third residual style;

[0063] S254: Input the first residual style, the second residual style, and the third residual style into the corresponding style tag layers respectively to obtain the global audio style vector, the local audio style vector, and the segment audio style vector.

[0064] In this embodiment, such as Figure 4As shown, considering the issue of information redundancy, the lower-scale levels will subtract the embedding style obtained from the previous scale level. Finally, residual styles for embedding at different scales can be obtained from the reference encoders at the above three scales, denoted as the first residual style R. global The second residual style vector R local and the third residual style vector R segment These three residual styles are fed into the corresponding style tagging layers, which output the corresponding style tags, providing style information for subsequent different scales. After processing by the style tagging layers, the corresponding global audio style vector Em, which is embedded globally, can be obtained. global Local audio style vector Em local and fragment audio style vector Em segment Ultimately, for each target audio segmented during the encoding phase, the corresponding multi-scale style is the sum of the embeddings of the three scale styles, Em. total =Em global +Em local +Em segmengt Em total That is, the first style embedding vector.

[0065] In this embodiment, considering hierarchical and contextual relationships, and taking into account the issue of information redundancy, the lower-scale level subtracts the embedding style obtained from the previous scale level. This effectively reduces the processing of redundant style information. Using this method of subtracting the output of the previous scale as a constraint reduces the repeated modeling of the same information. Finally, the residual styles of different scale embeddings can be obtained from the reference encoders of the above three scales, which is the first residual style R. global The second residual style vector R local and the third residual style vector R segment These residual styles are fed into the corresponding style label layer, which outputs the corresponding style labels to provide style information for the subsequent predictive encoder and construct the first style embedding vector.

[0066] In one embodiment, such as Figure 6 As shown, step S30, which involves performing style prediction on the target text to obtain the second style embedding vector, specifically includes the following steps:

[0067] S31: Extract the semantics of the target text as a local semantic sequence;

[0068] S32: Concatenate the context text of the target text with the target text to obtain a global semantic sequence by extracting the semantics of the concatenated text;

[0069] S33: Extract the semantic sequence of the sub-word set divided from the target text as a segment semantic sequence;

[0070] S34: Perform style prediction on the global semantic sequence, local semantic sequence and fragment semantic sequence respectively to obtain global text style vector, local text style vector and fragment text style vector;

[0071] S35: Superimpose the global text style vector, local text style vector, and fragment text style vector to obtain the total sentiment style variable of the target text as the second style embedding vector.

[0072] For steps S31-S35, when constructing the second-style embedding vector corresponding to the target text, it is first necessary to analyze the corresponding global semantic sequence, local semantic sequence, and fragment semantic sequence, such as... Figure 4 As shown, these three semantic sequences can be converted using a hierarchical text editor. The local semantic sequence is the sequence obtained by semantically converting the current target text. The global semantic sequence is the concatenated text obtained by semantically converting the target text and its context texts. Specifically, it involves concatenating 2n+1 texts (where n refers to the number of context texts related to the target text). The value of n can be set empirically and is not limited. The fragment semantic sequence refers to the Mel spectrum obtained by semantically converting the text fragments corresponding to the sub-words segmented based on the phoneme boundaries of the target text.

[0073] In its specific implementation, this embodiment introduces a hierarchical text editor that processes global, local, and fragment-scale semantics. It also introduces a global predictive encoder, a local predictive encoder, and a fragment predictive encoder for segmentation prediction based on semantic sequences. Notably, the predictive encoder consists of fully connected layers and activation functions, similar to the aforementioned reference encoder. The concatenated text is used as input to the global predictive encoder; the current target text is used as input to the local reference encoder; and the text fragments corresponding to the sub-words are used as input to the fragment predictive encoder. Thus, the global predictive encoder, the local predictive encoder, and the fragment predictive encoder each output a global text style vector P. global Local text style vector P local and fragment text style vector P segmengt .

[0074] It should be noted that there are multiple ways to segment words, such as various methods of spectral and phoneme or character alignment. For example, the MFA (Forced Alignment Tool) method used in Fastspeech2 maps speech and text together, identifying which segment of the spectrum each character corresponds to, thus obtaining the corresponding set of words.

[0075] After obtaining the global text style vector, local text style vector, and fragment text style vector, the total sentiment style variable of the target text is obtained based on these vectors. This total sentiment style variable will be used as the second style embedding vector. Specifically, the global text style vector, local text style vector, and fragment text style vector are superimposed to obtain the second style embedding vector.

[0076] In this embodiment, the style embeddings of these three scales attempt to recover the multi-scale speaking style in human speech by considering contextual information at different levels. At the same time, style prediction is performed on the text, and contextual and hierarchical information are also considered when predicting the style based on the text. These embeddings are superimposed to form the multi-scale style embedding of each target text in the current sentence.

[0077] In one embodiment, step S34, which involves performing style prediction on the global semantic sequence, local semantic sequence, and fragment semantic sequence to obtain a global text style vector, a local text style vector, and a fragment text style vector, includes: inputting the global semantic sequence, local semantic sequence, and fragment semantic sequence into a global style predictor, a local style predictor, and a fragment style predictor, respectively, to obtain a global text style vector, a local text style vector, and a fragment text style vector, respectively; wherein, the global text style serves as a style condition constraint for the local style predictor, and the local text style serves as a style condition constraint for the fragment style predictor.

[0078] In this embodiment, the global style embedding is used as a conditional constraint for the low-level style prediction encoder, and different scale embedded styles are sequentially generated from the multi-scale style predictor, including the global text style vector P. global Local text style vector P local and fragment text style vector P segmengt The predictor's training objective comes from the corresponding real style embeddings in the extractor, which further ensures the hierarchical relationship while improving the accuracy and relevance of style prediction. Considering the correlation between information at different levels, it combines contextual information between different levels to help synthesize emotional speech and improve the quality of emotional speech synthesis.

[0079] These three-scale style embeddings attempt to recover multi-scale speaking styles in human speech by considering different levels of contextual information. Finally, all embedded style vectors are superimposed to form the multi-scale style embedding for each segment of the current sentence. This is then processed by a differential adapter and decoder to obtain the final, emotionally rich synthesized speech. Here, the differential adapter predicts the pitch and duration required to synthesize correct speech. During training, the actual duration and pitch extracted from ground truth speech guide the optimization of the MSE loss, enabling the predictor to synthesize the duration of the correct pitch, thus ensuring correct prediction during inference.

[0080] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0081] In one embodiment, a speech synthesis device based on multi-scale style is provided, which corresponds one-to-one with the speech synthesis method based on multi-scale style in the above embodiments. For example... Figure 7 As shown, the multi-scale style-based speech synthesis device 10 includes an extraction module 101, a style prediction module 102, a style analysis module 103, a fusion module 104, and a synthesis module 105. Detailed descriptions of each functional module are as follows:

[0082] Extraction module 101 is used to extract the target audio and target text corresponding to the original speech;

[0083] Style prediction module 102 is used to perform style analysis on the target audio to obtain a first style embedding vector;

[0084] Style analysis module 103 is used to perform style prediction on the target text to obtain a second style embedding vector;

[0085] The fusion module 104 is used to fuse the first style embedding vector and the second style embedding vector to obtain the target style embedding vector;

[0086] The synthesis module 105 is used to synthesize target speech based on the target style embedding vector.

[0087] In one embodiment, the style analysis module 102 is specifically used for:

[0088] Extract the Mel spectrum of the target audio as a local Mel spectrum;

[0089] Obtain the Mel spectrum of the context speech of the target audio, and concatenate the Mel spectrum of the context audio and the Mel spectrum of the target audio to obtain the global Mel spectrum;

[0090] Extract the Mel spectra of the sub-audio segments from the target audio segment, which are divided according to the boundaries of the sub-word phonemes, as the segment Mel spectra;

[0091] Style encoding is performed on the global Mel spectrum, local Mel spectrum, and segment Mel spectrum respectively, and the encoded style information is input into the corresponding style tag layer to obtain the global audio style vector, local audio style vector, and segment audio style vector.

[0092] Based on the global audio style vector, local audio style vector, and segment audio style vector, the total emotional style variable of the target audio is obtained as the first style embedding vector.

[0093] In one embodiment, the style analysis module 102 is further specifically used for:

[0094] The global Mel spectrum is style-coded to obtain a global audio style as the first residual style. The local Mel spectrum and the segment Mel spectrum are style-coded respectively to obtain the local audio style and the segment audio style.

[0095] Subtracting the global audio style from the local audio style yields the second residual style;

[0096] Subtracting the local audio style from the segment audio style yields the third residual style;

[0097] The first residual style, the second residual style, and the third residual style are input into the corresponding style label layer to obtain the global audio style vector, the local audio style vector, and the segment audio style vector.

[0098] In one embodiment, the style prediction module 103 is specifically used for:

[0099] Extract the semantics of the target text as a local semantic sequence;

[0100] The context text of the target text is concatenated with the target text to form a spliced ​​text, and the semantics of the spliced ​​text are extracted to obtain a global semantic sequence;

[0101] Extract the semantic sequence of the sub-word set divided from the target text as a segment semantic sequence;

[0102] Style prediction is performed on the global semantic sequence, local semantic sequence, and fragment semantic sequence respectively to obtain global text style vector, local text style vector, and fragment text style vector;

[0103] The global text style vector, local text style vector, and fragment text style vector are superimposed to obtain the total sentiment style variable of the target text, which is used as the second style embedding vector.

[0104] In one embodiment, the style prediction module 103 is further specifically used for:

[0105] The global semantic sequence, local semantic sequence, and fragment semantic sequence are respectively input into the global style predictor, local style predictor, and fragment style predictor to obtain the global text style vector, local text style vector, and fragment text style vector, respectively.

[0106] The global text style serves as the style condition constraint for the local style predictor, and the local text style serves as the style condition constraint for the fragment style predictor.

[0107] In one embodiment, the fusion module 104 is specifically used for:

[0108] The first style embedding vector and the second style embedding vector are superimposed to obtain the target style embedding vector.

[0109] As can be seen, the embodiments of this application provide a speech synthesis device based on multi-scale style. Compared with traditional solutions, it proposes a multi-scale style extraction and embedding method, which fully extracts speech styles from different scales to highlight the style and emotion of synthesized speech data. It also proposes a multi-scale style prediction module that combines context, introduces speech style analysis and prediction at different scales, helps to express the emotion of synthesized speech, improves the synthesis quality of emotional speech, and can obtain the final emotional synthesized speech, solving the problem that traditional speech synthesis solutions have a strong machine feel and lack rich emotion.

[0110] Specific limitations regarding the multi-scale style-based speech synthesis device can be found in the limitations of the multi-scale style-based speech synthesis method above, and will not be repeated here. Each module in the aforementioned multi-scale style-based speech synthesis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0111] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores raw speech and synthesized speech. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a multi-scale style-based speech synthesis method. The computer-readable storage media includes volatile and / or non-volatile storage media.

[0112] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0113] Extract the target audio and target text corresponding to the original speech;

[0114] Perform style analysis on the target audio to obtain a first style embedding vector;

[0115] Perform style prediction on the target text to obtain a second style embedding vector;

[0116] The first style embedding vector and the second style embedding vector are fused to obtain the target style embedding vector;

[0117] The target speech is synthesized based on the target style embedding vector.

[0118] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0119] Extract the target audio and target text corresponding to the original speech;

[0120] Perform style analysis on the target audio to obtain a first style embedding vector;

[0121] Perform style prediction on the target text to obtain a second style embedding vector;

[0122] The first style embedding vector and the second style embedding vector are fused to obtain the target style embedding vector;

[0123] The target speech is synthesized based on the target style embedding vector.

[0124] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0125] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0126] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A multi-scale style based speech synthesis method, characterized by, include: Extract the target audio and target text corresponding to the original speech; Perform style analysis on the target audio to obtain a first style embedding vector; Perform style prediction on the target text to obtain a second style embedding vector; The first style embedding vector and the second style embedding vector are fused to obtain the target style embedding vector; Target speech is synthesized based on the target style embedding vector; The step of performing style analysis on the target audio to obtain a first style embedding vector includes: Extract the Mel spectrum of the target audio as a local Mel spectrum; Obtain the Mel spectrum of the context speech of the target audio, and concatenate the Mel spectrum of the context speech and the Mel spectrum of the target audio to obtain the global Mel spectrum; Extract the Mel spectra of the sub-audio segments from the target audio segment, which are divided according to the boundaries of the sub-word phonemes, as the segment Mel spectra; The global Mel spectrum is style-encoded to obtain a global audio style as the first residual style. The local Mel spectrum and the segment Mel spectrum are style-encoded to obtain local audio styles and segment audio styles, respectively. The global audio style is subtracted from the local audio style to obtain the second residual style. The local audio style is subtracted from the segment audio style to obtain the third residual style. The first residual style, the second residual style, and the third residual style are input into the corresponding style label layer to obtain the global audio style vector, the local audio style vector, and the segment audio style vector, respectively. Based on the global audio style vector, local audio style vector, and segment audio style vector, the total emotional style variable of the target audio is obtained as the first style embedding vector.

2. The multi-scale style based speech synthesis method of claim 1, wherein, The step of performing style prediction on the target text to obtain a second style embedding vector includes: Extract the semantics of the target text as a local semantic sequence; The context text of the target text is concatenated with the target text to form a spliced ​​text, and the semantics of the spliced ​​text are extracted to obtain a global semantic sequence; Extract the semantic sequence of the sub-word set divided from the target text as a segment semantic sequence; Style prediction is performed on the global semantic sequence, local semantic sequence, and fragment semantic sequence respectively to obtain global text style vector, local text style vector, and fragment text style vector; The global text style vector, local text style vector, and fragment text style vector are superimposed to obtain the total sentiment style variable of the target text, which is used as the second style embedding vector.

3. The multi-scale style based speech synthesis method of claim 2, wherein, The step of performing style prediction on the global semantic sequence, local semantic sequence, and fragment semantic sequence respectively to obtain global text style vector, local text style vector, and fragment text style vector includes: The global semantic sequence, local semantic sequence, and fragment semantic sequence are respectively input into the global style predictor, local style predictor, and fragment style predictor to obtain the global text style vector, local text style vector, and fragment text style vector, respectively. The global text style serves as the style condition constraint for the local style predictor, and the local text style serves as the style condition constraint for the fragment style predictor.

4. The multi-scale style based speech synthesis method of any one of claims 1-3, wherein, The first style embedding vector and the second style embedding vector are fused to obtain the target style embedding vector, including: The first style embedding vector and the second style embedding vector are superimposed to obtain the target style embedding vector.

5. A multi-scale style-based speech synthesis device, characterized by, include: The extraction module is used to extract the target audio and target text corresponding to the original speech; The style analysis module is used to perform style analysis on the target audio to obtain a first style embedding vector; The style prediction module is used to perform style prediction on the target text to obtain a second style embedding vector; A fusion module is used to fuse the first style embedding vector and the second style embedding vector to obtain a target style embedding vector; The synthesis module is used to synthesize target speech based on the target style embedding vector; The style prediction module is specifically used for: Extract the Mel spectrum of the target audio as a local Mel spectrum; Obtain the Mel spectrum of the context speech of the target audio, and concatenate the Mel spectrum of the context speech and the Mel spectrum of the target audio to obtain the global Mel spectrum; Extract the Mel spectra of the sub-audio segments from the target audio segment, which are divided according to the boundaries of the sub-word phonemes, as the segment Mel spectra; The global Mel spectrum is style-encoded to obtain a global audio style as the first residual style. The local Mel spectrum and the segment Mel spectrum are style-encoded to obtain local audio styles and segment audio styles, respectively. The global audio style is subtracted from the local audio style to obtain the second residual style. The local audio style is subtracted from the segment audio style to obtain the third residual style. The first residual style, the second residual style, and the third residual style are input into the corresponding style label layer to obtain the global audio style vector, the local audio style vector, and the segment audio style vector, respectively. Based on the global audio style vector, local audio style vector, and segment audio style vector, the total emotional style variable of the target audio is obtained as the first style embedding vector.

6. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the speech synthesis method based on multi-scale style as described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, wherein the computer program comprises the following steps of: receiving a request for a resource from a client; determining whether the client is authorized to access the resource; and if the client is authorized to access the resource, providing the resource to the client. When the computer program is executed by the processor, it implements the steps of the speech synthesis method based on multi-scale style as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Speech synthesis method and related device, electronic equipment and storage medium

    CN114283781A

  • Multi-language speech synthesis method and system based on hierarchical rhythm prediction

    CN115547293A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN116129862A