Speech synthesis method, model generation method
By predicting and adjusting speech features of the text before speech synthesis, the problem of a single style in existing speech synthesis models is solved, enabling flexible speech style adjustment and cost reduction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2022-09-22
- Publication Date
- 2026-04-24
AI Technical Summary
Existing speech synthesis models can only synthesize a single speech style, which cannot meet the diverse broadcasting needs, resulting in poor speech broadcasting performance in different scenarios. Furthermore, retraining speech synthesis models is costly and time-consuming.
By predicting the speech features of the text, the speech feature information of the target speech style is obtained, and the basic speech feature information of the speech synthesis model is adjusted using this information to generate speech in the target speech style.
It enables flexible adjustment of speech style without changing the TTS speaker's timbre, reduces the launch cycle and cost of adding new speech styles, and generates speech that matches the target speech style.
Smart Images

Figure CN115910028B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of speech synthesis technology, and in particular to a speech synthesis method, a training method for a speech feature prediction model, a speech broadcasting method, a computing device, and a computer storage medium. Background Technology
[0002] With the continuous expansion of the audio content market, voice broadcasting is being widely used in various scenarios.
[0003] In related technologies, speech synthesis models are typically used to achieve speech broadcasting. Specifically, the speech synthesis model takes text as input and synthesizes speech based on pre-configured basic speech feature information, such as speech rate, pitch, and volume, to obtain speech.
[0004] However, in the process of realizing the concept of this invention, the inventors discovered that the speech synthesis models in related technologies can only synthesize speech with a single speech style, which is not accurate enough. Summary of the Invention
[0005] This invention provides a speech synthesis method, a model generation method, a speech broadcasting method, an apparatus, a computing device, and a computer storage medium.
[0006] In a first aspect, embodiments of the present invention provide a speech synthesis method, comprising:
[0007] Obtain the text to be used for speech synthesis;
[0008] Speech feature prediction is performed on the text to obtain speech feature information of the target speech style;
[0009] The basic speech feature information of the speech synthesis model is adjusted using the aforementioned speech feature information to obtain the target speech feature information;
[0010] The text is input into the speech synthesis model to synthesize speech using the target speech feature information, thereby obtaining target speech in the target speech style.
[0011] Secondly, this invention provides a model generation method, including:
[0012] Obtain a speech sample set, which includes multiple speech samples belonging to the same target speech style;
[0013] Perform text recognition on the multiple speech samples respectively to determine the text samples corresponding to the multiple speech samples;
[0014] Feature extraction is performed on the multiple speech samples respectively to determine the speech sample feature information corresponding to each of the multiple speech samples;
[0015] A speech feature prediction model is trained using the text samples and speech sample feature information corresponding to the multiple speech samples respectively. The speech feature prediction model is used to predict the speech feature information of the text to be synthesized. The speech feature information is used to adjust the basic speech feature information of the speech synthesis model to obtain target speech feature information. The target speech feature information is used to synthesize speech from the input text to obtain target speech that matches the target speech style.
[0016] Thirdly, this invention provides a voice broadcasting method, including:
[0017] Get the text to be broadcast;
[0018] Speech feature prediction is performed on the text to be broadcast to obtain speech feature information of the target speech style;
[0019] The text to be played and the speech feature information are input into the speech synthesis model. The basic speech feature information of the speech synthesis model is adjusted based on the speech feature information to obtain the target speech feature information. The speech feature information is then used to synthesize the content to be played to obtain the broadcast speech with the target speech style.
[0020] Play the aforementioned audio message.
[0021] Fourthly, embodiments of the present invention provide a speech synthesis device, comprising:
[0022] The first text acquisition module is used to acquire the text to be used for speech synthesis;
[0023] The first feature prediction module is used to predict the speech features of the text and obtain speech feature information of the target speech style;
[0024] The first feature information adjustment module is used to adjust the basic speech feature information of the speech synthesis model using the speech feature information to obtain the target speech feature information.
[0025] The first speech synthesis module is used to input the text into the speech synthesis model, so as to use the target speech feature information to synthesize the text into speech and obtain the target speech in the target speech style.
[0026] Fifthly, embodiments of the present invention provide a model generation apparatus, comprising:
[0027] The sample acquisition module is used to acquire a speech sample set, which includes multiple speech samples belonging to the same target speech style.
[0028] The recognition module is used to perform text recognition on the plurality of speech samples respectively, and determine the text samples corresponding to the plurality of speech samples respectively;
[0029] The feature extraction module is used to extract features from the plurality of speech samples respectively, and determine the speech sample feature information corresponding to the plurality of speech samples respectively;
[0030] The training module is used to train a speech feature prediction model using the text samples and speech sample feature information corresponding to the multiple speech samples respectively. The speech feature prediction model is used to predict the speech feature information of the text to be synthesized. The speech feature information is used to adjust the basic speech feature information of the speech synthesis model to obtain target speech feature information. The target speech feature information is used to synthesize the input text to obtain target speech that matches the target speech style.
[0031] Sixthly, an embodiment of the present invention provides a voice broadcasting device, comprising:
[0032] The second text acquisition module is used to acquire the text to be broadcast.
[0033] The second feature prediction module is used to predict the speech features of the text to be broadcast in order to obtain speech feature information of the target speech style.
[0034] The second feature information adjustment module is used to input the text to be played and the speech feature information into the speech synthesis model, so as to adjust the basic speech feature information of the speech synthesis model based on the speech feature information to obtain target speech feature information, and use the target speech feature information to synthesize the speech of the content to be played to obtain the broadcast speech of the target speech style.
[0035] The voice playback module plays the broadcast voice message.
[0036] In a seventh aspect, an embodiment of the present invention provides a computing device, including a processing component and a storage component;
[0037] The storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement the speech synthesis method provided in the embodiments of the present invention, or to implement the model generation method provided in the embodiments of the present invention, or to implement the speech broadcasting method provided in the embodiments of the present invention.
[0038] Eighthly, in this embodiment of the invention, a computer storage medium is provided, storing a computer program. When the computer executes the program, it implements the speech synthesis method provided in this embodiment of the invention, or the model generation method provided in this embodiment of the invention, or the speech broadcasting method provided in this embodiment of the invention.
[0039] In this embodiment of the invention, before inputting the text to be synthesized into the speech synthesis model, the speech feature information of the text is first predicted to obtain the speech feature information of the target speech style. Further, the basic speech feature information of the speech synthesis model is adjusted using the obtained speech feature information of the target style to generate the target speech feature information. Then, the text is input into the speech synthesis model, so that the speech synthesis model uses the target speech feature information to synthesize the text and generate the target speech of the target style. Thus, the speech style output by the speech synthesis model is converted from the basic speech style corresponding to the basic speech feature information to the target speech style corresponding to the target speech feature information.
[0040] These or other aspects of the invention will become more apparent from the following description of the embodiments. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 A system architecture diagram is shown that can be applied to the technical solution of the embodiments of the present invention;
[0043] Figure 2 The flowchart illustrating a speech synthesis method provided in one embodiment of the present invention is shown in the schematic diagram.
[0044] Figure 3 The diagram illustrates a speech synthesis method provided in an embodiment of the present invention.
[0045] Figure 4 The flowchart illustrating a model generation method provided in one embodiment of the present invention is shown in the schematic diagram.
[0046] Figure 5 A waveform diagram is schematically shown in one embodiment of the present invention;
[0047] Figure 6 The flowchart illustrating a voice broadcasting method provided in one embodiment of the present invention is shown in the schematic diagram.
[0048] Figure 7 This schematic diagram illustrates a system in which the voice broadcasting method provided in the embodiments of the present invention can be applied;
[0049] Figure 8The diagram schematically illustrates a block diagram of a speech synthesis apparatus provided in one embodiment of the present invention;
[0050] Figure 9 The diagram schematically illustrates a block diagram of a model generation apparatus provided in one embodiment of the present invention;
[0051] Figure 10 The diagram schematically illustrates a block diagram of a voice broadcasting device provided in one embodiment of the present invention;
[0052] Figure 11 A block diagram of a computing device provided in one embodiment of the present invention is shown schematically. Detailed Implementation
[0053] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0054] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0055] Text-to-Speech (TTS) technology generally converts text into speech. In related technologies, speech synthesis typically involves first inviting voice actors to record high-quality speech samples in a professional recording studio, obtaining the corresponding text from the speech samples, and then training a neural network model. The input is text, and the output is speech, thus obtaining a trained speech synthesis model.
[0056] TTS speakers typically refer to the timbre and pronunciation style of the speech synthesized by a speech synthesis model. Speech synthesis models trained using speech libraries recorded by different voice actors have different basic speech features, such as speech rate, pitch, and volume, and their timbre and pronunciation style will also be different.
[0057] In the process of realizing the concept of this invention, the inventors discovered that TTS speakers in related technologies have a certain pronunciation style, which is often only applicable to specific scenarios. For example, in the basic speech feature information of most existing speech synthesis models, the speech rate is slow, the tone is flat, and the volume is low. Correspondingly, when TTS speakers are broadcasting, their pronunciation is relatively flat, and the voice lacks fluctuation. This pronunciation style is often only applicable to scenarios with short broadcast durations. In scenarios with long broadcast durations, this lack of intonation can easily cause auditory fatigue for users.
[0058] To imbue TTS (Text-to-Speech) speakers with rhythm, voice actors can record all texts with various emotions during voice library recording. The speech synthesis model is then trained using a voice library with multiple emotional styles. This method improves the expressiveness of the TTS speaker's voice, making the pronunciation more varied and rhythmic. However, this optimization method drastically increases the amount of training data required for the speech synthesis model, leading to higher costs and longer deployment cycles. Furthermore, while this optimization method enhances the expressiveness of the TTS speaker, it still limits its ability to deliver speech in only a single voice style.
[0059] As the audio content market continues to expand, voice broadcasting is widely used in various scenarios. Consequently, TTS speakers need to adopt different voice styles for different broadcasting scenarios, which cannot be achieved by directly using a speech synthesis model.
[0060] In related technologies, the following methods are commonly used to enrich the pronunciation style of TTS speakers:
[0061] 1. When a specific speech style is required for broadcasting, the voice actor records according to that characteristic speech style to generate a speech library of that style. This speech library is then used to train a speech synthesis model for that style. This method produces the best pronunciation results, but adding a new speech style requires retraining the speech synthesis model. Retraining the speech synthesis model requires collecting training samples again, which is costly and time-consuming, making it impossible to meet the needs of adding new speech styles in a timely manner.
[0062] 2. When a specific speech style is required for speech broadcasting, the prosody of the text is first manually adjusted using SSML (Speech Synthesis Markup Language), and then the adjusted text is input into the speech synthesis model. However, since the adjustment parameters in SSML lack a basis for setting, they can only be adjusted based on the engineer's experience, which is not only inefficient but also has poor results.
[0063] To address the technical challenge of understanding the speech style of newly added TTS speakers in related technologies, this invention provides a speech synthesis method. Before inputting the text to be synthesized into a speech synthesis model, the method first predicts the speech features of the text to obtain the speech features of the target speech style. Then, it adjusts the basic speech features of the speech synthesis model using the obtained target speech features to generate the target speech features. The text is then input into the speech synthesis model, allowing the model to synthesize the text using the target speech features and generate the target speech in the target style. Thus, the speech style output by the speech synthesis model is transformed from the basic speech style corresponding to the basic speech features to the target speech style corresponding to the target speech features.
[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0065] Figure 1 A system architecture diagram is shown that can be applied to the technical solution of the embodiments of the present invention. This system architecture may include a user terminal 101 and a server terminal 102. (Alternatively, it may include a server terminal and multiple user terminals; or it may include a server terminal, a first user terminal, and a second user terminal, etc.)
[0066] In this system, the user terminal 101 and the server terminal 102 establish a connection via a network. The network provides a communication link medium between the user terminal 101 and the server terminal 102. The network can include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0067] Client 101 can interact with server 102 via the network to receive or send messages, etc.
[0068] The user terminal 101 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. The user terminal 101 can be deployed on an electronic device and depends on the device to run or on certain apps within the device. Electronic devices can have displays and support information browsing, such as personal mobile terminals like smartphones, tablets, and personal computers. For ease of understanding, Figure 1 The user end is primarily represented by the device itself. Various other types of applications can also be configured on electronic devices, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platforms.
[0069] The server 102 may include servers that provide various services, such as a server for backend training that supports the model used on the client 101, or a server that processes interactive information sent by the client.
[0070] It should be noted that server 102 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0071] It should be noted that the speech synthesis method, model generation method, and speech broadcasting method provided in the embodiments of the present invention can be executed by the server 102, and the corresponding speech synthesis device, model generation device, and speech broadcasting device are respectively set in the server 102. However, in other embodiments of the present invention, the user terminal 101 may also have similar functions to the server 102, thereby executing the speech synthesis method, model generation method, and speech broadcasting method provided in the embodiments of the present invention. In other embodiments, the speech synthesis method, model generation method, and speech broadcasting method provided in the embodiments of the present invention may also be jointly executed by the user terminal 101 and the server 102.
[0072] It should be understood that Figure 1 The number of client and server instances shown is merely illustrative. Depending on implementation needs, there can be any number of client and server instances.
[0073] The implementation details of the technical solutions of the embodiments of the present invention will be described in detail below.
[0074] Figure 2 The schematic diagram illustrates a flowchart of a speech synthesis method according to an embodiment of the present invention, which may include the following steps:
[0075] 201, Obtain the text to be used for speech synthesis.
[0076] 202. Perform speech feature prediction on the text to obtain speech feature information of the target speech style.
[0077] 203. Adjust the basic speech feature information of the speech synthesis model using speech feature information to obtain the target speech feature information.
[0078] 204. Input the text into the speech synthesis model to synthesize the text using the target speech feature information and obtain the target speech in the target speech style.
[0079] According to embodiments of the present invention, the text to be synthesized can include any text in any language. For example, the language can include Chinese, English, Russian, German, etc. The text can include news text, entertainment text, sports commentary text, etc.
[0080] According to embodiments of the present invention, text to be synthesized can be obtained based on user input. The text input method may include, for example, voice input, input by character recognition of physical text, or input based on peripheral devices.
[0081] According to embodiments of the present invention, different texts usually require different speech styles for their playback, and the same text may also require different speech styles for playback in different language environments. Therefore, the speech features of the text can be predicted based on the text content and the desired target speech style to obtain the speech feature information of the target speech style.
[0082] According to embodiments of the present invention, the basic speech feature information may include feature parameters that influence the pronunciation style of a TTS speaker, generated after training a speech synthesis model using samples. By changing the values of the feature parameters in the basic speech feature information, the pronunciation style of the corresponding TTS speaker can be changed.
[0083] According to embodiments of the present invention, the implementation method of adjusting basic speech feature information using speech feature information may include, for example, replacing basic speech feature information with speech feature information to use speech feature information as target speech feature information, or adjusting the values of basic speech features using speech feature information based on basic speech feature information to obtain target speech feature information.
[0084] According to embodiments of the present invention, after inputting speech feature information into a speech synthesis model to adjust the basic speech feature information of the speech synthesis model to obtain target speech feature information, text can be input into the speech synthesis model, so that the speech synthesis model performs speech synthesis on the text based on the adjusted target speech feature information to generate target speech in the target style. However, this is not limited to this; speech feature information and text can also be input together into the speech synthesis model. After input, the basic speech information can first be adjusted using the speech feature information to generate target speech feature information, and then the text can be synthesized using the target speech feature information to output the target speech.
[0085] According to an embodiment of the present invention, target speech feature information can indicate the pronunciation mode of each unit to be pronounced in the text. Based on this, in one embodiment of the present invention, the speech synthesis model performs speech synthesis on the text based on the target speech feature information as follows:
[0086] Find the target pronunciation unit corresponding to each pronunciation unit in the text from multiple pronunciation units stored in the speech library;
[0087] For any unit to be pronounced, based on the target speech feature information corresponding to the unit to be pronounced, the target pronunciation mode that matches the target speech feature information is determined from the multiple pronunciation modes of the target pronunciation unit corresponding to the unit to be pronounced;
[0088] Multiple target pronunciation units that have already matched the target pronunciation mode are concatenated to obtain the target speech.
[0089] According to embodiments of the present invention, a speech database can store multiple pronunciation units. The pronunciation of these pronunciation units can be derived from the segmentation of natural speech. After segmentation, these pronunciation units are labeled with pronunciation symbols (including pronunciation markers, voiced / voiced segmentation, etc.). To obtain a more ideal synthesis effect, the speech database can store multiple pronunciations with different prosody (e.g., different tones, different emotions) for any pronunciation unit. Thus, after determining the target speech feature information corresponding to each pronunciation unit in the text, the speech database can be searched to generate the target speech.
[0090] In another embodiment of the present invention, the speech synthesis model performing speech synthesis on text based on target speech feature information can be specifically implemented as follows:
[0091] Based on the target speech feature information, generate the Mel spectrum corresponding to the text;
[0092] Convert the Mel spectrum into a waveform;
[0093] Use the waveform as the target speech.
[0094] According to an embodiment of the present invention, predicting speech features of text to obtain speech feature information of the target speech style can be specifically implemented as follows:
[0095] The speech feature prediction model is used to predict the speech features of the text to obtain the speech feature information of the target speech style; the speech feature prediction model is trained based on multiple speech samples belonging to the target speech style.
[0096] According to an embodiment of the present invention, the speech feature prediction model may include a neural network model trained using multiple speech samples, wherein the multiple speech samples may belong to the same target speech style.
[0097] According to an embodiment of the present invention, before training the speech feature prediction model, the speech style can be classified according to the application scenario. For example, the speech style can be divided into live-streaming e-commerce speech style, sports commentary speech style, news broadcasting speech style, human-computer interaction speech style, etc.
[0098] According to embodiments of the present invention, multiple speech samples belonging to a target speech style can be obtained by capturing recordings of the corresponding speech style. For example, speech samples in the style of news broadcasts can be obtained by capturing the audio or video of news programs. When video is captured, the audio in the video can be extracted, and the extracted audio can be used as a speech sample.
[0099] According to embodiments of the present invention, training samples for the speech feature prediction model can be obtained by crawling recordings belonging to the target speech style from the Internet. This eliminates the need for manual recording of the speech database and annotation of sample data, thereby reducing the training cost of the speech feature prediction model. Furthermore, the speech feature information predicted by the speech feature prediction model can be provided to multiple speech synthesis models, thus allowing the pronunciation style of the TTS speaker to be changed without altering the speaker's timbre.
[0100] According to an embodiment of the present invention, when it is necessary to change the speech style of a TTS speaker, the desired speech style can be determined first, and then audio corresponding to the desired speech style can be crawled from the Internet to obtain multiple speech samples. After obtaining the speech samples, a speech feature prediction model can be trained using the speech samples. Since the speech feature prediction model is trained using speech samples belonging to the same style, the speech feature prediction model has a strong feature extraction capability for that speech style, and can thus predict speech feature information matching that style from the text.
[0101] According to an embodiment of the present invention, by predicting the speech feature information of the text before using the speech synthesis model to synthesize the text, the pronunciation style of the target speech synthesized by the speech synthesis model is changed. This eliminates the need for recording studios and professional voice actors to record speech libraries, and also eliminates the need to retrain the speech synthesis model, thus shortening the launch cycle of new speech styles.
[0102] According to an embodiment of the present invention, the speech feature prediction model is used to predict the speech features of text to obtain the speech feature information of the target speech style, which can be specifically implemented as follows:
[0103] The text is segmented according to a preset segmentation granularity to obtain multiple segmentation units;
[0104] Multiple segmentation units are input into the speech feature prediction model, and the output is speech feature information corresponding to each segmentation unit.
[0105] According to embodiments of the present invention, the preset segmentation granularity may include one or more of character granularity, word granularity, phoneme granularity, and sentence granularity. Combining different segmentation granularities can improve the flexibility of prosodic control of the target speech, thereby making the target speech pronunciation more natural.
[0106] According to embodiments of the present invention, when performing granular segmentation of text, the finer the segmentation granularity, for example, if all text is segmented at the phoneme level, the greater the adjustable space of the prosody of the target speech. However, correspondingly, the difficulty of speech feature prediction also increases. Therefore, the preset segmentation granularity can be selected by those skilled in the art according to actual application needs, and the embodiments of the present invention do not specifically limit the preset segmentation granularity.
[0107] Specifically, adjusting the basic speech feature information of the speech synthesis model using speech feature information to obtain the target speech feature information can be achieved as follows:
[0108] By utilizing the speech feature information corresponding to multiple segmentation units, the basic speech feature information of the speech synthesis model is adjusted to obtain the target speech feature information.
[0109] According to an embodiment of the present invention, the basic speech feature information can be a general speech feature, that is, for any input text, the same basic speech feature information is used for speech synthesis. This is also one of the reasons why the speech generated by the speech synthesis model based on the basic speech feature information lacks variation and sounds mechanical.
[0110] According to an embodiment of the present invention, by inputting multiple segmentation units into a speech feature prediction model, speech features corresponding to each segmentation unit can be output. Typically, the speech features corresponding to each segmentation unit are different. Based on this, for each segmentation unit, the basic speech features can be adjusted based on its corresponding speech features to generate sub-target speech feature information corresponding to each segmentation unit. Then, the multiple sub-target speech feature information can be used together as the target speech feature information.
[0111] According to an embodiment of the present invention, the speech feature information includes at least one speech feature value.
[0112] According to an embodiment of the present invention, the target speech feature information is obtained by adjusting the basic speech feature information of the speech synthesis model using the speech feature information corresponding to multiple segmentation units respectively.
[0113] At least one speech feature value corresponding to each segmentation unit is arcregularized to determine multiple target speech feature values;
[0114] The average value of the speech feature is determined by averaging multiple target speech feature values.
[0115] Based on the average speech feature value and the target speech feature value corresponding to each segmentation unit, the speech feature ratio value corresponding to each segmentation unit is determined;
[0116] The speech feature ratio value corresponding to each segmentation unit is input into the speech synthesis model. By using the speech feature ratio value corresponding to each segmentation unit, the basic speech feature information of the speech synthesis model is adjusted to obtain the target speech feature information.
[0117] According to an embodiment of the present invention, determining the speech feature ratio value corresponding to each segmentation unit based on the average speech feature value and the target speech feature value corresponding to each segmentation unit can be specifically implemented as follows:
[0118] Divide the target speech feature value corresponding to each segmentation unit by the average speech feature value to generate the speech feature ratio value corresponding to each segmentation unit.
[0119] According to an embodiment of the present invention, the proportion of speech features corresponding to each segmentation unit can be represented by the following formulas (1) and (2).
[0120]
[0121]
[0122] Among them, V avg V represents the average speech feature value. scale V represents the proportion of speech features, N represents the number of segmentation units included in the text, and V represents the proportion of speech features. 反正则 This represents the feature value of the target speech.
[0123] According to embodiments of the present invention, for speech, there can be multiple dimensions of features that can affect speech style, such as pitch, speech rate, volume, etc., and each speech feature value can represent a dimension that can affect speech style.
[0124] According to embodiments of the present invention, the speech feature ratio value can be calculated sequentially for each speech feature value included in the speech feature information. For example, the speech feature information includes pitch feature values and volume feature values. Therefore, the pitch feature values of each segmentation unit in the text can first be arcregulated to generate a target pitch feature value; then, the pitch feature values corresponding to each segmentation unit in the text are counted; then, the multiple pitch feature values are summed and divided by the number of segmentation units to obtain an average pitch feature value; finally, the average pitch feature value and the target pitch feature value are used to determine the pitch ratio value corresponding to each segmentation unit. Similar to the pitch feature value, the volume ratio value corresponding to each segmentation unit can be determined using the same process.
[0125] According to an embodiment of the present invention, the pitch ratio value and the volume ratio value can be calculated in parallel or in sequence.
[0126] According to an embodiment of the present invention, the calculation method of antiregulation can be the inverse operation of regularization. The calculation method of regularization depends on the regularization calculation method adopted for the training samples during the preprocessing process when training the speech feature reservation model.
[0127] According to an embodiment of the present invention, after the regularization calculation method is determined, the calculation method can be derived so that when antiregulation is required, the derived regularization calculation method can be obtained, thereby determining the antiregulation calculation method. For example, the regularization calculation method can be represented by the following formula (3).
[0128]
[0129] Therefore, the calculation method of the antiregulation can be expressed by the following formula (4).
[0130] V 反正则 =V·V STD +V Mean (4)
[0131] Where V can represent the sample value of the sample, V Mean V can represent the mean of multiple samples. STD It can represent the standard deviation of multiple samples.
[0132] According to an embodiment of the present invention, the speech feature information includes at least one speech feature value.
[0133] According to an embodiment of the present invention, the target speech feature information is obtained by adjusting the basic speech feature information of the speech synthesis model using the speech feature information corresponding to multiple segmentation units respectively.
[0134] At least one speech feature value corresponding to each of the multiple segmentation units is input into the speech synthesis model, so as to adjust the basic speech feature information of the speech synthesis model by using the multiple speech feature values corresponding to the multiple segmentation units, and obtain the target speech feature information.
[0135] According to an embodiment of the present invention, the basic speech feature information corresponding to the speech feature information includes at least one basic speech feature value.
[0136] According to an embodiment of the present invention, after inputting at least one speech feature value corresponding to multiple segmentation units into the speech synthesis model, the corresponding basic speech feature values can be adjusted using the speech feature values.
[0137] For example, speech feature information includes pitch feature values and volume feature values. Correspondingly, basic speech feature information can include basic pitch feature values and basic volume feature values. Thus, after inputting pitch feature values into the speech synthesis model, the basic pitch feature values can be adjusted using the pitch feature values to generate target pitch feature values. Similarly, after inputting volume feature values into the speech synthesis model, the basic volume feature values can be adjusted using the volume feature values to generate target volume feature values.
[0138] In one embodiment of the present invention, after extracting the speech feature values corresponding to each segmentation unit of the text using the speech feature prediction model, the speech feature values can be directly input into the speech synthesis model. In another embodiment of the present invention, the speech feature values can be preprocessed first to convert them into corresponding speech feature ratio values, and then the speech feature ratio values can be input into the speech synthesis model.
[0139] According to an embodiment of the present invention, by converting speech feature values into speech feature ratio values and inputting them into the speech synthesis model, the speech synthesis model can be adjusted proportionally based on the basic speech feature values, making the generated target speech more natural and harmonious.
[0140] According to an embodiment of the present invention, adjusting the basic speech feature information of the speech synthesis model by utilizing the speech feature ratio value corresponding to each segmentation unit to obtain the target speech feature information can be specifically implemented as follows:
[0141] The speech feature ratio value corresponding to each segmentation unit is multiplied by the basic speech feature information to generate the first speech feature information corresponding to each segmentation unit.
[0142] Each first speech feature is added to the basic speech feature to determine the target speech feature.
[0143] According to an embodiment of the present invention, multiplying the speech feature ratio value corresponding to each segmentation unit with the basic speech feature information to generate the first speech feature information corresponding to each segmentation unit can be specifically implemented as follows:
[0144] Multiply the speech feature ratio value by the corresponding basic speech feature value to generate the first speech feature value corresponding to each segmentation unit;
[0145] Multiple first speech feature values are used as first speech feature information.
[0146] According to an embodiment of the present invention, for the dimension of speech rate, for example, the basic speech rate feature value is 100, and the speech rate feature ratio value is 0.3. Therefore, the first speech rate feature value obtained by multiplying the basic speech rate feature value and the speech rate feature ratio value is 30. The first speech rate feature value can characterize the magnitude of adjustment that needs to be made based on the basic speech rate feature value. After obtaining the first speech rate feature value, the first speech rate feature value can be added to the basic speech rate feature value to determine the target speech rate feature value as 130.
[0147] Figure 3 The diagram illustrates a speech synthesis method provided in an embodiment of the present invention.
[0148] like Figure 3 As shown, after obtaining the text to be synthesized, the text can be segmented according to a preset segmentation granularity to generate multiple segmentation units. For example, the text can be "very worth buying", and by segmenting the text, three segmentation units can be generated: "very", "worth", and "buy".
[0149] After granular segmentation, the three segmentation units can be input into the speech feature prediction model, and the output will be speech feature information corresponding to the three segmentation units respectively. The speech feature information can include three speech feature values, namely pitch feature value, speech rate feature value and volume feature value, which can be represented as T = [T1(t1, t2, t3), T2(t1, t2, t3), T3(t1, t2, t3)], where T represents text, T1, T2, and T3 each represent a segmentation unit, and t1, t2, and t3 each represent a speech feature value of the segmentation unit.
[0150] Furthermore, prosodic amplitude adjustment can be performed. Specifically, the three speech feature values corresponding to the three segmentation units can be converted into speech feature ratio values to generate T = [T1(s1, s2, s3), T2(s1, s2, s3), T3(s1, s2, s3)], where s1, s2, and s3 each represent a speech feature ratio value.
[0151] Furthermore, the speech feature ratio values corresponding to the three segmentation units can be input into the speech synthesis model so that the corresponding basic speech feature values can be adjusted using the speech feature ratio values to generate the target speech feature ratio values.
[0152] Finally, the three segmentation units can be input into the speech synthesis model to output the target speech.
[0153] Figure 4 The schematic diagram illustrates a flowchart of a model generation method according to an embodiment of the present invention, which may include the following steps:
[0154] 401. Obtain a speech sample set, which includes multiple speech samples belonging to the same target speech style.
[0155] 402. Perform text recognition on multiple speech samples to determine the text samples corresponding to each speech sample.
[0156] 403. Perform feature extraction on multiple speech samples respectively to determine the speech sample feature information corresponding to each of the multiple speech samples.
[0157] 404. A speech feature prediction model is trained using the text samples and speech sample feature information corresponding to multiple speech samples. The speech feature prediction model is used to predict the speech feature information of the text to be synthesized. The speech feature information is used to adjust the basic speech feature information of the speech synthesis model to obtain the target speech feature information. The target speech feature information is used to synthesize the input text to obtain the target speech that matches the target speech style.
[0158] According to an embodiment of the present invention, the speech feature prediction model may include a neural network model trained using multiple speech samples, wherein the multiple speech samples may belong to the same target speech style.
[0159] According to an embodiment of the present invention, before training the speech feature prediction model, the speech style can be classified according to the application scenario. For example, the speech style can be divided into live-streaming e-commerce speech style, sports commentary speech style, news broadcasting speech style, human-computer interaction speech style, etc.
[0160] According to embodiments of the present invention, the target speech style can correspond to the text type. For example, when the text is a product description for live-streaming e-commerce, the target speech style can be the live-streaming e-commerce speech style; when the text is a news broadcast script, the target speech style can be the news broadcast speech style.
[0161] According to embodiments of the present invention, multiple speech samples belonging to a target speech style can be obtained by capturing recordings of the corresponding speech style. For example, speech samples in the style of news broadcasts can be obtained by capturing the audio or video of news programs. When video is captured, the audio in the video can be extracted, and the extracted audio can be used as a speech sample.
[0162] According to embodiments of the present invention, training samples for the speech feature prediction model can be obtained by crawling recordings belonging to the target speech style from the Internet. This eliminates the need for manual recording of the speech database and annotation of sample data, thereby reducing the training cost of the speech feature prediction model. Furthermore, the speech feature information predicted by the speech feature prediction model can be provided to multiple speech synthesis models, thus allowing the pronunciation style of the TTS speaker to be changed without altering the speaker's timbre.
[0163] According to an embodiment of the present invention, when it is necessary to change the speech style of a TTS speaker, the desired speech style can be determined first, and then audio corresponding to the desired speech style can be crawled from the Internet to obtain multiple speech samples. After obtaining the speech samples, a speech feature prediction model can be trained using the speech samples. Since the speech feature prediction model is trained using speech samples belonging to the same style, the speech feature prediction model has a strong feature extraction capability for that speech style, and can thus predict speech feature information matching that style from the text.
[0164] According to an embodiment of the present invention, by predicting the speech feature information of the text before using the speech synthesis model to synthesize the text, the pronunciation style of the target speech synthesized by the speech synthesis model is changed. This eliminates the need for recording studios and professional voice actors to record speech libraries, and also eliminates the need to retrain the speech synthesis model, thus shortening the launch cycle of new speech styles.
[0165] According to an embodiment of the present invention, after performing text recognition on multiple speech samples in a speech sample set and determining the text samples corresponding to the multiple speech samples respectively, the model generation method further includes:
[0166] For any given text sample, the text sample is segmented according to a preset segmentation granularity to obtain multiple segmentation units.
[0167] According to embodiments of the present invention, the preset segmentation granularity may include one or more of character granularity, word granularity, phoneme granularity, and sentence granularity.
[0168] According to embodiments of the present invention, when performing granular segmentation of text, the finer the segmentation granularity, for example, if all text is segmented using factor granularity, the greater the adjustable space of the prosody of the target speech. However, correspondingly, the difficulty of speech feature prediction also increases. Therefore, the preset segmentation granularity can be selected by those skilled in the art according to actual application needs, and the embodiments of the present invention do not specifically limit the preset segmentation granularity.
[0169] The specific implementation of feature extraction for multiple speech samples in a speech sample set, to determine the speech sample feature information corresponding to each of the multiple speech samples, can be as follows:
[0170] For any given text sample, feature extraction is performed on the multiple segmentation units contained in the text sample to determine the speech sample feature information corresponding to each segmentation unit.
[0171] Training a speech feature prediction model using text samples and speech feature information corresponding to multiple speech samples can be specifically implemented as follows:
[0172] The speech feature prediction model is trained by taking multiple segmentation units of the text sample corresponding to any speech sample as model input and the speech sample feature information corresponding to the speech sample as model label.
[0173] According to an embodiment of the present invention, for any given text sample, feature extraction is performed on multiple segmentation units contained in the text sample to determine the speech sample feature information corresponding to each of the multiple segmentation units. Specifically, this can be achieved as follows:
[0174] For any segmentation unit in any text sample, the time axis of the segmentation unit is aligned with the corresponding speech sample in order to determine the time label of the segmentation unit;
[0175] Obtain the waveform of the speech sample;
[0176] Based on the time tags of the segmentation units, the waveforms corresponding to the segmentation units are determined from the waveform diagrams of the speech samples;
[0177] The speech sample feature information corresponding to the segmentation unit is determined based on the waveform.
[0178] According to an embodiment of the present invention, the segmentation unit is time-aligned on the speech sample, that is, the start time and end time of the pronunciation of the segmentation unit in the speech sample are determined, so that the start time and end time together serve as time labels corresponding to the segmentation unit.
[0179] Figure 5 A waveform diagram of one embodiment of the present invention is illustrated schematically.
[0180] Figure 5This can be a waveform of the speech sample corresponding to the text sample "Very worth buying". In this waveform, the horizontal axis represents time and the vertical axis represents amplitude.
[0181] The text sample "Very worth buying" can be divided into three segments: "Very", "Worth Buying", and "Buy" according to a preset segmentation granularity. By aligning the segments with the timeline, the start time of the segment "Very" can be determined as point a, the end time as point b, the start time of the segment "Worth Buying" as point c, the end time as point d, and the start time of the segment "Buy" as point e, and the end time as point f.
[0182] Based on this, it can be determined that in the waveform diagram, the waveform between point a and point b corresponds to the "very" unit, the waveform between point c and point d corresponds to the "worth" unit, and the waveform between point e and point f corresponds to the "buy" unit.
[0183] According to an embodiment of the present invention, determining the speech sample feature information corresponding to the segmentation unit based on the waveform can be specifically implemented as follows:
[0184] The speech rate characteristic value of the segmentation unit is determined based on the start and end times of the waveform;
[0185] The fundamental frequency of the waveform is extracted to determine the pitch characteristic value of the segmentation unit;
[0186] The volume characteristic value of the segmentation unit is determined based on the amplitude of the waveform;
[0187] Speech sample feature information is generated based on speech rate feature value, pitch feature value, and volume feature value.
[0188] According to an embodiment of the present invention, by subtracting the start time from the end time, the time required for the segmentation unit to be pronounced in the speech sample can be determined, and this time can be used as a speech rate feature value.
[0189] According to an embodiment of the present invention, the fundamental frequency extraction of the waveform can be performed using the time-domain method or frequency-domain method in the prior art, which will not be elaborated here.
[0190] According to an embodiment of the present invention, training a speech feature prediction model using the feature information of text samples and speech samples corresponding to multiple speech samples can be specifically implemented as follows:
[0191] For any given speech sample, input the corresponding text sample into the speech feature prediction model and output the predicted speech feature value.
[0192] Calculate the loss result between the predicted speech feature value and the speech feature value corresponding to the speech sample;
[0193] Adjust the model parameters of the speech feature prediction model based on the loss results until the loss results meet the preset conditions.
[0194] According to embodiments of the present invention, the calculation of the loss result can be achieved using a loss function, which may include, for example, a cross-entropy loss function, a squared loss function, etc.
[0195] According to an embodiment of the present invention, the preset conditions may include, for example, a loss function less than a preset threshold, or an iteration number greater than a preset round threshold.
[0196] Figure 6 The schematic diagram illustrates a flowchart of a voice broadcasting method according to an embodiment of the present invention, which may include the following steps:
[0197] 601, retrieve the text to be broadcast;
[0198] 602. Predict speech features of the text to be broadcast to obtain speech feature information of the target speech style;
[0199] 603. Input the text to be played and the speech feature information into the speech synthesis model, adjust the basic speech feature information of the speech synthesis model based on the speech feature information to obtain the target speech feature information, and use the target speech feature information to synthesize the speech of the content to be played to obtain the broadcast speech with the target speech style.
[0200] 604, play the announcement audio.
[0201] According to embodiments of the present invention, the method of obtaining the text to be broadcast can differ in different application scenarios. For example, in a live-streaming e-commerce scenario, the user can write the product description text, and then the device can receive the product description text uploaded by the user and use the received product description text as the text to be broadcast. For example, in a voice interaction scenario, the terminal device can receive questions or instructions raised by the user via voice. In response to receiving the questions or instructions input by the user, the terminal device can generate a response text, which can then be used as the text to be broadcast.
[0202] Figure 7 The diagram illustrates a system schematic that can apply the voice broadcasting method provided in the embodiments of the present invention.
[0203] like Figure 7 As shown, the system may include a terminal device 701 and a server 702.
[0204] Terminal device 701 can receive text 703 input by the user to be played. After acquiring the text 703, terminal device 701 can perform certain preprocessing on the text 703. Preprocessing may include, for example, recognizing the text 703 to determine the corresponding speech style. Recognizing the text 703 can be achieved by keyword recognition or semantic recognition.
[0205] After the terminal device 701 recognizes the text 703 to be broadcast, it can send the text 703 to be broadcast and the recognition result to the server 702.
[0206] After receiving the text to be broadcast 703 and the recognition result, the server 702 can determine the target speech feature prediction model 705 corresponding to the recognition result from multiple pre-trained speech feature prediction models 704 based on the recognition result. The text to be broadcast 703 is then input into the target speech feature prediction model 705 for speech feature prediction. Each of the multiple speech feature prediction models 704 can correspond to a different speech style. For example, the multiple speech feature prediction models 704 may include a speech feature prediction model for news broadcasting, a speech feature prediction model for sports commentary, and a speech feature prediction model for live broadcasting. Since the recognition result of the text to be broadcast 703 represents the text as a news broadcast script, the speech feature prediction model for the news broadcasting style can be determined as the target speech feature prediction model 705.
[0207] After the target speech feature prediction model outputs speech feature information, the server 702 can input the speech feature information into the speech synthesis model 706, thereby using the speech feature information to adjust the basic speech feature information and generate target speech feature information. Furthermore, the text to be played can be input into the speech synthesis model 706 to output the target speech 707.
[0208] After generating the target voice 707, the server can send the target voice 707 to the terminal device 701, so that the terminal device 701 can broadcast the target voice 707.
[0209] It should be noted that the speech feature prediction model and the speech synthesis model can also be deployed in the terminal device. Thus, after receiving the text to be played, the terminal device can use the speech feature prediction model to predict the speech feature information locally, and use the speech synthesis model to synthesize the target speech.
[0210] According to an embodiment of the present invention, Figure 7 The specific implementation details of the voice broadcasting method shown are as follows: Figure 2The speech synthesis methods shown are the same or similar; please refer to [link / reference]. Figure 2 The description of the speech synthesis method shown will not be repeated here.
[0211] Figure 8 The schematic diagram illustrates a block diagram of a speech synthesis device provided in one embodiment of the present invention, such as... Figure 8 As shown, the speech synthesis device 800 includes a first text acquisition module 801, a first feature prediction module 802, a first feature information adjustment module 803, and a first speech synthesis module 804.
[0212] The first text acquisition module 801 is used to acquire the text to be used for speech synthesis;
[0213] The first feature prediction module 802 is used to predict the speech features of the text and obtain speech feature information of the target speech style.
[0214] The first feature information adjustment module 803 is used to adjust the basic speech feature information of the speech synthesis model using the speech feature information to obtain the target speech feature information.
[0215] The first speech synthesis module 804 is used to input the text into the speech synthesis model, so as to use the target speech feature information to synthesize the text into speech and obtain the target speech in the target speech style. Figure 8 The aforementioned speech synthesis device can perform Figure 2 The implementation principle and technical effects of the speech synthesis method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the speech synthesis device in the above embodiments performs its operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0216] According to an embodiment of the present invention, the first feature prediction module 802 is specifically used for:
[0217] The speech feature prediction model is used to predict the speech features of the text to obtain the speech feature information of the target speech style; wherein, the speech feature prediction model is trained based on multiple speech samples belonging to the target speech style.
[0218] According to an embodiment of the present invention, the first feature prediction module 802 is specifically used for:
[0219] The text is segmented according to a preset segmentation granularity to obtain multiple segmentation units;
[0220] The multiple segmentation units are input into the speech feature prediction model, and speech feature information corresponding to each of the multiple segmentation units is output.
[0221] The process of adjusting the basic speech feature information of the speech synthesis model using the aforementioned speech feature information to obtain the target speech feature information includes:
[0222] By utilizing the speech feature information corresponding to the multiple segmentation units, the basic speech feature information of the speech synthesis model is adjusted to obtain the target speech feature information.
[0223] According to an embodiment of the present invention, the speech feature information includes at least one speech feature value.
[0224] According to an embodiment of the present invention, the first feature information adjustment module 803 is specifically used for:
[0225] At least one speech feature value corresponding to each segmentation unit is arcregularized to determine multiple target speech feature values;
[0226] The average value of the multiple target speech feature values is calculated to determine the average speech feature value.
[0227] Based on the average speech feature value and the target speech feature value corresponding to each segmentation unit, a speech feature ratio value corresponding to each segmentation unit is determined;
[0228] The speech feature ratio value corresponding to each segmentation unit is input into the speech synthesis model. By using the speech feature ratio value corresponding to each segmentation unit, the basic speech feature information of the speech synthesis model is adjusted to obtain the target speech feature information.
[0229] According to an embodiment of the present invention, the speech feature information includes at least one speech feature value.
[0230] According to an embodiment of the present invention, the first feature information adjustment module 803 is specifically used for:
[0231] At least one speech feature value corresponding to each of the multiple segmentation units is input into the speech synthesis model, so as to adjust the basic speech feature information of the speech synthesis model using the multiple speech feature values corresponding to the multiple segmentation units, and obtain the target speech feature information.
[0232] According to an embodiment of the present invention, the first feature information adjustment module 803 is specifically used for:
[0233] The speech feature ratio value corresponding to each segmentation unit is multiplied by the basic speech feature information to generate the first speech feature information corresponding to each segmentation unit.
[0234] Each of the first speech feature information is added to the basic speech feature information to determine the target speech feature information.
[0235] According to an embodiment of the present invention, the first feature information adjustment module 803 is specifically used for:
[0236] Divide the target speech feature value corresponding to each segmentation unit by the average speech feature value to generate the speech feature ratio value corresponding to each segmentation unit.
[0237] Figure 9 The schematic diagram illustrates a block diagram of a model generation apparatus provided in one embodiment of the present invention, such as... Figure 9 As shown, the model generation device 900 includes a sample acquisition module 901, a recognition module 902, a feature extraction module 903, and a training module 904.
[0238] The sample acquisition module 901 is used to acquire a speech sample set, which includes multiple speech samples belonging to the same target speech style.
[0239] The recognition module 902 is used to perform text recognition on the plurality of speech samples respectively, and determine the text samples corresponding to the plurality of speech samples respectively;
[0240] Feature extraction module 903 is used to extract features from the plurality of speech samples respectively and determine the speech sample feature information corresponding to the plurality of speech samples respectively;
[0241] The training module 904 is used to train a speech feature prediction model using the text samples and speech sample feature information corresponding to the multiple speech samples respectively. The speech feature prediction model is used to predict the speech feature information of the text to be synthesized. The speech feature information is used to adjust the basic speech feature information of the speech synthesis model to obtain target speech feature information. The target speech feature information is used to synthesize the input text to obtain target speech that matches the target speech style.
[0242] According to an embodiment of the present invention, the identification module 902 is specifically used for:
[0243] For any given text sample, the text sample is segmented according to a preset segmentation granularity to obtain multiple segmentation units;
[0244] According to an embodiment of the present invention, the feature extraction module 903 is specifically used for:
[0245] For any given text sample, feature extraction is performed on the multiple segmentation units contained in the text sample to determine the speech sample feature information corresponding to each of the multiple segmentation units.
[0246] According to an embodiment of the present invention, the training module 904 is specifically used for:
[0247] The speech feature prediction model is trained by taking multiple segmentation units of the text sample corresponding to any speech sample as model input and the speech sample feature information corresponding to the speech sample as model label.
[0248] According to an embodiment of the present invention, the feature extraction module 903 is specifically used for:
[0249] For any segmentation unit in any text sample, the segmentation unit is time-aligned in the corresponding speech sample to determine the time label of the segmentation unit;
[0250] Obtain the waveform of the speech sample;
[0251] Based on the time tag of the segmentation unit, the waveform corresponding to the segmentation unit is determined from the waveform diagram of the speech sample;
[0252] The speech sample feature information corresponding to the segmentation unit is determined based on the waveform.
[0253] According to an embodiment of the present invention, the feature extraction module 903 is specifically used for:
[0254] The speech rate characteristic value of the segmentation unit is determined based on the start time and end time of the waveform;
[0255] The fundamental frequency of the waveform is extracted to determine the pitch feature value of the segmentation unit;
[0256] The volume characteristic value of the segmentation unit is determined based on the amplitude of the waveform;
[0257] The speech sample feature information is generated based on the speech rate feature value, the pitch feature value, and the volume feature value.
[0258] According to an embodiment of the present invention, the training module 904 is specifically used for:
[0259] For any given speech sample, the corresponding text sample is input into the speech feature prediction model, and the predicted speech feature value is output.
[0260] Calculate the loss result between the predicted speech feature value and the speech feature value corresponding to the speech sample;
[0261] The model parameters of the speech feature prediction model are adjusted based on the loss result until the loss result meets the preset conditions.
[0262] Figure 9 The model generation device described above can perform Figure 4The implementation principle and technical effects of the model generation method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the model generation apparatus in the above embodiments performs its operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0263] Figure 10 The schematic diagram illustrates a block diagram of a voice broadcasting device provided in one embodiment of the present invention, such as... Figure 10 As shown, the voice broadcasting device 1000 includes a second text acquisition module 1001, a second feature prediction module 1002, a second feature information adjustment module 1003, and a voice playback module 1004.
[0264] The second text acquisition module 1001 is used to acquire the text to be broadcast;
[0265] The second feature prediction module 1002 is used to predict the speech features of the text to be broadcast in order to obtain speech feature information of the target speech style.
[0266] The second feature information adjustment module 1003 is used to input the text to be played and the speech feature information into the speech synthesis model, so as to adjust the basic speech feature information of the speech synthesis model based on the speech feature information to obtain target speech feature information, and use the target speech feature information to synthesize the speech of the content to be played to obtain the broadcast speech of the target speech style.
[0267] The voice playback module 1004 plays the broadcast voice.
[0268] Figure 10 The aforementioned voice broadcasting device can perform Figure 7 The implementation principle and technical effects of the voice broadcasting method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the voice broadcasting device in the above embodiments performs its operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0269] In one possible design, the speech synthesis device, model generation device, and speech broadcasting device provided in the embodiments of the present invention can be implemented as computing devices, such as... Figure 11 As shown, the computing device may include a storage component 1101 and a processing component 1102;
[0270] The storage component 1101 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 1102 to implement the speech synthesis method, model generation method, and speech broadcasting method provided in the embodiments of the present invention.
[0271] Of course, computing devices may also include other components, such as input / output interfaces and communication components. Input / output interfaces provide an interface between processing components and peripheral interface modules, which can be output devices, input devices, etc. Communication components are configured to facilitate wired or wireless communication between the computing device and other devices.
[0272] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.
[0273] When the computing device is a physical device, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device.
[0274] In practical applications, this computing device can be specifically deployed as a node in a message queue system, acting as a producer, consumer, relay server, or naming server in the message queue system.
[0275] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, can implement the speech synthesis method, model generation method, and speech broadcasting method provided in this invention.
[0276] This invention also provides a computer program product, including a computer program that, when executed by a computer, can implement the speech synthesis method, model generation method, and speech broadcasting method provided in this invention.
[0277] The processing component in the corresponding embodiments described above may include one or more processors to execute computer instructions to complete all or part of the steps in the method described above. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the method described above.
[0278] Storage components are configured to store various types of data to support operation within the device. Storage components can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0279] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0280] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0281] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0282] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech synthesis method, characterized in that, include: Obtain the text to be used for speech synthesis; Speech feature prediction is performed on the text to obtain speech feature information of the target speech style; The basic speech feature information of the speech synthesis model is adjusted using the aforementioned speech feature information to obtain the target speech feature information; The text is input into the speech synthesis model to synthesize the text using the target speech feature information, thereby obtaining the target speech in the target speech style; The step of predicting speech features from the text to obtain speech feature information of the target speech style includes: The speech feature prediction model is used to predict the speech features of the text to obtain the speech feature information of the target speech style; wherein, the speech feature prediction model is trained based on multiple speech samples belonging to the target speech style; The step of using a speech feature prediction model to predict the speech features of the text and obtaining speech feature information of the target speech style includes: The text is segmented according to a preset segmentation granularity to obtain multiple segmentation units; The multiple segmentation units are input into the speech feature prediction model, and speech feature information corresponding to each of the multiple segmentation units is output. The process of adjusting the basic speech feature information of the speech synthesis model using the aforementioned speech feature information to obtain the target speech feature information includes: By utilizing the speech feature information corresponding to the multiple segmentation units, the basic speech feature information of the speech synthesis model is adjusted to obtain the target speech feature information.
2. The method according to claim 1, characterized in that, The speech feature information includes at least one speech feature value. The step of adjusting the basic speech feature information of the speech synthesis model using the speech feature information corresponding to the plurality of segmentation units to obtain the target speech feature information includes: At least one speech feature value corresponding to each segmentation unit is arcregularized to determine multiple target speech feature values; The average value of the multiple target speech feature values is calculated to determine the average speech feature value. Based on the average speech feature value and the target speech feature value corresponding to each segmentation unit, a speech feature ratio value corresponding to each segmentation unit is determined; The speech feature ratio value corresponding to each segmentation unit is input into the speech synthesis model. By using the speech feature ratio value corresponding to each segmentation unit, the basic speech feature information of the speech synthesis model is adjusted to obtain the target speech feature information.
3. The method according to claim 1, characterized in that, The speech feature information includes at least one speech feature value. The step of adjusting the basic speech feature information of the speech synthesis model using the speech feature information corresponding to the plurality of segmentation units to obtain the target speech feature information includes: At least one speech feature value corresponding to each of the multiple segmentation units is input into the speech synthesis model, so as to adjust the basic speech feature information of the speech synthesis model using the multiple speech feature values corresponding to the multiple segmentation units, and obtain the target speech feature information.
4. The method according to claim 2, characterized in that, The step of adjusting the basic speech feature information of the speech synthesis model by utilizing the speech feature ratio value corresponding to each segmentation unit to obtain the target speech feature information includes: The speech feature ratio value corresponding to each segmentation unit is multiplied by the basic speech feature information to generate the first speech feature information corresponding to each segmentation unit. Each of the first speech feature information is added to the basic speech feature information to determine the target speech feature information.
5. The method according to claim 2, characterized in that, The step of determining the speech feature ratio value corresponding to each segmentation unit based on the average speech feature value and the target speech feature value corresponding to each segmentation unit includes: Divide the target speech feature value corresponding to each segmentation unit by the average speech feature value to generate the speech feature ratio value corresponding to each segmentation unit.
6. A model generation method, characterized in that, include: Obtain a speech sample set, which includes multiple speech samples belonging to the same target speech style; Perform text recognition on the multiple speech samples respectively to determine the text samples corresponding to the multiple speech samples; Feature extraction is performed on the multiple speech samples respectively to determine the speech sample feature information corresponding to each of the multiple speech samples; A speech feature prediction model is trained using the text samples and speech sample feature information corresponding to the multiple speech samples respectively. The speech feature prediction model is used to predict the speech feature information of the text to be synthesized. The speech feature information is used to adjust the basic speech feature information of the speech synthesis model to obtain target speech feature information. The target speech feature information is used to synthesize the input text to obtain target speech that matches the target speech style. The speech feature prediction model predicts the speech feature information of the text to be synthesized through the following operations: The text sample is segmented according to a preset segmentation granularity to obtain multiple segmentation units; The multiple segmentation units are input into the speech feature prediction model, and speech feature information corresponding to each of the multiple segmentation units is output. The adjustment of the basic speech feature information of the speech synthesis model to obtain the target speech feature information includes: By utilizing the speech feature information corresponding to the multiple segmentation units, the basic speech feature information of the speech synthesis model is adjusted to obtain the target speech feature information.
7. The method according to claim 6, characterized in that, The step of extracting features from multiple speech samples in the speech sample set and determining the speech sample feature information corresponding to each of the multiple speech samples includes: For any given text sample, feature extraction is performed on the multiple segmentation units contained in the text sample to determine the speech sample feature information corresponding to each of the multiple segmentation units. The step of training the speech feature prediction model using the text samples and speech feature information corresponding to the multiple speech samples includes: The speech feature prediction model is trained by taking multiple segmentation units of the text sample corresponding to any speech sample as model input and the speech sample feature information corresponding to the speech sample as model label.
8. The method according to claim 7, characterized in that, For any given text sample, feature extraction is performed on multiple segmentation units contained in the text sample to determine the speech sample feature information corresponding to each of the multiple segmentation units, including: For any segmentation unit in any text sample, the segmentation unit is time-aligned in the corresponding speech sample to determine the time label of the segmentation unit; Obtain the waveform of the speech sample; Based on the time tag of the segmentation unit, the waveform corresponding to the segmentation unit is determined from the waveform diagram of the speech sample; The speech sample feature information corresponding to the segmentation unit is determined based on the waveform.
9. The method according to claim 8, characterized in that, The step of determining the speech sample feature information corresponding to the segmentation unit based on the waveform includes: The speech rate characteristic value of the segmentation unit is determined based on the start time and end time of the waveform; The fundamental frequency of the waveform is extracted to determine the pitch feature value of the segmentation unit; The volume characteristic value of the segmentation unit is determined based on the amplitude of the waveform; The speech sample feature information is generated based on the speech rate feature value, the pitch feature value, and the volume feature value.
10. The method according to claim 6, characterized in that, The step of training a speech feature prediction model using the text samples and speech sample feature information corresponding to the multiple speech samples includes: For any given speech sample, the corresponding text sample is input into the speech feature prediction model, and the predicted speech feature value is output. Calculate the loss result between the predicted speech feature value and the speech feature value corresponding to the speech sample; The model parameters of the speech feature prediction model are adjusted based on the loss result until the loss result meets the preset conditions.
11. A voice broadcasting method, characterized in that, include: Get the text to be broadcast; Speech feature prediction is performed on the text to be broadcast to obtain speech feature information of the target speech style; The text to be broadcast and the speech feature information are input into the speech synthesis model. The basic speech feature information of the speech synthesis model is adjusted based on the speech feature information to obtain the target speech feature information. The target speech feature information is then used to synthesize the text to be broadcast to obtain the broadcast speech with the target speech style. Play the aforementioned audio message; The step of predicting the speech features of the text to be broadcast to obtain speech feature information of the target speech style includes: The speech feature prediction model is used to predict the speech features of the text to be broadcast in order to obtain the speech feature information of the target speech style; wherein, the speech feature prediction model is trained based on multiple speech samples belonging to the target speech style. The step of using a speech feature prediction model to predict the speech features of the text to be broadcast, and obtaining speech feature information of the target speech style, includes: The text is segmented according to a preset segmentation granularity to obtain multiple segmentation units; The multiple segmentation units are input into the speech feature prediction model, and speech feature information corresponding to each of the multiple segmentation units is output. The process of adjusting the basic speech feature information of the speech synthesis model using the aforementioned speech feature information to obtain the target speech feature information includes: By utilizing the speech feature information corresponding to the multiple segmentation units, the basic speech feature information of the speech synthesis model is adjusted to obtain the target speech feature information.
12. A computing device, characterized in that, This includes processing components and storage components; The storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement the speech synthesis method as described in any one of claims 1-5, or the model generation method as described in any one of claims 6-10, or the speech broadcasting method as described in claim 11.
13. A computer storage medium, characterized in that, The device stores a computer program, which, when executed by a computer, implements the speech synthesis method as described in any one of claims 1-5, or the model generation method as described in any one of claims 6-10, or the speech broadcasting method as described in claim 11.
Citation Information
Patent Citations
Model training and speech synthesis method and device, equipment and medium
CN111883101A