Training text-to-speech model, text-to-speech method, device and equipment
By introducing phoneme sequences and granular hierarchical structural annotation information into the text-to-speech model, the problem of unnatural speech prosody in existing technologies is solved, and more natural speech feature prediction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2024-06-19
- Publication Date
- 2026-07-24
AI Technical Summary
Existing text-to-speech models cannot effectively utilize the structural information of text at different granular levels, resulting in unnatural predicted speech prosody, which is especially evident in non-autoregressive algorithm frameworks.
Structural segmentation information, including phoneme sequences and granular-level structural annotations, is introduced into the input data of the text-to-speech model. Prosodic symbols are inserted through a prosodic model, and speech feature prediction is performed by combining prosodic features at the phoneme and granular levels.
It improves the naturalness of speech prosody, making the predicted speech features more coherent in the text structure, and solves the problem of unnatural speech prosody in existing technologies.
Smart Images

Figure CN119007706B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a method, apparatus, and device for training a text-to-speech model and text-to-speech conversion. Background Technology
[0002] Text-to-speech (TTS) technology is widely used in various fields, such as voice assistants, e-books, navigation systems, and automated customer service. A key focus of TTS technology is to ensure that the rhythm of the speech derived from text is as natural as possible (close to the rhythm of human speech). This rhythm includes at least the cadence of the speech sounds.
[0003] The current approach involves inserting prosodic symbols into the text using a prosodic model to indicate which positions in the text sequence require how long a pause should be. Then, a text-to-speech model is used to convert the text with inserted prosodic symbols into speech with a certain rhythm.
[0004] Based on this, a more effective technical solution is provided, which makes the rhythm of speech converted from text more natural. Summary of the Invention
[0005] This specification provides an embodiment of a method for training a text-to-speech model, including:
[0006] Obtain a text sample and its structural partitioning information; wherein the structural partitioning information represents a structural partitioning at at least one granular level.
[0007] The text sample is inserted with prosodic symbols using a prosodic model, and the text sample with inserted prosodic symbols is converted into a phoneme sequence.
[0008] Based on the phoneme sequence and the structural division information, at least one granularity level of structural annotation information is obtained; wherein, for any granularity level, the structural annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the phoneme sequence belongs;
[0009] The phoneme sequence and the structural annotation information at at least one granular level are used as input to the text-to-speech model to train the text-to-speech model.
[0010] This specification provides an embodiment of a text-to-speech method, including:
[0011] Obtain the target text to be processed, and the target structure partitioning information of the target text; wherein, the structure partitioning information represents a structure partitioning at at least one granular level;
[0012] Prosodic symbols are inserted into the target text using a prosodic model, and the target text with inserted prosodic symbols is converted into a target phoneme sequence;
[0013] Based on the target phoneme sequence and the target structure partitioning information, at least one granularity level of target structure annotation information is obtained; wherein, for any granularity level, the target structure annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the target phoneme sequence belongs.
[0014] The target phoneme sequence and the target structure annotation information at at least one granularity level are input into the text-to-speech model, and the predicted speech features are output.
[0015] This specification provides an apparatus for training a text-to-speech model, comprising:
[0016] The acquisition module acquires text samples and structural partitioning information of the text samples; wherein the structural partitioning information represents structural partitioning at at least one granular level.
[0017] The conversion module uses a prosodic model to insert prosodic symbols into the text sample and converts the text sample with inserted prosodic symbols into a phoneme sequence.
[0018] The processing module obtains at least one granularity level of structural annotation information based on the phoneme sequence and the structural partitioning information; wherein, for any granularity level, the structural annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the phoneme sequence belongs.
[0019] The training module takes the phoneme sequence and the structural annotation information at least one granular level as input to the text-to-speech model and trains the text-to-speech model.
[0020] This specification provides an embodiment of a text-to-speech device, comprising:
[0021] The acquisition module acquires the target text to be processed, and the target structure partitioning information of the target text; wherein, the structure partitioning information represents a structure partitioning at at least one granular level;
[0022] The conversion module uses a prosodic model to insert prosodic symbols into the target text and converts the target text with inserted prosodic symbols into a target phoneme sequence.
[0023] The processing module obtains at least one granularity level of target structure annotation information based on the target phoneme sequence and the target structure partitioning information; wherein, for any granularity level, the target structure annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the target phoneme sequence belongs.
[0024] The prediction module inputs the target phoneme sequence and the target structure annotation information at at least one granularity level into the text-to-speech model and outputs predicted speech features.
[0025] This specification also provides a computer program product that stores at least one instruction adapted to be loaded by a processor and executed in accordance with the above-described method steps.
[0026] This specification also provides a storage medium storing a computer program adapted to be loaded by a processor and to execute the steps of the method described above.
[0027] This specification also provides an electronic device, including a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method described above.
[0028] In the technical solution of this specification, the composition of the input data of the text-to-speech model is redefined. The input data includes not only the phoneme sequence corresponding to the text with inserted prosodic symbols, but also structural annotation information that can represent the structural division of the text at at least one granular level. This allows the text-to-speech model to refer not only to the prosody of the text at the phoneme level, but also to the prosody of the text at the granular level of single words, phrases, and sentences during the prediction of speech features. This makes the predicted speech features have the continuity of pronunciation in the text structure, and the prosody is more natural.
[0029] It should be noted that this disclosure pertains to technical solutions in the field of artificial intelligence, and the privacy data used in the implementation of this solution has been authorized by all parties. Attached Figure Description
[0030] Figure 1 This is a flowchart illustrating a method for training a text-to-speech model provided in this publication;
[0031] Figure 2 This is a flowchart illustrating a text-to-speech method disclosed herein;
[0032] Figure 3 This is a schematic diagram of the structure of a device for training a text-to-speech model provided in this disclosure.
[0033] Figure 4 is a schematic structural diagram of a text-to-speech device provided by the present disclosure;
[0034] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of this specification. Specific embodiments
[0035] To make the purpose, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this specification.
[0036] In the field of TTS, the prosody of speech usually includes at least the pronunciation rhythm. The pronunciation rhythm of speech can be understood as the pause duration between each single word in the text. Currently, a common practice is to train a prosody model whose function is to insert prosody symbols into the text to indicate which position in the text sequence requires a pronunciation pause for how long.
[0037] For example, assume the text is "Well, thinking, don't pay service fees, security deposits, etc.", then what we are concerned about is how many pause durations need to be waited after each single word is pronounced before the pronunciation of the next single word can be carried out.
[0038] In practical applications, there are various prosody models, which are generally the same in principle. Here, a kind of prosody symbols defined by a prosody model is introduced. The prosody symbols used by this prosody model include: the first prosody symbol (denoted as #1), the second prosody symbol (denoted as #2), the third prosody symbol (denoted as #3), and the fourth prosody symbol (denoted as #4).
[0039] Among them, the pronunciation pause durations represented by the first prosody symbol, the second prosody symbol, the third prosody symbol, and the fourth prosody symbol increase in sequence.
[0040] #1 represents the shortest pause duration, which is usually applicable to the pause after the pronunciation of a single character, a single word, or a single phrase. #2 and #3 represent longer pause durations (#3 is slightly longer than #2), which are usually applicable to the pause after the pronunciation of partial short sentences in a complete sentence. #4 represents the longest pause duration, which is usually applicable to the pause after the pronunciation of a complete sentence. It is easy to understand that the text may include multiple complete sentences, and objectively, #4 can also separate different sentences.
[0041] Continuing with the above example, to make the rhythm of the speech corresponding to this text natural, after the prosody model inserts prosody symbols, it might be "En #2thinking #1Don't #1Pay #1Service fee #3Security deposit, etc. #4".
[0042] One available technical solution is to convert the text with inserted prosody symbols mentioned above into a phoneme sequence (the above prosody symbols will also be regarded as a phoneme), and then use the phoneme sequence as the sole input of the text-to-speech model to train the text-to-speech model.
[0043] In the above technical solution, when the text-to-speech model analyzes the phoneme sequence, it only analyzes the relationship features between adjacent phonemes at the phoneme level and uses this relationship feature as a reference for predicting speech features. However, text毕竟 has a certain hierarchical structure of granularity division. Here, the granularity level can be a single letter, a single syllable, a single character, a single phrase, a single sentence, etc. And the hierarchical structure of text at a certain granularity level usually determines the prosody when humans orally express the text. For example, when humans orally express text, the pause after pronouncing a single character is relatively short, the pause after pronouncing a single phrase is relatively long, and the pause after pronouncing a single sentence is even longer.
[0044] However, in the above technical solution, the text-to-speech model cannot extract the information about the structural division of the text at a certain granularity level from the input data, that is, it cannot analyze which phonemes belong to the same granularity unit and which phonemes do not belong to the same granularity unit. Therefore, when predicting speech features, it cannot refer to the association between the structural division of the text at a certain granularity level and the prosody. This may lead to the problem that the speech prosody obtained from the predicted speech features is not natural enough, such as the pause after pronouncing a single character is too long and the pause after pronouncing a single phrase is too short.
[0045] Especially, if the algorithm framework adopted by the text-to-speech model is a non-autoregressive algorithm framework (such as FastSpeech2), then the above problem is more prominent.
[0046] Based on this, in the technical solution provided in this disclosure, the composition of the input data of the text-to-speech model is redefined. The input data not only includes the phoneme sequence corresponding to the text with inserted prosody symbols, but also includes the structural annotation information that can represent the structural division of the text at at least one granularity level. Thus, during the process of the text-to-speech model predicting speech features, it can not only refer to the prosody of the text at the phoneme level, but also refer to the prosody of the text at the granularity levels such as single words, phrases, sentences, etc. (explicitly controlling the relationship between phonemes from the perspective of text structure meaning). This can make the speech prosody obtained from the predicted speech features have the coherence of pronunciation in text structure and the prosody is more natural.
[0047] The technical solution of this disclosure is described in detail below with reference to the accompanying drawings.
[0048] Figure 1 This is a flowchart illustrating a method for training a text-to-speech model provided in this disclosure, including the following steps:
[0049] S100: Obtain a text sample and the structural division information of the text sample.
[0050] Easy to understand Figure 1 The method described herein is for a single text sample. However, in actual model training, multiple text samples are typically used, and each text sample is trained using... Figure 1 The process is carried out according to the method flow shown.
[0051] It is also easy to understand that during the model training phase, the training of the model using text samples is done iteratively. After each iteration, the model parameters need to be adjusted based on the difference between the model prediction results (predicted speech features) and the standard results corresponding to the text samples (speech features with standard, assumed prosody). This part will not be elaborated in this disclosure.
[0052] S102: Insert prosodic symbols into the text sample using a prosodic model, and convert the text sample with inserted prosodic symbols into a phoneme sequence.
[0053] S104: Based on the phoneme sequence and the structural division information, obtain at least one level of structural annotation information.
[0054] S106: Use the phoneme sequence and the structural annotation information at least one granular level as input to the text-to-speech model to train the text-to-speech model.
[0055] The structural partitioning information of the text is used to represent the structural partitioning at at least one granularity level. In the implementation scheme, several granularity levels can be defined, and the text can be structurally partitioned at each granularity level. Each partitioned unit is called a granularity unit.
[0056] In some embodiments, multiple granularity levels can be set. It is easy to understand that introducing even a single granularity level of structural annotation information into the model input can optimize the predictive ability of the text-to-speech model. The optimal number of granularity levels introduced will yield the greatest optimization effect.
[0057] As an example, the granularity levels may include a single - character granularity level, a phrase granularity level, and a sentence granularity level. It should be noted here that for languages represented by Chinese, a single - character granularity can be understood as a single Chinese character. For languages represented by English, a single - character granularity can be understood as a single word. Therefore, the single - character granularity level in this disclosure can also be referred to as the single - word granularity level.
[0058] The phrase granularity can be understood as a single phrase, or a combination of multiple phrases, or a combination of a phrase and a single word or character. In Chinese, a phrase can be an idiom, a common saying, an independent short phrase within a complete sentence, etc. The sentence granularity can be understood as a single complete sentence.
[0059] In some embodiments, the phrase granularity level may specifically include a prosodic - phrase granularity level. Thus, in the step of obtaining the structural - division information of the prosodic - phrase granularity level of the text sample, the structural - division information of the prosodic - phrase granularity level of the text sample can be specifically determined according to the prosodic - phrase symbols inserted in the text sample.
[0060] Furthermore, in the case where the prosodic symbols used in the prosodic model include a first prosodic symbol, a second prosodic symbol, a third prosodic symbol, and a fourth prosodic symbol, and the pronunciation pause durations represented by the first prosodic symbol, the second prosodic symbol, the third prosodic symbol, and the fourth prosodic symbol increase in sequence, it can be determined that the second prosodic symbol and the third prosodic symbol belong to prosodic - phrase symbols.
[0061] For example, continuing with the above example, the text with inserted prosodic - phrase symbols is "En #2 thinking #1 don't #1 pay #1 service fee #3 deposit, etc. #4". Then, #3 is identified as a prosodic - phrase symbol. According to the position of #3, it can be judged that "don't pay service fee" is a granularity unit at the prosodic - phrase granularity level.
[0062] In addition, in step S104 above, for any granularity level, the structural - annotation information of this granularity level indicates: at this granularity level, the granularity unit to which each phoneme in the phoneme sequence belongs.
[0063] As an example, for any granularity level, the structural - annotation information of this granularity level is the structural - annotation sequence of this granularity level; each element in the structural - annotation sequence corresponds one - to - one with each element in the phoneme sequence; the different elements in the structural - annotation sequence are B, I, E, S respectively; B represents the starting phoneme of the granularity unit, I represents the internal phoneme of the granularity unit, E represents the ending phoneme of the granularity unit, and S represents an independent phoneme that is alone as a granularity unit.
[0064] For example, the phoneme sequence corresponding to "For example, “en #2 thinking #1 don't #1 pay #1 service fee #3 deposit, etc. #4” is “#4, en4, #2, TH, TH1, NG, K, IN0, NG, #1, b, u2, y, ao4, #1, zh, iii1, f, u4, #1, f, u2, w, u4, f, ei4, #3, b, ao3, zh, eng4, j, in1, d, eng3, #4”.
[0065] The structural annotation sequence of the above phoneme sequence at the single - character granularity level is:
[0066] “S, S, S, B, I, I, I, I, E, S, B, E, B, E, S, B, E, B, E, S, B, E, B, E, B, E, S, B, E, B, E, B, E, B, E, S”.
[0067] The structural annotation sequence of the above phoneme sequence at the phrase granularity level is:
[0068] “S, S, S, B, I, I, I, I, E, S, B, I, I, E, S, B, I, I, E, S, B, I, I, I, I, E, S, B, I, I, I, I, I, I, E, S”.
[0069] The structural annotation sequence of the above phoneme sequence at the sentence granularity level is:
[0070] “S, B, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, I, S”.
[0071] It is easy to understand that the above - mentioned phoneme sequence and structural annotation sequence can be mapped into mapping features represented by digital identifiers after being input into the model, which is convenient for the model to analyze.
[0072] In some embodiments, during the process of training a text - to - speech model, the encoded feature corresponding to the phoneme sequence can be spliced with the mapping feature corresponding to the structural annotation information to obtain a spliced feature; wherein, the encoded feature is the mapping feature corresponding to the phoneme sequence after passing through an encoder. Then, the spliced feature is input into a speech feature predictor, and the predictor can finally output the predicted speech feature.
[0073] Figure 2 It is a schematic flowchart of a text - to - speech method provided by the present disclosure, including the following steps:
[0074] S200: Obtain the target text to be processed and the target structure division information of the target text.
[0075] S202: Insert prosodic symbols into the target text using a prosodic model, and convert the target text with inserted prosodic symbols into a target phoneme sequence.
[0076] S204: Based on the target phoneme sequence and the target structure partitioning information, obtain at least one level of target structure annotation information.
[0077] S206: Input the target phoneme sequence and the target structure annotation information at least one granularity level into the text-to-speech model, and output the predicted speech features.
[0078] As is easy to understand, the model prediction stage involves using the target text that needs to be converted into speech, obtaining the target phoneme sequence and at least one level of target structure annotation information corresponding to the target text, inputting the text-to-speech model, and outputting predicted speech features.
[0079] In the Figure 1 The explanation of the method flow shown has already provided a detailed introduction to the relevant principles in the model analysis and prediction process, which can be understood by referring to the previous text. Figure 2 The principles behind the method flow shown will not be elaborated further.
[0080] Furthermore, this disclosure also provides a schematic diagram of the structure of an apparatus for training a text-to-speech model, such as... Figure 3 As shown, it includes:
[0081] The acquisition module 301 acquires a text sample and structural partitioning information of the text sample; wherein the structural partitioning information represents a structural partitioning at at least one granular level.
[0082] The conversion module 302 uses a prosody model to insert prosodic symbols into the text sample and converts the text sample with inserted prosodic symbols into a phoneme sequence.
[0083] Processing module 303 obtains at least one granularity level of structural annotation information based on the phoneme sequence and the structural division information; wherein, for any granularity level, the structural annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the phoneme sequence belongs;
[0084] The training module 304 uses the phoneme sequence and the structural annotation information at least one granular level as input to the text-to-speech model to train the text-to-speech model.
[0085] This disclosure also provides a schematic diagram of the structure of a text-to-speech device, including:
[0086] The acquisition module 401 acquires the target text to be processed and the target structure partitioning information of the target text; wherein, the structure partitioning information represents a structure partitioning at at least one granular level;
[0087] The conversion module 402 uses a prosody model to insert prosodic symbols into the target text and converts the target text with inserted prosodic symbols into a target phoneme sequence.
[0088] Processing module 403 obtains at least one granularity level of target structure annotation information based on the target phoneme sequence and the target structure partitioning information; wherein, for any granularity level, the target structure annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the target phoneme sequence belongs.
[0089] The prediction module 404 inputs the target phoneme sequence and the target structure annotation information at least one granular level into the text-to-speech model.
[0090] The above-described apparatus embodiments correspond to the method embodiments, and detailed descriptions can be found in the description of the method embodiments section, which will not be repeated here. The apparatus embodiments are derived based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments; detailed descriptions can be found in the corresponding method embodiments.
[0091] This specification also provides a computer storage medium that can store multiple instructions adapted for loading by a processor and executing the methods of this disclosure.
[0092] This specification also provides a computer program product that stores at least one instruction, which is loaded by the processor and executes the method of the embodiments of this disclosure.
[0093] The embodiments in this specification also provide Figure 5 The diagram shows the structure of the electronic device. At the hardware level, this electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for various tasks. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to implement the aforementioned voice activity detection method.
[0094] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0095] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0096] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0097] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0098] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0099] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0100] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0101] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0102] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0103] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0104] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0105] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0106] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0107] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0108] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0109] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0110] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for training a text-to-speech model, comprising: Obtain a text sample and its structural partitioning information; wherein the structural partitioning information represents a structural partitioning at at least one granular level. Prosodic symbols are inserted into the text sample using a prosodic model. The prosodic symbols used in the prosodic model include: a first prosodic symbol, a second prosodic symbol, a third prosodic symbol, and a fourth prosodic symbol. The duration of the pronunciation pauses represented by the first prosodic symbol, the second prosodic symbol, the third prosodic symbol, and the fourth prosodic symbol increases sequentially. The second prosodic symbol and the third prosodic symbol are prosodic phrase symbols. Based on the prosodic phrase symbols already inserted in the text sample, determine the structural division information of the prosodic phrase granularity level of the text sample; The text sample with inserted prosodic symbols is converted into a phoneme sequence; Based on the phoneme sequence and the structural division information, at least one granularity level of structural annotation information is obtained; wherein, for any granularity level, the structural annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the phoneme sequence belongs; In the process of training the text-to-speech model, the encoded features corresponding to the phoneme sequence are concatenated with the mapping features corresponding to the structural annotation information to obtain concatenated features; wherein, the encoded features are obtained by passing the mapping features corresponding to the phoneme sequence through an encoder; the concatenated features are input into a speech feature predictor to train the text-to-speech model.
2. The method as described in claim 1, wherein, Granularity levels include single-word granularity level, phrase granularity level, and sentence granularity level.
3. The method as described in claim 1, wherein, For any given granularity level, the structural annotation information for that granularity level is the structural annotation sequence for that granularity level; Each element in the structure annotation sequence corresponds one-to-one with each element in the phoneme sequence; the different elements in the structure annotation sequence are B, I, E, and S, respectively. B represents the starting phoneme of a granular unit, I represents the internal phoneme of a granular unit, E represents the ending phoneme of a granular unit, and S represents an independent phoneme that functions as a granular unit on its own.
4. The method according to any one of claims 1-3, wherein the algorithmic framework of the text-to-speech model is a non-autoregressive algorithmic framework.
5. The method as described in claim 4, wherein the algorithm framework of the text-to-speech model specifically includes FastSpeech2.
6. A text-to-speech method, comprising: Obtain the target text to be processed, and the target structure partitioning information of the target text; wherein, the structure partitioning information represents a structure partitioning at at least one granular level; Prosodic symbols are inserted into the target text using a prosodic model, and the target text with inserted prosodic symbols is converted into a target phoneme sequence; Based on the target phoneme sequence and the target structure partitioning information, at least one granularity level of target structure annotation information is obtained; wherein, for any granularity level, the target structure annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the target phoneme sequence belongs. The target phoneme sequence and the target structure annotation information at at least one granularity level are input into the text-to-speech model, and the predicted speech features are output; the text-to-speech model is trained based on the method described in any one of claims 1-5.
7. An apparatus for training a text-to-speech model, comprising: The acquisition module acquires text samples and structural partitioning information of the text samples; wherein the structural partitioning information represents structural partitioning at at least one granular level. The conversion module inserts prosodic symbols into the text sample using a prosodic model. The prosodic symbols used in the prosodic model include: a first prosodic symbol, a second prosodic symbol, a third prosodic symbol, and a fourth prosodic symbol. The duration of the pronunciation pauses represented by the first, second, third, and fourth prosodic symbols increases sequentially. The second and third prosodic symbols are prosodic phrase symbols. Based on the inserted prosodic phrase symbols in the text sample, the module determines the structural division information of the prosodic phrase granularity level of the text sample. Finally, the module converts the text sample with inserted prosodic symbols into a phoneme sequence. The processing module obtains at least one granularity level of structural annotation information based on the phoneme sequence and the structural partitioning information; wherein, for any granularity level, the structural annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the phoneme sequence belongs. In the training module, during the training of the text-to-speech model, the encoded features corresponding to the phoneme sequence are concatenated with the mapping features corresponding to the structural annotation information to obtain concatenated features; wherein, the encoded features are obtained by passing the mapping features corresponding to the phoneme sequence through an encoder; the concatenated features are input into a speech feature predictor to train the text-to-speech model.
8. A text-to-speech device, comprising: The acquisition module acquires the target text to be processed, and the target structure partitioning information of the target text; wherein, the structure partitioning information represents a structure partitioning at at least one granular level; The conversion module uses a prosodic model to insert prosodic symbols into the target text and converts the target text with inserted prosodic symbols into a target phoneme sequence. The processing module obtains at least one granularity level of target structure annotation information based on the target phoneme sequence and the target structure partitioning information; wherein, for any granularity level, the target structure annotation information of that granularity level represents: at that granularity level, the granularity unit to which each phoneme in the target phoneme sequence belongs. The prediction module inputs the target phoneme sequence and the target structure annotation information at at least one granularity level into the text-to-speech model and outputs predicted speech features; the text-to-speech model is trained based on the method described in any one of claims 1-5.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method as claimed in any one of claims 1 to 5.
11. A computer program product having at least one instruction stored thereon, characterized in that, When the at least one instruction is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.