Speech synthesis method and device, equipment, storage medium and program product

By using two-layer sub-models to process semantic features and acoustic features in the speech synthesis technology, the problem of low authenticity of speech synthesis in the prior art is solved, and a speech synthesis effect that is more suitable for human voices is achieved.

CN120048245APending Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311603829.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing speech synthesis technologies are difficult to generate speech that is comparable to real-person pronunciations, and the synthesized audio has low authenticity.

Method used

By obtaining semantic features and acoustic features, inputting them into the two-layer sub-model in the speech synthesis model for processing, generating synthetic audio that is the same as the reference speech timbre.

Benefits of technology

The authenticity of speech synthesis is improved, so that the generated synthetic audio is more in line with the actual vocal characteristics of the object corresponding to the selected tone.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048245A_ABST
    Figure CN120048245A_ABST
Patent Text Reader

Abstract

The invention discloses a speech synthesis method and device, equipment, a storage medium and a program product, and belongs to the technical field of speech synthesis. The method comprises the following steps: acquiring semantic features and acoustic features, wherein the semantic features are used for representing features of text information corresponding to an audio to be synthesized; embedding the semantic features into the acoustic features through a first-layer sub-model in the speech synthesis model to obtain intermediate acoustic features; inputting the intermediate acoustic features into a second-layer sub-model in the speech synthesis model to obtain synthesized audio features; and generating a synthetic audio with the same tone as the reference voice based on the synthetic audio feature. According to the method, the sound output by the speech synthesis model better conforms to the actual sound production characteristics of the object corresponding to the selected tone, and the trueness of speech synthesis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of speech synthesis, and in particular, to a speech synthesis method, apparatus, device, storage medium, and program product. Background Art

[0002] With the development of deep learning, speech synthesis technology has achieved leapfrog development. Lifelike and natural speech synthesis technology has been applied to speech interaction systems such as mobile phone voice assistants, smart speakers, and in-vehicle computers of cars. At the same time, users' demands for speech synthesis technology are increasing day by day, and the technical requirements are also increasing. Users not only hope that the synthesized speech can be comparable to that of real people, but also hope to have various timbres, even the timbres of family members and friends.

[0003] In related technologies, during speech synthesis, a text sequence is input into a trained speech synthesis model, and the speech synthesis model will automatically generate a synthesized audio according to the text. Specifically, after inputting a text sequence, the speech synthesis model first maps the text sequence to corresponding audio features, and then converts the audio features into sounds that we can understand through the speech synthesis model, that is, into synthesized audio.

[0004] However, the above method can only rigidly simulate real speech, and the generated synthesized audio does not conform to the characteristics of actual human speech, and the authenticity of the synthesized audio is poor. Summary of the Invention

[0005] The present application provides a speech synthesis method, apparatus, device, storage medium, and program product, and the technical solutions are as follows:

[0006] According to one aspect of the present application, a speech synthesis method is provided, and the method includes:

[0007] Obtain a semantic feature and an acoustic feature, where the semantic feature is used to represent the feature of the text information corresponding to the audio to be synthesized, and the acoustic feature is the feature of the acoustic information corresponding to the reference speech, and the reference speech refers to the speech of the object corresponding to the selected timbre;

[0008] Embed the semantic feature into the acoustic feature through the first-layer sub-model in the speech synthesis model to obtain an intermediate acoustic feature, where the first-layer sub-model is used to embed the semantic feature into the acoustic feature, and the intermediate acoustic feature is used to represent the feature obtained after embedding the semantic feature into the acoustic feature;

[0009] Input the intermediate acoustic feature into the second-layer sub-model in the speech synthesis model to obtain a synthesized audio feature, where the second-layer sub-model is used to synthesize the synthesized audio feature based on the intermediate acoustic feature, and the synthesized audio feature is used to represent the feature corresponding to the audio to be synthesized;

[0010] Generate a synthetic audio with the same timbre as the reference speech based on the synthetic audio features.

[0011] According to one aspect of the present application, a method for training a speech synthesis model is provided, the method comprising:

[0012] Obtain sample semantic features, sample acoustic features, and sample audio, where the sample semantic features are features representing the text information corresponding to the audio to be synthesized, the sample acoustic features are features of the acoustic information corresponding to the sample reference speech, and the sample reference speech refers to the speech of the object corresponding to the selected timbre;

[0013] Embed the sample semantic features into the sample acoustic features through the first-layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features, where the first-layer sub-model is used to embed the sample semantic features into the sample acoustic features, and the sample intermediate acoustic features are used to represent the features obtained after embedding the sample semantic features into the sample acoustic features;

[0014] Input the sample intermediate acoustic features into the second-layer sub-model in the speech synthesis model to obtain synthetic audio features, where the second-layer sub-model is used to synthesize the synthetic audio features based on the sample intermediate acoustic features, and the synthetic audio features are used to represent the features corresponding to the audio to be synthesized;

[0015] Generate a synthetic audio with the same timbre as the sample reference speech based on the synthetic audio features;

[0016] Calculate the training loss of the speech synthesis model based on the sample audio and the synthetic audio;

[0017] Update the model parameters of the speech synthesis model according to the training loss.

[0018] According to one aspect of the present application, a speech synthesis device is provided, the device comprising:

[0019] An acquisition module, configured to acquire semantic features and acoustic features, where the semantic features are features representing the text information corresponding to the audio to be synthesized, the acoustic features are features of the acoustic information corresponding to the reference speech, and the reference speech refers to the speech of the object corresponding to the selected timbre;

[0020] A feature processing module, configured to embed the semantic features into the acoustic features through the first-layer sub-model in the speech synthesis model to obtain intermediate acoustic features, where the first-layer sub-model is used to embed the semantic features into the acoustic features, and the intermediate acoustic features are used to represent the features obtained after embedding the semantic features into the acoustic features;

[0021] The feature processing module is configured to input the intermediate acoustic features into a second-layer sub-model in the speech synthesis model to obtain synthesized audio features. The second-layer sub-model is used to synthesize the synthesized audio features based on the intermediate acoustic features, and the synthesized audio features are used to represent the features corresponding to the audio to be synthesized.

[0022] The generation module is configured to generate a synthesized audio with the same timbre as the reference speech based on the synthesized audio features.

[0023] According to one aspect of the present application, there is provided a training device for a speech synthesis model. The device includes:

[0024] An acquisition module is configured to acquire sample semantic features, sample acoustic features, and sample audio. The sample semantic features are used to represent the features of the semantic text information corresponding to the audio to be synthesized. The sample acoustic features are the features of the acoustic information corresponding to the sample reference speech, and the sample reference speech refers to the speech of the object corresponding to the selected timbre.

[0025] The feature processing module is configured to embed the sample semantic features into the sample acoustic features through a first-layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features. The first-layer sub-model is used to embed the sample semantic features into the sample acoustic features, and the sample intermediate acoustic features are used to represent the features obtained after embedding the sample semantic features into the sample acoustic features.

[0026] The feature processing module is configured to input the sample intermediate acoustic features into a second-layer sub-model in the speech synthesis model to obtain synthesized audio features. The second-layer sub-model is used to synthesize the synthesized audio features based on the sample intermediate acoustic features, and the synthesized audio features are used to represent the features corresponding to the audio to be synthesized.

[0027] The generation module is configured to generate a synthesized audio with the same timbre as the sample reference speech based on the synthesized audio features;

[0028] The calculation module is configured to calculate the training loss of the speech synthesis model based on the sample audio and the synthesized audio;

[0029] The update module is configured to update the model parameters of the speech synthesis model according to the training loss.

[0030] According to another aspect of the present application, there is provided a computer device. The computer device includes: a processor and a memory. At least one computer program is stored in the memory, and at least one computer program is loaded and executed by the processor to implement the speech synthesis method described in the above aspect, or the training method of the speech synthesis model described in the above aspect.

[0031] According to another aspect of the present application, there is provided a computer storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the speech synthesis method described in the above aspect, or the training method of the speech synthesis model described in the above aspect.

[0032] According to another aspect of the present application, there is provided a computer program product, the computer program product includes a computer program, the computer program is stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device, so that the computer device executes the speech synthesis method described in the above aspect, or the training method of the speech synthesis model described in the above aspect.

[0033] The beneficial effects brought by the technical solution provided by the present application at least include:

[0034] By obtaining the semantic feature and the acoustic feature corresponding to the reference speech; inputting the semantic feature and the acoustic feature into the first-layer sub-model in the speech synthesis model to perform feature embedding, obtaining an intermediate acoustic feature; inputting the intermediate acoustic feature into the second-layer sub-model in the speech synthesis model to perform speech synthesis, obtaining a synthesized audio feature; decoding the synthesized audio feature to obtain a synthesized audio with the same timbre as the reference speech. The present application processes the semantic feature and the acoustic feature through two layers of sub-models in the speech synthesis model, so that the finally generated synthesized audio feature can learn both the semantic feature and the acoustic feature sufficiently. Compared with the method of rigidly imitating real speech, the voice output by the speech synthesis model in this method is more in line with the actual vocal characteristics of the object corresponding to the selected timbre, improving the authenticity of speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0036] Figure 1 is a schematic diagram of a speech synthesis method provided by an exemplary embodiment of the present application;

[0037] Figure 2 is a schematic diagram of the architecture of a computer system provided by an exemplary embodiment of the present application;

[0038] Figure 3 is a flowchart of a speech synthesis method provided by an exemplary embodiment of the present application;

[0039] Figure 4 It is a flowchart of a speech synthesis method provided by an exemplary embodiment of the present application;

[0040] Figure 5 It is a schematic diagram of obtaining intermediate acoustic features provided by an exemplary embodiment of the present application;

[0041] Figure 6 It is a schematic diagram of obtaining synthesized audio features provided by an exemplary embodiment of the present application;

[0042] Figure 7 It is a framework diagram of generating a training system of a speech synthesis model and training the speech synthesis model provided by an exemplary embodiment of the present application;

[0043] Figure 8 It is a flowchart of a method for training a speech synthesis model provided by an exemplary embodiment of the present application;

[0044] Figure 9 It is a flowchart of a method for training a speech synthesis model provided by an exemplary embodiment of the present application;

[0045] Figure 10 It is a block diagram of a speech synthesis device provided by an exemplary embodiment of the present application;

[0046] Figure 11 It is a block diagram of a device for training a speech synthesis model provided by an exemplary embodiment of the present application;

[0047] Figure 12 It is a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0048] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings. Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0049] The terms used in the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. The singular forms "a", "the", and "that" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0050] It should be understood that although terms such as first, second, and third may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other.

[0051] For the sake of easy understanding, several terms related to this application are explained below.

[0052] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0053] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0054] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.

[0055] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of technical network systems require a large amount of computing and storage resources, such as video websites, picture-based websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.

[0056] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called a "cloud". From the user's perspective, the resources in the "cloud" are infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for as needed.

[0057] As a cloud computing basic capability provider, a cloud computing resource pool will be established, referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform, in which various types of virtual resources are deployed for external customers to choose to use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.

[0058] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on the PaaS layer. SaaS can also be deployed directly on IaaS. PaaS is a platform for software operation, such as databases, Web (World Wide Web) containers, etc. SaaS is a variety of business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.

[0059] Computer vision (CV) is a science that studies how to make machines "see". To put it more specifically, it refers to machine vision such as using cameras and computers to replace human eyes to identify and measure targets, and further perform image processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and mapping, and common biometric recognition technology.

[0060] Machine learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. Pre-trained models are the latest development results of deep learning, integrating the above technologies.

[0061] An embodiment of the present application provides a schematic diagram of a speech synthesis method, as Figure 1 shown. This method can be executed by a computer device, which can be a terminal or a server, and a speech synthesis model is set in the computer device.

[0062] Exemplarily, the computer device acquires semantic features and acoustic features; the computer device inputs the semantic features and acoustic features into the first-layer sub-model in the speech synthesis model to obtain intermediate acoustic features, and the first-layer sub-model is used to embed the semantic features into the acoustic features; the computer device inputs the intermediate acoustic features into the second-layer sub-model in the speech synthesis model to obtain synthesized audio features, and the second-layer sub-model is used to synthesize the synthesized audio features based on the intermediate acoustic features; the computer device obtains a synthesized audio with the same timbre as the reference speech based on the synthesized audio features.

[0063] Semantic features are features used to represent the semantic information corresponding to the audio to be synthesized; or, semantic features are features used to represent the text information corresponding to the audio to be synthesized.

[0064] Optionally, the acquisition method of semantic features includes at least one of the following methods, but is not limited thereto:

[0065] · Performing speech recognition on the speech signal corresponding to the audio to be synthesized to obtain the text content; extracting features from the text content to obtain semantic features;

[0066] · Performing character recognition and feature extraction on the characters in the text content to obtain semantic features; optionally, the text content includes at least one of phrases, sentences, paragraphs, and texts composed of characters and / or symbols, etc. Optionally, the language types of the text content are not limited to Chinese and English and can be any one or more languages;

[0067] · Directly extracting features from the speech signal to obtain semantic features used to represent the semantics of the speech signal.

[0068] An acoustic feature is a feature of acoustic information corresponding to a reference speech.

[0069] Optionally, the acoustic feature includes at least one of a recording environment feature, a timbre feature, and a prosody duration feature, but is not limited thereto. The embodiments of the present application do not make specific limitations thereon.

[0070] Among them, the recording environment feature is used to characterize the feature of the reference speech recording environment; the timbre feature is used to characterize the timbre feature of the object corresponding to the selected timbre; the prosody duration feature is used to characterize the prosody feature in the reference speech. For example, the prosody duration feature is used to characterize features such as the pitch, duration, and pitch height of the object corresponding to the selected timbre when speaking, or the prosody duration feature is used to characterize the cadence feature of the object corresponding to the selected timbre when speaking.

[0071] The reference speech refers to the speech of the object corresponding to the selected timbre.

[0072] Optionally, the reference speech is the speech of the object corresponding to the selected timbre; or, the reference speech is the speech when the object corresponding to the selected timbre reads a text; or, the reference speech is an audio segment containing the sound of the object corresponding to the selected timbre, but is not limited thereto. For example, the reference speech is a recording of person A reading a text.

[0073] The intermediate acoustic feature is used to represent the feature obtained by embedding the semantic feature into the acoustic feature.

[0074] The synthesized audio feature is used to represent the feature corresponding to the audio to be synthesized.

[0075] As Figure 1 shown, the computer device obtains the reference speech 10, inputs the reference speech 10 into an acoustic feature extraction network to extract features, and obtains the initial acoustic feature 20; the computer device quantifies the initial acoustic feature 20 to obtain the acoustic feature 30. The initial acoustic feature 20 before quantization is a one-dimensional matrix, and the acoustic feature 30 after quantization is a two-dimensional matrix.

[0076] The initial acoustic feature 20 refers to the continuous feature extracted for the reference speech 10.

[0077] For example, the initial acoustic feature 20 before quantization can be expressed as [a1, a2, a3], and the acoustic feature 30 after quantization is a three-dimensional matrix, which can be expressed as:

[0078] It should be noted that the dimension of a matrix refers to the number of rows of the matrix.

[0079] After obtaining the quantized acoustic features 30, the computer device adds up the feature values in the same column dimension of the acoustic features 30 to obtain one-dimensional acoustic features 50; the computer device concatenates the semantic features 40 and the one-dimensional acoustic features 50 to obtain a concatenated feature, where the concatenated feature refers to the feature obtained by concatenating the semantic features 40 and the one-dimensional acoustic features 50; the computer device inputs the concatenated feature into the first-layer sub-model 60 to obtain intermediate acoustic features 70.

[0080] For example, the computer device adds up the feature values in the first column dimension of the acoustic features 30. For instance, it adds the numerical values of a11, a12, and a13 to obtain the first acoustic feature value A1 in the one-dimensional acoustic features 50; similarly, it adds the numerical values of a21, a22, and a23 to obtain the second acoustic feature value A2 in the one-dimensional acoustic features 50; it adds the numerical values of a31, a32, and a33 to obtain the third acoustic feature value A3 in the one-dimensional acoustic features 50. Therefore, the obtained one-dimensional acoustic features 50 can be represented as [A1, A2, A3].

[0081] In some embodiments, the semantic features 40 can be represented as [s1, s2, s3]. The computer device concatenates the semantic features 40 and the one-dimensional acoustic features 50 to obtain a concatenated feature, and the obtained concatenated feature can be represented as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the concatenated feature into the first-layer sub-model 60 to obtain the first intermediate acoustic feature value h1 in the intermediate acoustic features 70; the computer device inputs the first synthesized audio feature value corresponding to the first intermediate acoustic feature value h1 and the concatenated feature into the first-layer sub-model 60 to predict the intermediate acoustic feature value, and obtains the second intermediate acoustic feature value h2, where the first synthesized audio feature value is predicted by inputting the first intermediate acoustic feature value into the second-layer sub-model 80; and so on, until the number of output intermediate acoustic feature values is equal to the number of feature values in the one-dimensional acoustic features; the computer device combines the generated intermediate acoustic feature values to obtain the intermediate acoustic features 70, which can be represented as [h1, h2, h3, Eos], where Eos is used to represent the end symbol.

[0082] After obtaining the intermediate acoustic features 70, the computer device sequentially inputs the intermediate acoustic feature values in the intermediate acoustic features 70 into the second-layer sub-model 80 to predict the synthesized audio feature values in the synthesized audio features 90, and obtains the synthesized audio features 90; the computer device decodes the synthesized audio features 90 to obtain a synthesized audio 100 with the same timbre as the reference speech 10.

[0083] In some embodiments, the computer device inputs the first intermediate acoustic feature value h1 in the intermediate acoustic features 70 into the second-layer sub-model 80 to predict the synthesized audio feature value in the synthesized audio features 90, and obtains the first synthesized audio feature value in the synthesized audio features 90; the computer device inputs the generated first synthesized audio feature value and the splicing feature into the first-layer sub-model 60 to obtain the second intermediate acoustic feature value h2; the computer device inputs the generated second intermediate acoustic feature value h2 into the second-layer sub-model 80 to perform speech synthesis, and obtains the second synthesized audio feature value in the synthesized audio features 90; and so on, until no intermediate acoustic feature value is input into the second-layer sub-model 80, and the computer device merges the generated synthesized audio feature values to obtain the synthesized audio features 90.

[0084] In summary, the method provided in this embodiment obtains the semantic features and the acoustic features corresponding to the reference speech; inputs the semantic features and the acoustic features into the first-layer sub-model in the speech synthesis model to perform feature embedding, and obtains intermediate acoustic features; inputs the intermediate acoustic features into the second-layer sub-model in the speech synthesis model to perform speech synthesis, and obtains synthesized audio features; and decodes the synthesized audio features to obtain a synthesized audio with the same timbre as the reference speech. This application processes the semantic features and the acoustic features through two layers of sub-models in the speech synthesis model, so that the finally generated synthesized audio features can learn both the semantic features and the acoustic features. Compared with the method of rigidly imitating real speech, the voice output by the speech synthesis model in this method is more in line with the actual vocal characteristics of the object corresponding to the selected timbre, and improves the authenticity of speech synthesis.

[0085] Figure 2 The schematic diagram of the architecture of a computer system provided by an embodiment of the present application is shown. The computer system may include: a terminal 100 and a server 200.

[0086] The terminal 100 may be an electronic device such as a mobile phone, a tablet computer, a vehicle-mounted terminal (carputer), a wearable device, a personal computer (PC), a vehicle-mounted terminal, an aircraft, a self-service vending terminal, etc. A client for running the target application program may be installed in the terminal 100. The target application program may be an application program for reference speech synthesis, or other application programs provided with speech synthesis functions. The present application does not make any limitation thereto. In addition, the present application does not make any limitation to the form of the target application program, including but not limited to an application program (App), a small program, etc. installed in the terminal 100, and may also be in the form of a web page.

[0087] The server 200 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, Content Delivery Network (CDN), and cloud server providing basic cloud computing services such as big data. The server 200 can be the background server of the above target application, used to provide background services for the client of the target application.

[0088] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system. This can only be achieved through cloud computing.

[0089] In some embodiments, the above server can also be implemented as a node in a blockchain system. Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity (anti-counterfeiting) of the information and generate the next block. Blockchain can include the blockchain underlying platform, platform product service layer, and application service layer.

[0090] Communication can be carried out between the terminal 100 and the server 200 through a network, such as a wired or wireless network.

[0091] In the method for training a voice synthesis method and a voice synthesis model provided in the embodiments of the present application, the execution subject of each step can be a computer device, and the computer device refers to an electronic device with data computing, processing, and storage capabilities. Figure 2Taking the implementation of the shown solution as an example, the speech synthesis method and the training method of the speech synthesis model can be executed by the terminal 100 (for example, the client of the target application installed and running in the terminal 100 executes the speech synthesis method and the training method of the speech synthesis model), or can be executed by the server 200, or can be executed by the interaction and cooperation between the terminal 100 and the server 200. This application does not make any limitation in this regard.

[0092] Figure 3 FIG. 4 is a flowchart of a speech synthesis method provided by an exemplary embodiment of this application. This method can be executed by a computer device, and the computer device can be a terminal or a server, and a speech synthesis model is set in the computer device. This method includes:

[0093] Step 302: Obtain semantic features and acoustic features.

[0094] The semantic features are features used to represent the semantic information corresponding to the audio to be synthesized; or, the semantic features are features used to represent the text information corresponding to the audio to be synthesized.

[0095] Optionally, the obtaining manner of the semantic features includes: performing speech recognition on the speech signal corresponding to the audio to be synthesized to obtain the text content; extracting features from the text content to obtain the semantic features; or, performing character recognition and feature extraction on the characters in the text content to obtain the semantic features; or, directly extracting features from the speech signal to obtain the semantic features used to represent the semantics of the speech signal.

[0096] Optionally, the text content includes at least one of phrases, sentences, paragraphs, and texts composed of characters and / or symbols, etc. Optionally, the language types of the text content are not limited to Chinese and English, and can be any one or more languages.

[0097] The acoustic features are features of the acoustic information corresponding to the reference speech.

[0098] Optionally, the acoustic features include at least one of recording environment features, timbre features, and prosody duration features, but are not limited thereto, and this application embodiment does not make specific limitations in this regard.

[0099] Among them, the recording environment features are used to characterize the features of the reference speech recording environment; the timbre features are used to characterize the timbre features of the object corresponding to the selected timbre; the prosody duration features are used to characterize the prosody features in the reference speech. For example, the prosody duration features are used to characterize features such as pitch, duration, and pitch height of the object corresponding to the selected timbre when speaking, or the prosody duration features are used to characterize the cadence features of the object corresponding to the selected timbre when speaking.

[0100] The reference speech refers to the speech of the object corresponding to the selected timbre.

[0101] Optionally, the reference speech is the speech of the object corresponding to the selected timbre; or, the reference speech is the speech when the object corresponding to the selected timbre reads the text; or, the reference speech is an audio segment containing the voice of the object corresponding to the selected timbre, but is not limited thereto. For example, the reference speech is the recording of person A reading a text.

[0102] Optionally, the reference speech can be a speech with a length of 3 seconds.

[0103] In some embodiments, the acoustic features can be obtained by feature extraction through an acoustic encoder (also referred to as an acoustic feature extraction network) in the speech synthesis model.

[0104] Optionally, the acoustic encoder can be implemented by a Convolutional Neural Network (CNN), a Transformer neural network, or a Conformer network. For example, the acoustic encoder can be implemented using a four-layer Transformer neural network. Of course, the acquisition of acoustic features is not limited to this implementation.

[0105] In some embodiments, the semantic features can be obtained by feature extraction through a semantic encoder in the speech synthesis model.

[0106] Optionally, the semantic encoder can be implemented by a Self-Supervised Learning (SSL) encoder and a k-means clustering algorithm. Of course, the acquisition of semantic features is not limited to this implementation.

[0107] In some embodiments, the computer device can splice the reference speech and the speech signal corresponding to the audio to be synthesized to obtain a spliced speech signal, and the computer device extracts the semantic features and acoustic features by performing feature extraction on the spliced speech signal.

[0108] Optionally, obtain the speech signal corresponding to the reference speech with a first time length and the audio to be synthesized with a second time length; the computer device splices the speech signal corresponding to the reference speech and the audio to be synthesized to obtain a spliced speech signal; the computer device extracts the acoustic features from the reference speech with the first time length through an acoustic feature extraction network; the computer device extracts the semantic features from the speech signal corresponding to the audio to be synthesized with the second time length through a semantic feature extraction network.

[0109] For example, obtain the speech signals corresponding to a 3 - second reference speech and a 7 - second audio - to - be - synthesized. The computer device splices the 3 - second reference speech and the 7 - second speech signal to obtain a 10 - second spliced speech signal. The computer device extracts acoustic features from the 3 - second reference speech through an acoustic feature extraction network. The computer device extracts semantic features from the 7 - second speech signal through a semantic feature extraction network. The computer device finally obtains the synthesized audio based on the semantic features and acoustic features, that is, a synthesized audio that can be expressed in the voice of the object corresponding to the selected timbre through a short 3 - second reference speech. For instance, if person A wants to imitate person B to read an article, only need to obtain the 3 - second speech of person B as the reference speech and the speech signal of person A reading an article. After splicing the 3 - second speech of person B and the speech signal of person A reading an article, the computer device inputs them into the speech synthesis model provided in the embodiment of the present application, and then person A's voice reading an article in the voice of person B can be obtained.

[0110] Step 304: Embed the semantic features into the acoustic features through the first - layer sub - model in the speech synthesis model to obtain intermediate acoustic features.

[0111] The intermediate acoustic features are used to represent the features obtained after embedding the semantic features into the acoustic features.

[0112] Embedding the semantic features into the acoustic features means that through the first - layer sub - model, the acoustic features can learn the semantic information or text information in the semantic features, that is, can obtain the text information that the audio - to - be - synthesized wants to express.

[0113] The speech synthesis model includes a first - layer sub - model and a second - layer sub - model. Among them, the model structures of the first - layer sub - model and the second - layer sub - model are the same, but at least one of their network parameters, execution tasks, and number of network layers is different.

[0114] The first - layer sub - model is used to embed the semantic features into the acoustic features.

[0115] Optionally, the first - layer sub - model can adopt at least one of the attention network Transformer, the pre - trained language representation model (Bidirectional Encoder Representation from Transformers, BERT), and the recurrent neural network (Recurrent Neural Networks, RNN), but is not limited thereto. The embodiments of the present application do not make specific limitations on this.

[0116] Step 306: Input the intermediate acoustic features into the second - layer sub - model in the speech synthesis model to obtain synthesized audio features.

[0117] The synthetic audio feature is used to represent the feature corresponding to the audio to be synthesized.

[0118] The second-layer sub-model is used to synthesize the synthetic audio feature based on the intermediate acoustic feature.

[0119] Optionally, the second-layer sub-model may adopt at least one of the attention network Transformer, BERT, and RNN, but is not limited thereto, and the embodiments of the present application do not make specific limitations thereon.

[0120] Step 308: Generate a synthetic audio with the same timbre as the reference speech based on the synthetic audio feature.

[0121] The synthetic audio is an audio that imitates the voice of the object corresponding to the selected timbre (the human voice in the reference speech) to speak out the text information to be expressed.

[0122] Exemplarily, the computer device decodes the synthetic audio feature, so as to obtain a synthetic audio with the same timbre as the reference speech, that is, to obtain a synthetic audio expressed in the human voice of the reference speech.

[0123] In summary, the method provided in this embodiment obtains the semantic feature and the acoustic feature corresponding to the reference speech; inputs the semantic feature and the acoustic feature into the first-layer sub-model in the speech synthesis model to perform feature embedding, and obtains the intermediate acoustic feature; inputs the intermediate acoustic feature into the second-layer sub-model in the speech synthesis model to perform speech synthesis, and obtains the synthetic audio feature; decodes the synthetic audio feature to obtain a synthetic audio with the same timbre as the reference speech. The present application processes the semantic feature and the acoustic feature through two layers of sub-models in the speech synthesis model, so that the finally generated synthetic audio feature can learn both the semantic feature and the acoustic feature sufficiently. Compared with the method of rigidly imitating real speech, the voice output by the speech synthesis model in this method is more in line with the actual vocal characteristics of the object corresponding to the selected timbre, and improves the authenticity of speech synthesis.

[0124] Figure 4 It is a flowchart of a speech synthesis method provided by an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be a terminal or a server, and a speech synthesis model is set in the computer device. This method includes:

[0125] Step 402: Obtain the semantic feature and the acoustic feature.

[0126] The semantic feature is a feature used to represent the semantic information corresponding to the audio to be synthesized; or, the semantic feature is a feature used to represent the text information corresponding to the audio to be synthesized.

[0127] Optionally, the semantic features can be obtained in the following ways: performing speech recognition on the speech signal to obtain the text content; extracting features from the text content to obtain semantic features; or, performing character recognition and feature extraction on the characters in the text content to obtain semantic features; or, directly extracting features from the speech signal to obtain semantic features representing the semantics of the speech signal.

[0128] The acoustic features are the features of the acoustic information corresponding to the reference speech. The acoustic features can characterize the vocal characteristics, reading characteristics of the object corresponding to the selected timbre, or the characteristics of the reference speech recording environment, etc.

[0129] Optionally, the acoustic features include at least one of recording environment features, timbre features, and prosody duration features, but are not limited thereto. The embodiments of the present application do not make specific limitations in this regard.

[0130] Among them, the recording environment features are used to characterize the features of the reference speech recording environment; the timbre features are used to characterize the timbre features of the object corresponding to the selected timbre; the prosody duration features are used to characterize the prosody features in the reference speech. For example, the prosody duration features are used to characterize the pitch, length, pitch height, etc. of the object corresponding to the selected timbre when speaking, or the prosody duration features are used to characterize the cadence features of the object corresponding to the selected timbre when speaking.

[0131] The reference speech refers to the speech of the object corresponding to the selected timbre.

[0132] Optionally, the reference speech is the speech of the object corresponding to the selected timbre; or, the reference speech is the speech when the object corresponding to the selected timbre reads the text; or, the reference speech is an audio containing the voice of the object corresponding to the selected timbre, but is not limited thereto. For example, the reference speech is the recording of person A reading a text.

[0133] In some embodiments, the computer device obtains the reference speech; the computer device inputs the reference speech into the acoustic feature extraction network in the speech synthesis model to extract features and obtains the acoustic features.

[0134] Optionally, the computer device inputs the reference speech into the acoustic feature extraction network to extract features and obtains the initial acoustic features; the computer device quantifies the initial acoustic features to obtain the acoustic features.

[0135] Optionally, the computer device performs Fourier transform on the speech signal corresponding to each time granularity in the reference speech to obtain the Mel spectrogram corresponding to the reference speech; the computer device inputs the Mel spectrogram into the acoustic feature extraction network to extract features and obtains the initial acoustic features; the computer device quantifies the initial acoustic features to obtain the acoustic features.

[0136] The initial acoustic features refer to the continuous features extracted from the reference speech.

[0137] Since the initial acoustic features are continuous features, for the convenience of subsequent calculations, it is necessary to quantize the initial acoustic features to obtain the acoustic features. Among them, the initial acoustic features before quantization are a one-dimensional matrix, and the acoustic features after quantization are an n-dimensional matrix, where n is a positive integer greater than 1. The acoustic features after quantization have higher expressiveness than the initial acoustic features before quantization, that is, the multi-dimensional matrix can represent the acoustic features more fully.

[0138] For example, the initial acoustic features before quantization can be expressed as [a1, a2, a3], and the acoustic features after quantization are a three-dimensional matrix, which can be expressed as:

[0139] The steps of quantizing the initial acoustic features into acoustic features include: the computer device quantizes the m-th acoustic feature value in the initial acoustic features to obtain the first quantized acoustic feature value in the m-th column dimension of the acoustic features, where m is a positive integer; the computer device quantizes the residual between the first quantized acoustic feature value and the m-th acoustic feature value to obtain the second quantized acoustic feature value in the m-th column dimension; the computer device quantizes the residual between the k-th quantized acoustic feature value and the (k - 1)-th quantized acoustic feature value to obtain the (k + 1)-th quantized acoustic feature value in the m-th column dimension, where k is a positive integer greater than 1; the computer device combines the first k + 1 quantized acoustic feature values in the m-th column dimension to obtain the quantized acoustic feature value in the m-th column dimension of the acoustic features; repeat the above steps, combine the quantized acoustic feature values of each column dimension, and obtain the acoustic features.

[0140] For example, the computer device quantizes the first acoustic feature value a1 in the initial acoustic features to obtain the first quantized acoustic feature value a11 in the first column dimension of the acoustic features; the computer device quantizes the residual between the first acoustic feature value a1 and the first quantized acoustic feature value a11 in the first column dimension to obtain the second quantized acoustic feature value a12 in the first column dimension; the computer device quantizes the residual between the second acoustic feature value a12 and the first quantized acoustic feature value a11 in the first column dimension to obtain the third quantized acoustic feature value a13 in the first column dimension; combine the quantized acoustic feature values in the first column dimension to obtain the acoustic feature in the first column dimension of the acoustic features; perform the above steps on other column dimensions to obtain the acoustic features in other column dimensions of the acoustic features; combine the quantized acoustic feature values of each column dimension to obtain the acoustic features.

[0141] Optionally, the dimension of the acoustic features can be set artificially according to needs.

[0142] In some embodiments, the ways to obtain the reference speech include at least one of the following situations:

[0143] 1. The computer device receives a reference voice. For example, the terminal is the terminal that initiates audio recording. The terminal records audio and, after the recording ends, uses the audio as the reference voice.

[0144] 2. The computer device obtains the reference voice from the stored database.

[0145] It should be noted that the above methods for obtaining the reference voice are only illustrative examples, and the embodiments of the present application do not limit this.

[0146] Step 404: Concatenate the semantic feature and the one-dimensional acoustic feature to obtain a concatenated feature; input the concatenated feature into the first-layer sub-model to obtain an intermediate acoustic feature.

[0147] The concatenated feature refers to the feature obtained by concatenating the semantic feature and the one-dimensional acoustic feature.

[0148] The acoustic feature is the feature of the acoustic information corresponding to the reference voice. The acoustic feature is an n-dimensional matrix, where n is a positive integer greater than 1.

[0149] Exemplarily, the computer device adds the feature values of the same column dimension in the acoustic feature to obtain a one-dimensional acoustic feature.

[0150] For example, the acoustic feature is a three-dimensional matrix, which can be expressed as: The computer device adds the feature values of the first column dimension in the acoustic feature. For example, it adds the values of a11, a12, and a13 numerically to obtain the first acoustic feature value A1 in the one-dimensional acoustic feature; similarly, it adds the values of a21, a22, and a23 numerically to obtain the second acoustic feature value A2 in the one-dimensional acoustic feature; it adds the values of a31, a32, and a33 numerically to obtain the third acoustic feature value A3 in the one-dimensional acoustic feature. Thus, the obtained one-dimensional acoustic feature can be expressed as [A1, A2, A3].

[0151] In some embodiments, the semantic features may be represented as [s1, s2, s3]. The computer device splices the semantic features and the one-dimensional acoustic features to obtain the spliced features, and the obtained spliced features may be represented as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the spliced features into the first-layer sub-model to predict the intermediate acoustic feature values in the intermediate acoustic features, and obtains the first intermediate acoustic feature value in the intermediate acoustic features; the computer device inputs the (i-1)th synthetic audio feature value corresponding to the (i-1)th intermediate acoustic feature value and the spliced features into the first-layer sub-model to predict the intermediate acoustic feature values, and obtains the ith intermediate acoustic feature value, and the (i-1)th synthetic audio feature value is obtained by inputting the (i-1)th intermediate acoustic feature value into the second-layer sub-model for prediction; the computer device loops the previous step until the number of output intermediate acoustic feature values is equal to the number of feature values in the one-dimensional acoustic features; the computer device combines the ith intermediate acoustic feature value with the previous (i-1) intermediate acoustic feature values to obtain the intermediate acoustic features, where i is a positive integer greater than 1.

[0152] Exemplarily, such as Figure 5Schematic diagram for obtaining intermediate acoustic features shown. The computer device adds the feature values of the same column dimension in the acoustic feature 501 to obtain a one-dimensional acoustic feature 503. The semantic feature 502 can be expressed as [s1, s2, s3]. The computer device splices the semantic feature 502 and the one-dimensional acoustic feature 503 to obtain a spliced feature 507. The obtained spliced feature 507 can be expressed as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the spliced feature 507 into the first-layer sub-model 504 to obtain the first intermediate acoustic feature value h1 in the intermediate acoustic feature 505. The computer device inputs the first synthetic audio feature value corresponding to the generated first intermediate acoustic feature value h1 and the spliced feature 507 into the first-layer sub-model 504 to predict the intermediate acoustic feature value, and obtains the second intermediate acoustic feature value h2, where the first synthetic audio feature value is obtained by inputting the first intermediate acoustic feature value h1 into the second-layer sub-model 506 for prediction. The computer device inputs the second synthetic audio feature value corresponding to the generated second intermediate acoustic feature value h2 and the spliced feature 507 into the first-layer sub-model 504 to predict the intermediate acoustic feature value, and obtains the third intermediate acoustic feature value h3, where the second synthetic audio feature value is obtained by inputting the second intermediate acoustic feature value h2 into the second-layer sub-model 506 for prediction. And so on, until the number of intermediate acoustic feature values output by the first-layer sub-model 504 is equal to the number of feature values in the one-dimensional acoustic feature 503. For example, the one-dimensional acoustic feature 503 can be expressed as [Bos, A1, A2, A3], which includes 3 feature values. When the first-layer sub-model 504 outputs 3 intermediate acoustic feature values, the first-layer sub-model 504 ends the prediction. The computer device merges the generated intermediate acoustic feature values to obtain the intermediate acoustic feature 505, which can be expressed as [h1, h2, h3, Eos].

[0153] The speech synthesis model includes a first-layer sub-model and a second-layer sub-model. Among them, the model structures of the first-layer sub-model and the second-layer sub-model are the same, but at least one of their network parameters, execution tasks, and number of network layers is different.

[0154] The first-layer sub-model is used to embed the semantic feature into the acoustic feature.

[0155] Optionally, the first-layer sub-model can adopt at least one of an attention network Transformer, a pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT), and a recurrent neural network (Recurrent Neural Networks, RNN), but is not limited thereto. The embodiments of the present application do not make specific limitations in this regard.

[0156] The semantic feature values in the semantic features are discrete numbers. To facilitate the calculation of the first-layer sub-model, it is necessary to perform mathematical processing on the semantic feature values to convert them into vectors. The vectorization processing formula for semantic features can be expressed as:

[0157] E(s t1 ) = E s (s t1 ) + PE g (t)

[0158] Among them, E(s t1 ) refers to the vectorized semantic features, s t1 refers to the semantic feature values, E s is the embedding function for semantic feature values, and PE g is used to represent the position embedding function of the first-layer sub-model. t1 is used to represent the time point or time position, 1 ≤ t1 ≤ T 1 , and T 1 is used to represent the time length corresponding to the semantic features.

[0159] Similarly, the acoustic feature values in the one-dimensional acoustic features are discrete numbers. To facilitate the calculation of the first-layer sub-model, it is necessary to perform mathematical processing on the one-dimensional acoustic features to convert them into vectors. The vectorization processing formula for one-dimensional acoustic features can be expressed as:

[0160]

[0161] Among them, E(a t2 ) refers to the vectorized one-dimensional acoustic features, refers to the acoustic feature values, refers to the one-dimensional acoustic feature values, E s is the embedding function for acoustic feature values, and PE g is used to represent the position embedding function of the first-layer sub-model. t2 is used to represent the time point or time position, 1 ≤ t2 ≤ T 2 , and T2 is used to represent the time length corresponding to the one-dimensional acoustic features.

[0162] The calculation formula for the first-layer sub-model can be expressed as:

[0163] h t = GlobalTransformer(E(s t1 ), E(a t2 ))

[0164] = GlobalTransformer(s 1 , …, s T1 , a 1 , …, aT2 )

[0165] Among them, E(s t1 ) is the quantized semantic feature, and E(a t2 ) is the quantized one-dimensional acoustic feature. h t is the intermediate acoustic feature. GlobalTransformer is used to represent the first-layer sub-model, where 1 ≤ t ≤ T 1 +T 2 .

[0166] Step 406: Input the intermediate acoustic feature values in the intermediate acoustic feature into the second-layer sub-model in sequence to obtain the synthesized audio feature.

[0167] The synthesized audio feature is used to represent the feature corresponding to the audio to be synthesized.

[0168] The second-layer sub-model is used to synthesize the synthesized audio feature based on the acoustic feature.

[0169] Optionally, the second-layer sub-model may adopt at least one of the attention network Transformer, BERT, and RNN, but is not limited thereto. The embodiments of the present application do not make specific limitations on this.

[0170] Exemplarily, the computer device inputs the intermediate acoustic feature values in the intermediate acoustic feature into the second-layer sub-model in sequence to obtain the synthesized audio feature.

[0171] Exemplarily, the computer device inputs the first intermediate acoustic feature value in the intermediate acoustic feature into the second-layer sub-model to predict the synthesized audio feature value in the synthesized audio feature, and obtains the first synthesized audio feature value in the synthesized audio feature. The synthesized audio feature value refers to the feature value in the synthesized audio feature. The computer device inputs the jth intermediate acoustic feature value generated from the (j - 1)th synthesized audio feature value into the second-layer sub-model to predict the synthesized audio feature value, and obtains the jth synthesized audio feature value. The jth intermediate acoustic feature value is predicted by inputting the (j - 1)th synthesized audio feature value into the first-layer sub-model. Repeat the previous step until the number of output synthesized audio feature values is equal to the number of intermediate acoustic feature values in the intermediate acoustic feature. The computer device combines the jth synthesized audio feature value with the previous (j - 1) synthesized audio feature values to obtain the synthesized audio feature, where j is a positive integer greater than 1.

[0172] Exemplarily, such as Figure 6Schematic diagram for obtaining synthetic audio features. The computer device inputs the first intermediate acoustic feature value h1 in the intermediate acoustic features 602 into the second-layer sub-model 603 to predict the synthetic audio feature value in the synthetic audio features 604, and obtains the first synthetic audio feature value in the synthetic audio features 604. The first synthetic audio feature value can be expressed as: [m11, m12, m13]; the computer device inputs the generated first synthetic audio feature value and the splicing feature into the first-layer sub-model 601 to obtain the second intermediate acoustic feature value h2; the computer device inputs the generated second intermediate acoustic feature value h2 into the second-layer sub-model 603 to predict the synthetic audio feature value, and obtains the second synthetic audio feature value. The second synthetic audio feature value can be expressed as: [m21, m22, m23]; the computer device inputs the generated first synthetic audio feature value, the second synthetic audio feature value and the splicing feature into the first-layer sub-model 601 to obtain the third intermediate acoustic feature value h3; the computer device inputs the generated third intermediate acoustic feature value h3 into the second-layer sub-model 603 to predict the synthetic audio feature value, and obtains the third synthetic audio feature value. The third synthetic audio feature value can be expressed as: [m31, m32, m33]; and so on, until the number of output synthetic audio feature values is equal to the number of intermediate acoustic feature values in the intermediate acoustic features 602 or until no intermediate acoustic feature value is input into the second-layer sub-model 603. The computer device merges the generated synthetic audio feature values to obtain the synthetic audio features 604.

[0173] The acoustic feature values in the acoustic features to be input into the second-layer sub-model are discrete numbers. For the convenience of calculation of the second-layer sub-model, it is necessary to perform mathematical processing on the acoustic features to convert them into vectors. The vectorization processing formula of the acoustic features can be expressed as:

[0174]

[0175] Where, refers to the vectorized acoustic features, refers to the acoustic feature value, 1 ≤ q ≤ D, q is the dimension of the acoustic feature value, D is the total dimension of the acoustic features, E a is the embedding function for the acoustic feature value, PE l is used to represent the position embedding function of the second-layer sub-model, and t is used to represent the time point or time position.

[0176] The calculation formula of the second-layer sub-model can be expressed as:

[0177]

[0178] Where, refers to the vectorized acoustic features, h tRefers to the intermediate acoustic feature, Refers to the acoustic feature value. LocalTransformer is used to represent the second-layer sub-model, where 1 ≤ t ≤ T 2 .

[0179] Step 408: Generate a synthetic audio with the same timbre as the reference speech based on the synthetic audio features.

[0180] The synthetic audio is an audio that imitates the voice of the object corresponding to the selected timbre (the human voice in the reference speech) to speak the text information to be expressed.

[0181] Exemplarily, the computer device decodes the synthetic audio features to obtain a synthetic audio with the same timbre as the reference speech, that is, obtains a synthetic audio expressed in the human voice of the reference speech.

[0182] In some embodiments, the synthetic audio features can be obtained by feature decoding through a decoder in the speech synthesis model.

[0183] Optionally, the decoder can be implemented by a Convolutional Neural Network (CNN), a Transformer neural network, or a Conformer network enhanced by convolution. For example, the decoder can be implemented by a six-layer Transformer neural network, and of course, the structure of the decoder is not limited to this implementation.

[0184] To verify the synthesis effect of the speech synthesis method provided by the embodiments of the present application, the present application compares the synthesis effect of the speech synthesis model provided by the embodiments of the present application with the synthesis effect of the models in the related art. The evaluation metrics selected in the embodiments of the present application are the Word Error Rate (WER), the Speech Similarity (SPK), and the Speech Quality (DNSMOS). Through these three metrics, the synthesis advantages of the speech synthesis model provided by the embodiments of the present application are reflected, as shown in the comparison results of the synthesis effects of the speech synthesis models in Table 1.

[0185] Table 1 Comparison results of the synthesis effects of the speech synthesis models

[0186]

[0187] Among them, it is divided into three groups of experiments. Related Model 1 and Related Model 2 refer to the models in the related art, and the speech synthesis model refers to the model provided by the embodiments of the present application. It can be seen from the table that the speech or audio synthesized by the speech synthesis model has the lowest word error rate, the highest speech similarity, and the best speech quality.

[0188] In summary, the method provided in this embodiment obtains semantic features and acoustic features corresponding to the reference speech; inputs the semantic features and acoustic features into the first-layer sub-model in the speech synthesis model for feature embedding to obtain intermediate acoustic features; inputs the intermediate acoustic features into the second-layer sub-model in the speech synthesis model for speech synthesis to obtain synthetic audio features; and decodes the synthetic audio features to obtain a synthetic audio with the same timbre as the reference speech. This application processes semantic features and acoustic features through two layers of sub-models in the speech synthesis model, enabling the finally generated synthetic audio features to learn both semantic features and acoustic features sufficiently. Compared with the method of rigidly imitating real speech, the voice output by the speech synthesis model in this method is more in line with the actual vocal characteristics of the object corresponding to the selected timbre, improving the authenticity of speech synthesis.

[0189] The method provided in this embodiment embeds semantic features and acoustic features through the first-layer sub-model in the speech synthesis model, enabling the semantic features to be imitated and expressed to be fused with the acoustic features, and generating more in line with the actual vocal characteristics of the object corresponding to the selected timbre based on the semantic features to be expressed and the acoustic features to be imitated, improving the authenticity of speech synthesis.

[0190] The method provided in this embodiment restores the acoustic information of the intermediate acoustic features output by the first-layer sub-model through the second-layer sub-model in the speech synthesis model, and sends the restoration result back to the first-layer sub-model for prediction. After cyclic prediction, synthetic audio features are finally obtained. By processing semantic features and acoustic features in a cyclic manner, the finally generated synthetic audio features can learn both semantic features and acoustic features sufficiently. Compared with the method of rigidly imitating real speech, the voice output by the speech synthesis model in this method is more in line with the actual vocal characteristics of the object corresponding to the selected timbre, improving the authenticity of speech synthesis.

[0191] The method provided in this embodiment extracts and quantifies the features of the obtained reference speech, enabling the acoustic features to more fully represent the acoustic aspects of the features, making the voice of the speaker more in line with the actual vocal characteristics of the object corresponding to the selected timbre, and improving the authenticity of speech synthesis.

[0192] The method provided in this embodiment integrates semantic features and acoustic features through two layers of sub-models, not only reducing the computational cost but also effectively learning the interaction relationship between semantic features and acoustic features.

[0193] The training method of the speech synthesis model involved in this application can be implemented based on the training system of the speech synthesis model. This solution includes the generation stage of the training system of the speech synthesis model and the training stage of the speech synthesis model. Figure 7It is a framework diagram for generating a training system of a speech synthesis model and training the speech synthesis model shown in an exemplary embodiment of the present application. As Figure 7 shown, in the stage of generating the training system of the speech synthesis model, after the training system generation device 710 of the speech synthesis model obtains the training system of the speech synthesis model through a preset training sample data set, the training result of the speech synthesis model is generated based on the training system of the speech synthesis model. In the training stage of the speech synthesis model, the training device 720 of the speech synthesis model processes the input audio signal based on the training system of the speech synthesis model to obtain the training result of the speech synthesis model.

[0194] Among them, the above-mentioned training system generation device 710 of the speech synthesis model and the training device 720 of the speech synthesis model can be computer devices. For example, the computer device can be a fixed computer device such as a personal computer or a server, or the computer device can also be a mobile computer device such as a tablet computer or an e-book reader.

[0195] Optionally, the above-mentioned training system generation device 710 of the speech synthesis model and the training device 720 of the speech synthesis model can be the same device, or the training system generation device 710 of the speech synthesis model and the training device 720 of the speech synthesis model can also be different devices. And when the training system generation device 710 of the speech synthesis model and the training device 720 of the speech synthesis model are different devices, the training system generation device 710 of the speech synthesis model and the training device 720 of the speech synthesis model can be the same type of device. For example, the training system generation device 710 of the speech synthesis model and the training device 720 of the speech synthesis model can both be servers; or the training system generation device 710 of the speech synthesis model and the training device 720 of the speech synthesis model can also be different types of devices. For example, the training device 720 of the speech synthesis model can be a personal computer or a terminal, and the training system generation device 710 of the speech synthesis model can be a server, etc. The specific types of the training system generation device 710 of the speech synthesis model and the training device 720 of the speech synthesis model are not limited in the embodiments of the present application.

[0196] The above embodiments have described the speech synthesis method. Next, the training method of the speech synthesis model will be described.

[0197] Figure 8 It is a flowchart of the training method of the speech synthesis model provided by an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be a terminal or a server, and a speech synthesis model is set in the computer device. This method includes:

[0198] Step 802: Obtain the sample semantic features, sample acoustic features, and the sample audio corresponding to the sample semantic features.

[0199] The sample semantic features are the features representing the semantic information corresponding to the audio to be synthesized; or, the sample semantic features are the features representing the text information corresponding to the audio to be synthesized.

[0200] Optionally, the method for obtaining the sample semantic features includes: performing speech recognition on the speech signal corresponding to the audio to be synthesized to obtain the text content; extracting features from the text content to obtain the sample semantic features; or, performing character recognition and feature extraction on the characters in the text content to obtain the sample semantic features; or, directly extracting features from the speech signal to obtain the sample semantic features representing the semantics of the speech signal.

[0201] Optionally, the text content includes at least one of phrases, sentences, paragraphs, and passages composed of characters and / or symbols, etc. Optionally, the language types of the text content are not limited to Chinese and English, and can be any one or more languages.

[0202] The sample acoustic features are the features of the acoustic information corresponding to the sample reference speech.

[0203] Optionally, the sample acoustic features include at least one of recording environment features, timbre features, and prosody duration features, but are not limited thereto, and the embodiments of the present application do not make specific limitations thereto.

[0204] Among them, the recording environment features are used to characterize the features of the recording environment of the sample reference speech; the timbre features are used to characterize the timbre features of the object corresponding to the selected timbre; the prosody duration features are used to characterize the prosody features in the sample reference speech. For example, the prosody duration features are used to characterize features such as the pitch, duration, and pitch height of the object corresponding to the selected timbre when speaking, or the prosody duration features are used to characterize the cadence features of the object corresponding to the selected timbre when speaking.

[0205] The sample reference speech refers to the speech of the object corresponding to the selected timbre.

[0206] Optionally, the sample reference speech is the speech of the object corresponding to the selected timbre; or, the sample reference speech is the speech when the object corresponding to the selected timbre reads the text; or, the sample reference speech is an audio containing the voice of the object corresponding to the selected timbre, but is not limited thereto. For example, the sample reference speech is the recording of person A reading a text.

[0207] Optionally, the sample reference speech can be a 3-second-long speech.

[0208] The sample audio refers to the audio obtained by the object corresponding to the selected timbre for the text information or semantic information.

[0209] In some embodiments, the sample acoustic features can be obtained by feature extraction through an acoustic encoder (also known as an acoustic feature extraction network) in a speech synthesis model.

[0210] Optionally, the acoustic encoder can be implemented by a Convolutional Neural Network (CNN), a Transformer neural network, or a Conformer network. For example, the acoustic encoder can be implemented using a four-layer Transformer neural network. Of course, the acquisition of acoustic features is not limited to this implementation.

[0211] In some embodiments, the sample semantic features can be obtained by feature extraction through a semantic encoder in a speech synthesis model.

[0212] Optionally, the semantic encoder can be implemented by a Self-Supervised Learning (SSL) encoder and a k-means clustering algorithm. Of course, the acquisition of semantic features is not limited to this implementation.

[0213] In some embodiments, the computer device can splice the speech signals corresponding to the sample reference speech and the audio to be synthesized to obtain a spliced speech signal, and the computer device extracts features from the spliced speech signal to obtain the sample semantic features and the sample acoustic features.

[0214] Optionally, obtain the speech signals corresponding to the sample reference speech with a first time length and the audio to be synthesized with a second time length; the computer device splices the speech signals corresponding to the sample reference speech and the audio to be synthesized to obtain a spliced speech signal; the computer device extracts features from the reference speech with the first time length through an acoustic feature extraction network to obtain the sample acoustic features; the computer device extracts features from the speech signals corresponding to the audio to be synthesized with the second time length through a semantic feature extraction network to obtain the sample semantic features.

[0215] For example, obtain the speech signals corresponding to a 3 - second sample reference speech and a 7 - second audio - to - be - synthesized. The computer device splices the 3 - second sample reference speech and the 7 - second speech signal to obtain a 10 - second spliced speech signal. The computer device extracts features from the 3 - second sample reference speech through an acoustic feature extraction network to obtain sample acoustic features. The computer device extracts features from the 7 - second speech signal through a semantic feature extraction network to obtain sample semantic features. The computer device finally obtains the synthesized audio based on the sample semantic features and the sample acoustic features, that is, a synthesized audio that can be expressed in the voice of the object corresponding to the selected timbre through a short 3 - second sample reference speech. For example, if person A wants to imitate person B to read an article, only need to obtain the 3 - second speech of person B as the reference speech and the speech signal of person A reading an article. After splicing the 3 - second speech of person B and the speech signal of person A reading an article, the computer device inputs them into the speech synthesis model provided by the embodiment of the present application, and then person A can be obtained reading an article in the voice of person B.

[0216] Step 804: Embed the sample semantic features into the sample acoustic features through the first - layer sub - model in the speech synthesis model to obtain sample intermediate acoustic features.

[0217] The sample intermediate acoustic features are used to represent the features obtained after embedding the sample semantic features into the sample acoustic features.

[0218] Embedding the sample semantic features into the sample acoustic features means that through the first - layer sub - model, the sample acoustic features can learn the semantic information or text information in the sample semantic features, that is, can obtain the text information that the audio - to - be - synthesized wants to express.

[0219] The speech synthesis model includes a first - layer sub - model and a second - layer sub - model. Among them, the model structures of the first - layer sub - model and the second - layer sub - model are the same, but at least one of their network parameters, execution tasks, and number of network layers is different.

[0220] The first - layer sub - model is used to embed semantic features into acoustic features.

[0221] Optionally, the first - layer sub - model can adopt at least one of an attention network Transformer, a pre - trained language representation model (Bidirectional Encoder Representation from Transformers, BERT), and a recurrent neural network (Recurrent Neural Networks, RNN), but is not limited thereto. The embodiments of the present application do not make specific limitations on this.

[0222] Step 806: Input the sample intermediate acoustic features into the second - layer sub - model in the speech synthesis model to obtain synthesized audio features.

[0223] The synthetic audio features are used to represent the features corresponding to the audio to be synthesized.

[0224] The second-layer sub-model is used to synthesize synthetic audio features based on the sample intermediate acoustic features.

[0225] Optionally, the second-layer sub-model may adopt at least one of the attention network Transformer, BERT, and RNN, but is not limited thereto, and the embodiments of the present application do not make specific limitations in this regard.

[0226] Step 808: Generate a synthetic audio with the same timbre as the sample reference speech based on the synthetic audio features.

[0227] The synthetic audio is an audio that imitates the voice of the object corresponding to the selected timbre (the human voice in the sample reference speech) to speak out the text information to be expressed.

[0228] Exemplarily, the computer device decodes the synthetic audio features, so as to obtain a synthetic audio with the same timbre as the sample reference speech, that is, obtain a synthetic audio expressed in the human voice of the sample reference speech.

[0229] Step 810: Calculate the training loss of the speech synthesis model based on the sample audio and the synthetic audio.

[0230] Exemplarily, the computer device calculates the training loss of the speech synthesis model based on the sample audio and the synthetic audio.

[0231] The training loss refers to the difference value between the input and output of the speech synthesis model, and the performance of the speech synthesis model is measured by the training loss.

[0232] Step 812: Update the model parameters of the speech synthesis model according to the training loss.

[0233] Exemplarily, the computer device updates the model parameters of the speech synthesis model according to the training loss.

[0234] The model parameter update means updating the network parameters in the speech synthesis model, or updating the network parameters of each network module in the model, or updating the network parameters of each network layer in the model, but is not limited thereto, and the embodiments of the present application do not make limitations in this regard.

[0235] In summary, the method provided in this embodiment obtains sample semantic features, sample acoustic features, and sample audio; embeds the sample semantic features into the sample acoustic features through the first-layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features; inputs the sample intermediate acoustic features and the sample acoustic features into the second-layer sub-model in the speech synthesis model to obtain synthesized audio features; generates synthesized audio with the same timbre as the sample reference speech based on the synthesized audio features; calculates the training loss of the speech synthesis model based on the sample audio and the synthesized audio; and updates the model parameters of the speech synthesis model according to the training loss. By processing the semantic features and acoustic features, the synthesized audio features finally generated in this application can not only learn the semantic features but also fully learn the acoustic features, and the speech synthesis model is trained by the difference between the sample audio and the synthesized audio. Based on this, the synthesis effect of the speech synthesis model can be improved.

[0236] Figure 9 FIG. 4 is a flowchart of a method for training a speech synthesis model provided by an exemplary embodiment of the present application. This method can be executed by a computer device, which can be a terminal or a server, and a speech synthesis model is set in the computer device. The method includes:

[0237] Step 902: Obtain sample semantic features, sample acoustic features, and sample audio corresponding to the sample semantic features.

[0238] The sample semantic features are features used to represent the semantic information corresponding to the audio to be synthesized; or, the sample semantic features are features used to represent the text information corresponding to the audio to be synthesized.

[0239] Optionally, the obtaining method of the sample semantic features includes: performing speech recognition on the speech signal to obtain the text content; extracting features from the text content to obtain the sample semantic features; or, performing character recognition and feature extraction on the characters in the text content to obtain the sample semantic features; or, directly extracting features from the speech signal to obtain the sample semantic features used to represent the semantics of the speech signal.

[0240] The sample acoustic features are features of the acoustic information corresponding to the sample reference speech. The sample acoustic features can characterize the vocal characteristics, reading characteristics of the object corresponding to the selected timbre, or the characteristics of the reference speech recording environment, etc.

[0241] Optionally, the sample acoustic features include at least one of recording environment features, timbre features, and prosody duration features, but are not limited thereto, and the embodiments of the present application do not make specific limitations on this.

[0242] Among them, the recording environment feature is used to characterize the feature of the sample reference speech recording environment; the timbre feature is used to characterize the timbre feature of the object corresponding to the selected timbre; the prosody duration feature is used to characterize the prosody feature in the reference speech. For example, the prosody duration feature is used to characterize the pitch, duration, pitch height, etc. of the object corresponding to the selected timbre when speaking, or the prosody duration feature is used to characterize the cadence feature of the object corresponding to the selected timbre when speaking.

[0243] The sample reference speech refers to the speech of the object corresponding to the selected timbre.

[0244] Optionally, the sample reference speech is the speech of the object corresponding to the selected timbre; or, the sample reference speech is the speech when the object corresponding to the selected timbre reads a text; or, the sample reference speech is an audio segment containing the voice of the object corresponding to the selected timbre, but not limited to this. For example, the sample reference speech is the recording of person A reading a text.

[0245] In some embodiments, the computer device obtains the sample reference speech; the computer device inputs the sample reference speech into the acoustic feature extraction network in the speech synthesis model to extract features, and obtains acoustic features.

[0246] Optionally, the computer device inputs the sample reference speech into the acoustic feature extraction network to extract features, and obtains the initial sample acoustic features; the computer device quantizes the initial sample acoustic features to obtain the sample acoustic features.

[0247] The initial sample acoustic features refer to the continuous features extracted for the sample reference speech.

[0248] Since the initial sample acoustic features are continuous features, in order to facilitate subsequent calculations, it is necessary to quantize the initial sample acoustic features to obtain the sample acoustic features. Among them, the initial sample acoustic features before quantization are a one-dimensional matrix, and the sample acoustic features after quantization are an n-dimensional matrix, where n is a positive integer greater than 1. The sample acoustic features after quantization have higher expressiveness than the initial sample acoustic features before quantization, that is, the multi-dimensional matrix can more fully represent the sample acoustic features.

[0249] For example, the initial sample acoustic features before quantization can be expressed as [a1, a2, a3], and the sample acoustic features after quantization are a three-dimensional matrix, which can be expressed as:

[0250] The steps of quantifying the initial sample acoustic features into sample acoustic features include: The computer device quantifies the m-th acoustic feature value in the initial sample acoustic features to obtain the first quantified acoustic feature value in the first column dimension of the sample acoustic features, where m is a positive integer; The computer device quantifies the residual between the first quantified acoustic feature value and the m-th acoustic feature value to obtain the second quantified acoustic feature value in the m-th column dimension; The computer device quantifies the residual between the k-th quantified acoustic feature value and the (k - 1)-th quantified acoustic feature value to obtain the (k + 1)-th quantified acoustic feature value in the m-th column dimension, where k is a positive integer greater than 1; The computer device combines the first k + 1 quantified acoustic feature values in the m-th column dimension to obtain the quantified acoustic feature value in the m-th column dimension of the sample acoustic features; Repeat the above steps, combine the quantified acoustic feature values of each column dimension, and obtain the sample acoustic features.

[0251] For example, the computer device quantifies the first acoustic feature value a1 in the initial sample acoustic features to obtain the first quantified acoustic feature value a11 in the first column dimension of the sample acoustic features; The computer device quantifies the residual between the first acoustic feature value a1 and the first quantified acoustic feature value a11 in the first column dimension to obtain the second quantified acoustic feature value a12 in the first column dimension; The computer device quantifies the residual between the second acoustic feature value a12 and the first quantified acoustic feature value a11 in the first column dimension to obtain the third quantified acoustic feature value a13 in the first column dimension; Combine the quantified acoustic feature values in the first column dimension to obtain the acoustic feature in the first column dimension of the sample acoustic features; Perform the above steps on other column dimensions to obtain the acoustic features of other column dimensions in the sample acoustic features; Combine the quantified acoustic feature values of each column dimension to obtain the sample acoustic features.

[0252] Optionally, the dimension of the sample acoustic features can be set artificially according to needs.

[0253] In some embodiments, the ways of obtaining the sample reference speech include at least one of the following situations:

[0254] 1. The computer device receives the sample reference speech. For example, the terminal is the terminal that initiates audio recording. The terminal records audio and, after the recording ends, uses this audio as the sample reference speech.

[0255] 2. The computer device obtains the sample reference speech from a stored database.

[0256] It should be noted that the above ways of obtaining the sample reference speech are only illustrative examples, and the embodiments of the present application are not limited thereto.

[0257] Step 904: Concatenate the sample semantic feature and the one-dimensional sample acoustic feature to obtain a sample concatenated feature; input the sample concatenated feature into the first-layer sub-model to obtain a sample intermediate acoustic feature.

[0258] The sample concatenated feature refers to the feature obtained by concatenating the sample semantic feature and the one-dimensional sample acoustic feature.

[0259] The sample acoustic feature is the feature of the acoustic information corresponding to the sample reference speech. The sample acoustic feature is an n-dimensional matrix, where n is a positive integer greater than 1.

[0260] Exemplarily, the computer device adds the feature values of the same column dimension in the sample acoustic feature to obtain a one-dimensional sample acoustic feature.

[0261] For example, the sample acoustic feature is a three-dimensional matrix, which can be expressed as: The computer device adds the feature values of the first column dimension in the sample acoustic feature. For example, it adds the numerical values of a11, a12, and a13 to obtain the first acoustic feature value A1 in the one-dimensional sample acoustic feature; similarly, it adds the numerical values of a21, a22, and a23 to obtain the second acoustic feature value A2 in the one-dimensional sample acoustic feature; it adds the numerical values of a31, a32, and a33 to obtain the third acoustic feature value A3 in the one-dimensional sample acoustic feature. Thus, the obtained one-dimensional sample acoustic feature can be expressed as [A1, A2, A3].

[0262] In some embodiments, the sample semantic feature can be expressed as [s1, s2, s3]. The computer device concatenates the sample semantic feature and the one-dimensional sample acoustic feature to obtain a sample concatenated feature, and the obtained sample concatenated feature can be expressed as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the sample concatenated feature into the first-layer sub-model to predict the sample intermediate acoustic feature value in the sample intermediate acoustic feature, and obtains the first sample intermediate acoustic feature value in the sample intermediate acoustic feature; the computer device inputs the (i - 1)-th synthesized audio feature value corresponding to the (i - 1)-th sample intermediate acoustic feature value and the concatenated feature into the first-layer sub-model to predict the intermediate acoustic feature value, and obtains the i-th sample intermediate acoustic feature value. The (i - 1)-th synthesized audio feature value is obtained by inputting the (i - 1)-th sample intermediate acoustic feature value into the second-layer sub-model for prediction; the computer device loops the previous step until the number of output sample intermediate acoustic feature values is equal to the number of feature values in the one-dimensional sample acoustic feature; the computer device combines the i-th sample intermediate acoustic feature value with the previous (i - 1) sample intermediate acoustic feature values to obtain the sample intermediate acoustic feature, where i is a positive integer greater than 1.

[0263] Exemplarily, the computer device adds the feature values of the same column dimension in the sample acoustic features to obtain one-dimensional sample acoustic features. The sample semantic features can be expressed as [s1, s2, s3]. The computer device splices the sample semantic features and the one-dimensional sample acoustic features to obtain sample spliced features, and the obtained sample spliced features can be expressed as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the sample spliced features into the first-layer sub-model to obtain the first sample intermediate acoustic feature value in the sample intermediate acoustic features; the computer device inputs the first synthesized audio feature value corresponding to the generated first sample intermediate acoustic feature value and the sample spliced features into the first-layer sub-model to predict the sample intermediate acoustic feature value, and obtains the second sample intermediate acoustic feature value, where the first synthesized audio feature value is predicted by inputting the first sample intermediate acoustic feature value into the second-layer sub-model; the computer device inputs the second synthesized audio feature value corresponding to the generated second sample intermediate acoustic feature value and the sample spliced features into the first-layer sub-model to predict the sample intermediate acoustic feature value, and obtains the third intermediate acoustic feature value, where the second synthesized audio feature value is predicted by inputting the second intermediate acoustic feature value into the second-layer sub-model; and so on, until the number of sample intermediate acoustic feature values output by the first-layer sub-model is equal to the number of feature values in the one-dimensional sample acoustic features. For example, the one-dimensional sample acoustic features can be expressed as [Bos, A1, A2, A3], which includes 3 feature values. When the first-layer sub-model 504 outputs 3 sample intermediate acoustic feature values, the first-layer sub-model ends the prediction; the computer device merges the generated sample intermediate acoustic feature values to obtain sample intermediate acoustic features, which can be expressed as [h1, h2, h3, Eos].

[0264] The speech synthesis model includes a first-layer sub-model and a second-layer sub-model. Among them, the model structures of the first-layer sub-model and the second-layer sub-model are the same, but at least one of their network parameters, execution tasks, and number of network layers is different.

[0265] The first-layer sub-model is used to embed the sample semantic features into the sample acoustic features.

[0266] Optionally, the first-layer sub-model can adopt at least one of an attention network Transformer, a pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT), and a recurrent neural network (Recurrent Neural Networks, RNN), but is not limited thereto. The embodiments of the present application do not make specific limitations on this.

[0267] Step 906: Input the sample intermediate acoustic feature values in the sample intermediate acoustic features into the second-layer sub-model in sequence to obtain synthetic audio features.

[0268] The synthetic audio features are used to represent the features corresponding to the audio to be synthesized.

[0269] The second-layer sub-model is used to synthesize synthetic audio features based on the sample intermediate acoustic features.

[0270] Optionally, the second-layer sub-model may adopt at least one of the attention network Transformer, BERT, and RNN, but is not limited thereto. The embodiments of the present application do not make specific limitations in this regard.

[0271] Exemplarily, the computer device inputs the sample intermediate acoustic feature values in the sample intermediate acoustic features into the second-layer sub-model in sequence to obtain synthetic audio features.

[0272] Exemplarily, the computer device inputs the first sample intermediate acoustic feature value in the sample intermediate acoustic features into the second-layer sub-model to predict the synthetic audio feature value in the synthetic audio features, and obtains the first synthetic audio feature value in the synthetic audio features. The synthetic audio feature value refers to the feature value in the synthetic audio features. The computer device inputs the jth sample intermediate acoustic feature value generated from the (j - 1)th synthetic audio feature value into the second-layer sub-model to predict the synthetic audio feature value, and obtains the jth synthetic audio feature value. The jth sample intermediate acoustic feature value is predicted by inputting the (j - 1)th sample synthetic audio feature value into the first-layer sub-model. Repeat the previous step until the number of output synthetic audio feature values is equal to the number of sample intermediate acoustic feature values in the sample intermediate acoustic features. The computer device combines the jth synthetic audio feature value with the previous (j - 1) synthetic audio feature values to obtain the synthetic audio features, where j is a positive integer greater than 1.

[0273] Exemplarily, the computer device inputs the first sample intermediate acoustic feature value in the sample intermediate acoustic features into the second-layer sub-model to predict the synthetic audio feature value in the synthetic audio features, and obtains the first synthetic audio feature value in the synthetic audio features; the computer device inputs the generated first synthetic audio feature value and the sample splicing features into the first-layer sub-model, and obtains the second sample intermediate acoustic feature value; the computer device inputs the generated second sample intermediate acoustic feature value into the second-layer sub-model to predict the synthetic audio feature value, and obtains the second synthetic audio feature value; the computer device inputs the generated first synthetic audio feature value, the second synthetic audio feature value and the sample splicing features into the first-layer sub-model, and obtains the third sample intermediate acoustic feature value; the computer device inputs the generated third sample intermediate acoustic feature value into the second-layer sub-model to predict the synthetic audio feature value, and obtains the third synthetic audio feature value; and so on, until the number of the output synthetic audio feature values is equal to the number of the intermediate acoustic feature values in the intermediate acoustic features or until there is no sample intermediate acoustic feature value input into the second-layer sub-model, and the computer device merges the generated synthetic audio feature values to obtain the synthetic audio features.

[0274] Step 908: Generate a synthetic audio with the same timbre as the sample reference speech based on the synthetic audio features.

[0275] The synthetic audio is an audio that imitates the voice of the object corresponding to the selected timbre (the human voice in the reference speech) to speak out the text information to be expressed.

[0276] Exemplarily, the computer device decodes the synthetic audio features, so as to obtain a synthetic audio with the same timbre as the sample reference speech, that is, to obtain a synthetic audio expressed in the human voice of the sample reference speech.

[0277] In some embodiments, the synthetic audio features can be obtained by feature decoding through a decoder in the speech synthesis model.

[0278] Optionally, the decoder can be implemented by a Convolutional Neural Network (CNN), a Transformer neural network, or a Conformer network enhanced by convolution. For example, the decoder can be implemented by a six-layer Transformer neural network, and of course the structure of the decoder is not limited to this implementation.

[0279] Step 910: Calculate the training loss of the speech synthesis model based on the sample audio and the synthetic audio.

[0280] Exemplarily, the computer device calculates the training loss of the speech synthesis model based on the sample audio and the synthetic audio.

[0281] The training loss refers to the difference value between the input and output of the speech synthesis model, and the performance of the speech synthesis model is measured by the training loss.

[0282] Step 912: Update the model parameters of the speech synthesis model according to the training loss.

[0283] Exemplarily, the computer device updates the model parameters of the speech synthesis model according to the training loss.

[0284] Model parameter update means updating the network parameters in the speech synthesis model, or updating the network parameters of each network module in the model, or updating the network parameters of each network layer in the model, but not limited to this, and the embodiments of the present application do not make limitations in this regard.

[0285] Based on the loss function value, the loss function value is used as a training metric to update the model parameters of the first-layer sub-model and the second-layer sub-model in the speech synthesis model until the loss function value converges, so as to obtain a trained speech synthesis model.

[0286] The convergence of the loss function value means that the loss function value no longer changes, or the error difference between two adjacent iterations during the training of the speech synthesis model is less than a preset value, or the number of training times of the speech synthesis model reaches at least one of the preset times, but not limited to this, and the embodiments of the present application do not make limitations in this regard.

[0287] Optionally, the target condition satisfied by the training can be that the number of training iterations of the initial model reaches the target number, and those skilled in the art can preset the number of training iterations. Or, the target condition satisfied by the training can be that the loss value satisfies the target threshold condition, such as the loss value is less than 0.00001, but not limited to this, and the embodiments of the present application do not make limitations in this regard.

[0288] In summary, the method provided in this embodiment obtains sample semantic features, sample acoustic features, and sample audio; embeds the sample semantic features into the sample acoustic features through the first-layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features; inputs the sample intermediate acoustic features into the second-layer sub-model in the speech synthesis model to obtain synthesized audio features; generates a synthesized audio with the same timbre as the sample reference speech based on the synthesized audio features; calculates the training loss of the speech synthesis model based on the sample audio and the synthesized audio; updates the model parameters of the speech synthesis model according to the training loss. By processing the semantic features and acoustic features, the present application enables the finally generated synthesized audio features to not only learn the semantic features but also fully learn the acoustic features, and trains the speech synthesis model through the difference between the sample audio and the synthesized audio. Based on this, the synthesis effect of the speech synthesis model can be improved.

[0289] Figure 10The figure shows a schematic structural diagram of a speech synthesis device provided by an exemplary embodiment of the present application. The device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:

[0290] An acquisition module 1001, configured to acquire semantic features and acoustic features. The semantic features are used to represent the features of the text information corresponding to the audio to be synthesized, and the acoustic features are the features of the acoustic information corresponding to the reference speech. The reference speech refers to the speech of the object corresponding to the selected timbre;

[0291] A feature processing module 1002, configured to embed the semantic features into the acoustic features through the first-layer sub-model in the speech synthesis model to obtain intermediate acoustic features. The first-layer sub-model is used to embed the semantic features into the acoustic features, and the intermediate acoustic features are used to represent the features obtained after embedding the semantic features into the acoustic features;

[0292] The feature processing module 1002 is further configured to input the intermediate acoustic features into the second-layer sub-model in the speech synthesis model to obtain synthesized audio features. The second-layer sub-model is used to synthesize the synthesized audio features based on the intermediate acoustic features, and the synthesized audio features are used to represent the features corresponding to the audio to be synthesized;

[0293] A generation module 1003, configured to generate a synthesized audio with the same timbre as the reference speech based on the synthesized audio features.

[0294] In some embodiments, the feature processing module 1002 is further configured to add the feature values of the same column dimension in the acoustic features to obtain a one-dimensional acoustic feature; splice the semantic features and the one-dimensional acoustic feature to obtain a spliced feature, where the spliced feature refers to the feature obtained by splicing the semantic features and the one-dimensional acoustic feature; and input the spliced feature into the first-layer sub-model to obtain the intermediate acoustic features.

[0295] In some embodiments, the feature processing module 1002 is further configured to input the spliced feature into the first-layer sub-model to predict the intermediate acoustic feature value in the intermediate acoustic feature, so as to obtain the first intermediate acoustic feature value in the intermediate acoustic feature, where the intermediate acoustic feature refers to the feature predicted by the first-layer sub-model based on the spliced feature, and the intermediate acoustic feature value refers to the feature value in the intermediate acoustic feature; input the (i-1)-th synthetic audio feature value corresponding to the (i-1)-th intermediate acoustic feature value and the spliced feature into the first-layer sub-model to predict the intermediate acoustic feature value, so as to obtain the i-th intermediate acoustic feature value, where the (i-1)-th synthetic audio feature value is obtained by inputting the (i-1)-th intermediate acoustic feature value into the second-layer sub-model; repeat the previous step until the number of the output intermediate acoustic feature values is equal to the number of the feature values in the one-dimensional acoustic feature; combine the i-th intermediate acoustic feature value with the previous i-1 intermediate acoustic feature values to obtain the intermediate acoustic feature value, and i is a positive integer greater than 1.

[0296] In some embodiments, the feature processing module 1002 is further configured to sequentially input the intermediate acoustic feature values in the intermediate acoustic feature into the second-layer sub-model to obtain the synthetic audio feature.

[0297] In some embodiments, the feature processing module 1002 is further configured to input the first intermediate acoustic feature value in the intermediate acoustic feature into the second-layer sub-model to predict the synthetic audio feature value in the synthetic audio feature, so as to obtain the first synthetic audio feature value in the synthetic audio feature, where the synthetic audio feature value refers to the feature value in the synthetic audio feature; input the j-th intermediate acoustic feature value generated from the (j-1)-th synthetic audio feature value into the second-layer sub-model to predict the synthetic audio feature value, so as to obtain the j-th synthetic audio feature value, where the j-th intermediate acoustic feature value is obtained by inputting the (j-1)-th synthetic audio feature value into the first-layer sub-model; repeat the previous step until the number of the output synthetic audio feature values is equal to the number of the intermediate acoustic feature values in the intermediate acoustic feature; combine the j-th synthetic audio feature value with the previous j-1 synthetic audio feature values to obtain the synthetic audio feature, and j is a positive integer greater than 1.

[0298] In some embodiments, the acquisition module 1001 is further configured to acquire the reference speech; extract features from the reference speech to obtain the acoustic feature.

[0299] In some embodiments, the device further includes a calculation module 1004, and the calculation module 1004 is further configured to extract features from the reference speech input into the acoustic feature extraction network to obtain initial acoustic features, where the initial acoustic features refer to continuous features extracted for the reference speech; and quantize the initial acoustic features to obtain the acoustic features.

[0300] In some embodiments, the calculation module 1004 is further configured to quantize the m-th acoustic feature value in the initial acoustic features to obtain the first quantized acoustic feature value in the m-th column dimension of the acoustic features, where m is a positive integer; quantize the residual between the first quantized acoustic feature value and the m-th acoustic feature value to obtain the second quantized acoustic feature value in the m-th column dimension; quantize the residual between the k-th quantized acoustic feature value and the (k - 1)-th quantized acoustic feature value to obtain the (k + 1)-th quantized acoustic feature value in the m-th column dimension, where k is a positive integer greater than 1; combine the first k + 1 quantized acoustic feature values in the m-th column dimension to obtain the quantized acoustic feature value in the m-th column dimension of the acoustic features; repeat the above steps, and combine the quantized acoustic feature values of each column dimension to obtain the acoustic features.

[0301] Figure 11 The figure shows a schematic structural diagram of a speech synthesis device provided by an exemplary embodiment of the present application. The device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:

[0302] An acquisition module 1101, configured to acquire a sample semantic feature, a sample acoustic feature, and a sample audio, where the sample semantic feature is used to represent the feature of the text information corresponding to the audio to be synthesized, the sample acoustic feature is the feature of the acoustic information corresponding to the sample reference speech, and the sample reference speech refers to the speech of the object corresponding to the selected timbre;

[0303] A feature processing module 1102, configured to embed the sample semantic feature into the sample acoustic feature through the first-layer sub-model in the speech synthesis model to obtain a sample intermediate acoustic feature, where the first-layer sub-model is used to embed the sample semantic feature into the sample acoustic feature, and the sample intermediate acoustic feature is used to represent the feature obtained after embedding the sample semantic feature into the sample acoustic feature;

[0304] The feature processing module 1102 is further configured to input the sample intermediate acoustic feature into the second-layer sub-model in the speech synthesis model to obtain a synthesized audio feature, where the second-layer sub-model is used to synthesize the synthesized audio feature based on the sample intermediate acoustic feature, and the synthesized audio feature is used to represent the feature corresponding to the audio to be synthesized;

[0305] A generation module 1103, configured to generate a synthesized audio with the same timbre as the sample reference speech based on the synthesized audio features;

[0306] A calculation module 1104, configured to calculate a training loss of the speech synthesis model based on the sample audio and the synthesized audio;

[0307] An update module 1105, configured to update model parameters of the speech synthesis model according to the training loss.

[0308] In some embodiments, the feature processing module 1102 is further configured to add the feature values of the same column dimension in the sample acoustic features to obtain one-dimensional sample acoustic features; splice the sample semantic features and the one-dimensional sample acoustic features to obtain sample splicing features, where the sample splicing features refer to the features obtained by splicing the sample semantic features and the one-dimensional sample acoustic features; and input the sample splicing features into the first-layer sub-model to obtain the sample intermediate acoustic features.

[0309] In some embodiments, the feature processing module 1102 is further configured to input the sample splicing features into the first-layer sub-model to predict the sample intermediate acoustic feature value in the sample intermediate acoustic features, and obtain the first sample intermediate acoustic feature value in the sample intermediate acoustic features, where the intermediate acoustic features refer to the features predicted by the first-layer sub-model based on the splicing features, and the sample intermediate acoustic feature value refers to the feature value in the sample intermediate acoustic features; input the (i-1)th synthesized audio feature value corresponding to the (i-1)th sample intermediate acoustic feature value and the sample splicing features into the first-layer sub-model to predict the sample intermediate acoustic feature value, and obtain the ith sample intermediate acoustic feature value, where the (i-1)th synthesized audio feature value is obtained by inputting the (i-1)th sample intermediate acoustic feature value into the second-layer sub-model; repeat the previous step until the number of the output sample intermediate acoustic feature values is equal to the number of the feature values in the one-dimensional sample acoustic features; and combine the ith sample intermediate acoustic feature value with the previous (i-1) sample intermediate acoustic feature values to obtain the sample intermediate acoustic features, where i is a positive integer greater than 1.

[0310] In some embodiments, the feature processing module 1102 is further configured to sequentially input the sample intermediate acoustic feature values in the sample intermediate acoustic features into the second-layer sub-model to obtain the synthesized audio features.

[0311] In some embodiments, the feature processing module 1102 is further configured to input the first sample intermediate acoustic feature value in the sample intermediate acoustic features into the second-layer sub-model to predict the synthetic audio feature value in the synthetic audio features, so as to obtain the first synthetic audio feature value in the synthetic audio features; input the jth sample intermediate acoustic feature value generated from the (j-1)th synthetic audio feature value into the second-layer sub-model to predict the synthetic audio feature value, so as to obtain the jth synthetic audio feature value in the synthetic audio features, where the jth sample intermediate acoustic feature value is obtained by inputting the (j-1)th sample synthetic audio feature value into the first-layer sub-model for prediction; repeat the previous step until the number of the output synthetic audio feature values is equal to the number of the sample intermediate acoustic feature values in the sample intermediate acoustic features; merge the jth synthetic audio feature value with the previous j-1 synthetic audio feature values to obtain the synthetic audio features, where j is a positive integer greater than 1.

[0312] In some embodiments, the obtaining module 1101 is further configured to obtain the sample reference speech; input the sample reference speech into the acoustic feature extraction network in the speech synthesis model to extract features, so as to obtain the sample acoustic features.

[0313] In some embodiments, the calculation module 1104 is further configured to input the sample reference speech into the acoustic feature extraction network to extract features, so as to obtain the initial sample acoustic features, where the initial sample acoustic features refer to the continuous features extracted for the sample reference speech; quantize the initial sample acoustic features to obtain the sample acoustic features.

[0314] In some embodiments, the calculation module 1104 is further configured to quantize the mth acoustic feature value in the initial sample acoustic features to obtain the first quantized acoustic feature value in the mth column dimension of the sample acoustic features, where m is a positive integer; quantize the residual between the first sample quantized acoustic feature value and the mth sample acoustic feature value to obtain the second sample quantized acoustic feature value in the mth column dimension; quantize the residual between the kth quantized acoustic feature value and the (k-1)th quantized acoustic feature value to obtain the (k+1)th quantized acoustic feature value in the mth column dimension, where k is a positive integer greater than 1; merge the first k+1 quantized acoustic feature values in the mth column dimension to obtain the quantized acoustic feature value in the mth column dimension of the sample acoustic features; repeat the above steps, and merge the quantized acoustic feature values of each column dimension to obtain the sample acoustic features.

[0315] Figure 12The block diagram of a computer device 1200 shown in an exemplary embodiment of the present application is presented. This computer device can be implemented as the server in the above solution of the present application. The computer device 1200 includes a Central Processing Unit (CPU) 1201, a system memory 1204 including a Random Access Memory (RAM) 1202 and a Read-Only Memory (ROM) 1203, and a system bus 1205 connecting the system memory 1204 and the central processing unit 1201. The computer device 1200 also includes a mass storage device 1206 for storing an operating system 1209, application programs 1210, and other program modules 1211.

[0316] The mass storage device 1206 is connected to the central processing unit 1201 through a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1206 and its associated computer-readable medium provide non-volatile storage for the computer device 1200. That is to say, the mass storage device 1206 can include computer-readable media (not shown) such as a hard disk or a Compact Disc Read-Only Memory (CD-ROM) drive.

[0317] Without loss of generality, the computer-readable medium can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, Erasable Programmable Read Only Memory (EPROM), Electrically-Erasable Programmable Read-Only Memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, Digital Versatile Disc (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art know that the computer storage media is not limited to the above several types. The above system memory 1204 and mass storage device 1206 can be collectively referred to as memory.

[0318] According to various embodiments of the present disclosure, the computer device 1200 may also operate by connecting to a remote computer on a network through a network such as the Internet. That is, the computer device 1200 may be connected to the network 1208 through the network interface unit 1207 connected to the system bus 1205. Or rather, the network interface unit 1207 may also be used to connect to other types of networks or remote computer systems (not shown).

[0319] The memory further includes at least one segment of computer program. The at least one segment of computer program is stored in the memory, and the central processing unit 1201 implements all or part of the steps in the speech synthesis method or the training method of the speech synthesis model shown in the above various embodiments by executing the at least one segment of program.

[0320] An embodiment of the present application further provides a computer device, which includes a processor and a memory. At least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the speech synthesis method or the training method of the speech synthesis model provided in the above method embodiments.

[0321] An embodiment of the present application further provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by the processor to implement the speech synthesis method or the training method of the speech synthesis model provided in the above method embodiments.

[0322] An embodiment of the present application further provides a computer program product, which includes a computer program stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device, so that the computer device executes to implement the speech synthesis method or the training method of the speech synthesis model provided in the above method embodiments.

[0323] It can be understood that in the specific implementation manner of the present application, for data, historical data, portraits, and other user data processing related to user identity or characteristics, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0324] It should be noted that, unless otherwise clearly defined herein, all terms used in the claims are to be interpreted according to their ordinary meanings in the technical field. Unless otherwise expressly stated, all references to "an element, device, component, equipment, step, etc." are to be construed liberally as referring to at least one instance of the element, device, component, equipment, step, etc. Unless expressly stated, the steps of any method disclosed herein need not be performed in the exact order disclosed.

[0325] It should be understood that "a plurality" as referred to herein means two or more. "And / or" describes the associative relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0326] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The storage media mentioned above can be a read-only memory, a magnetic disk, an optical disk, etc.

[0327] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A speech synthesis method, It is characterized in that The method comprises: Acquire semantic features and acoustic features, wherein the semantic features are used to represent features of text information corresponding to the audio to be synthesized, and the acoustic features are features of acoustic information corresponding to a reference voice, wherein the reference voice refers to the voice of an object corresponding to the selected timbre; The semantic feature is embedded into the acoustic feature through a first layer sub-model in a speech synthesis model to obtain an intermediate acoustic feature, wherein the first layer sub-model is used to embed the semantic feature into the acoustic feature, and the intermediate acoustic feature is used to represent the feature obtained after the semantic feature is embedded into the acoustic feature; Inputting the intermediate acoustic features into a second layer sub-model in the speech synthesis model to obtain synthesized audio features, wherein the second layer sub-model is used to synthesize the synthesized audio features based on the intermediate acoustic features, and the synthesized audio features are used to represent features corresponding to the audio to be synthesized; A synthesized audio having the same timbre as the reference voice is generated based on the synthesized audio features.

2. The method according to claim 1, It is characterized in that The method of embedding the semantic feature into the acoustic feature through the first layer sub-model in the speech synthesis model to obtain the intermediate acoustic feature includes: Adding the feature values ​​of the same column dimension in the acoustic feature to obtain a one-dimensional acoustic feature; splicing the semantic feature and the one-dimensional acoustic feature to obtain a spliced ​​feature, where the spliced ​​feature refers to a feature obtained by splicing the semantic feature and the one-dimensional acoustic feature; The concatenated features are input into the first layer sub-model to obtain the intermediate acoustic features.

3. The method according to claim 2, It is characterized in that The step of inputting the concatenated features into the first layer sub-model to obtain the intermediate acoustic features comprises: Inputting the concatenated feature into the first layer sub-model to predict the intermediate acoustic feature value in the intermediate acoustic feature, and obtaining the first intermediate acoustic feature value in the intermediate acoustic feature, wherein the intermediate acoustic feature refers to the feature predicted by the first layer sub-model based on the concatenated feature, and the intermediate acoustic feature value refers to the feature value in the intermediate acoustic feature; Inputting the i-1th synthetic audio feature value corresponding to the i-1th intermediate acoustic feature value and the concatenated feature into the first layer sub-model to predict the intermediate acoustic feature value, thereby obtaining the i-th intermediate acoustic feature value, wherein the i-1th synthetic audio feature value is obtained by inputting the i-1th intermediate acoustic feature value into the second layer sub-model for prediction; Repeat the previous step until the number of the intermediate acoustic eigenvalues ​​outputted is equal to the number of eigenvalues ​​in the one-dimensional acoustic feature; The i-th intermediate acoustic eigenvalue is combined with the first i-1 intermediate acoustic eigenvalues ​​to obtain the intermediate acoustic eigenvalue, where i is a positive integer greater than 1.

4. The method according to claim 1, It is characterized in that The step of inputting the intermediate acoustic features into the second layer sub-model in the speech synthesis model to obtain synthesized audio features includes: The intermediate acoustic feature values ​​in the intermediate acoustic feature are sequentially input into the second layer sub-model to obtain the synthesized audio feature.

5. The method according to claim 4, It is characterized in that The step of sequentially inputting the intermediate acoustic feature values ​​in the intermediate acoustic feature into the second layer sub-model to obtain the synthesized audio feature comprises: Inputting a first intermediate acoustic feature value in the intermediate acoustic features into the second layer sub-model to predict a synthetic audio feature value in the synthetic audio features, thereby obtaining a first synthetic audio feature value in the synthetic audio features, wherein the synthetic audio feature value refers to a feature value in the synthetic audio features; Inputting the jth intermediate acoustic eigenvalue generated by the j-1th synthetic audio eigenvalue into the second layer sub-model to predict the synthetic audio eigenvalue, thereby obtaining the jth synthetic audio eigenvalue, wherein the jth intermediate acoustic eigenvalue is obtained by inputting the j-1th synthetic audio eigenvalue into the first layer sub-model for prediction; Repeat the previous step until the number of the outputted synthetic audio feature values ​​is equal to the number of the intermediate acoustic feature values ​​in the intermediate acoustic feature; The j-th synthetic audio feature value is combined with the first j-1 synthetic audio feature values ​​to obtain the synthetic audio feature, where j is a positive integer greater than 1.

6. The method according to any one of claims 1 to 5, It is characterized in that The obtaining of acoustic features corresponding to the reference speech includes: Acquiring the reference speech; Extract features from the reference speech to obtain the acoustic features.

7. The method according to claim 6, It is characterized in that The extracting features from the reference speech to obtain the acoustic features comprises: Inputting the sample reference speech into the acoustic feature extraction network to extract features and obtain initial acoustic features, wherein the initial acoustic features refer to continuous features extracted from the reference speech; The initial acoustic feature is quantified to obtain the acoustic feature.

8. The method according to claim 7, It is characterized in that The quantifying the initial acoustic feature to obtain the acoustic feature includes: quantizing the mth acoustic feature value in the initial acoustic feature to obtain the first quantized acoustic feature value of the mth column dimension in the acoustic feature, where m is a positive integer; quantizing a residual between the first quantized acoustic eigenvalue and the m-th acoustic eigenvalue to obtain a second quantized acoustic eigenvalue in the m-th column dimension; quantizing a residual between the kth quantized acoustic eigenvalue and the k-1th quantized acoustic eigenvalue to obtain the k+1th quantized acoustic eigenvalue in the mth column dimension, where k is a positive integer greater than 1; Merging the first k+1 quantized acoustic feature values ​​in the m-th column dimension to obtain the quantized acoustic feature value of the m-th column dimension in the acoustic feature; Repeat the above steps to combine the quantized acoustic feature values ​​of each column dimension to obtain the acoustic feature.

9. A training method for a speech synthesis model, It is characterized in that The method comprises: Acquire sample semantic features, sample acoustic features and sample audio, wherein the sample semantic features are used to represent features of text information corresponding to the audio to be synthesized, and the sample acoustic features are features of acoustic information corresponding to the sample reference speech, and the sample reference speech refers to the speech of an object corresponding to the selected timbre; The sample semantic feature is embedded into the sample acoustic feature through the first layer sub-model in the speech synthesis model to obtain a sample intermediate acoustic feature, wherein the first layer sub-model is used to embed the sample semantic feature into the sample acoustic feature, and the sample intermediate acoustic feature is used to represent the feature obtained after the sample semantic feature is embedded into the sample acoustic feature; Inputting the intermediate acoustic features of the sample into the second layer sub-model in the speech synthesis model to obtain synthesized audio features, wherein the second layer sub-model is used to synthesize the synthesized audio features based on the intermediate acoustic features of the sample, and the synthesized audio features are used to represent features corresponding to the audio to be synthesized; Generate a synthesized audio having the same timbre as the sample reference speech based on the synthesized audio feature; Calculating the training loss of the speech synthesis model based on the sample audio and the synthesized audio; The model parameters of the speech synthesis model are updated according to the training loss.

10. The method according to claim 9, It is characterized in that The embedding of the sample semantic features into the sample acoustic features through the first layer sub-model in the speech synthesis model to obtain the sample intermediate acoustic features includes: Adding the feature values ​​of the same column dimension in the sample acoustic feature to obtain a one-dimensional sample acoustic feature; Splicing the sample semantic feature and the one-dimensional sample acoustic feature to obtain a sample splicing feature, where the sample splicing feature refers to a feature obtained by splicing the sample semantic feature and the one-dimensional sample acoustic feature; The sample concatenation features are input into the first layer sub-model to obtain the sample intermediate acoustic features.

11. The method according to claim 10, It is characterized in that The step of inputting the sample concatenation feature into the first layer sub-model to obtain the sample intermediate acoustic feature comprises: Inputting the sample concatenation feature into the sample intermediate acoustic feature value in the sample intermediate acoustic feature predicted by the first layer sub-model, to obtain the first sample intermediate acoustic feature value in the sample intermediate acoustic feature, wherein the intermediate acoustic feature refers to the feature predicted by the first layer sub-model based on the concatenation feature, and the sample intermediate acoustic feature value refers to the feature value in the sample intermediate acoustic feature; Inputting the i-1th synthetic audio feature value corresponding to the i-1th sample intermediate acoustic feature value and the sample concatenation feature into the first layer sub-model to predict the sample intermediate acoustic feature value, thereby obtaining the i-th sample intermediate acoustic feature value, wherein the i-1th synthetic audio feature value is obtained by inputting the i-1th sample intermediate acoustic feature value into the second layer sub-model for prediction; Repeat the previous step until the number of the outputted intermediate acoustic feature values ​​of the sample is equal to the number of feature values ​​in the one-dimensional sample acoustic feature; The intermediate acoustic feature value of the i-th sample is combined with the intermediate acoustic feature values ​​of the first i-1 samples to obtain the intermediate acoustic feature of the sample, where i is a positive integer greater than 1.

12. The method according to claim 9, It is characterized in that The step of inputting the sample intermediate acoustic features into the second layer sub-model in the speech synthesis model to obtain synthesized audio features includes: The sample intermediate acoustic feature values ​​in the sample intermediate acoustic feature are sequentially input into the second layer sub-model to obtain the synthesized audio feature.

13. The method according to claim 12, It is characterized in that The step of sequentially inputting the sample intermediate acoustic feature values ​​in the sample intermediate acoustic feature into the second layer sub-model to obtain the synthesized audio feature comprises: Inputting a first sample intermediate acoustic feature value in the sample intermediate acoustic features into the second layer sub-model to predict a synthetic audio feature value in the synthetic audio features, thereby obtaining a first synthetic audio feature value in the synthetic audio features; Inputting the j-th sample intermediate acoustic feature value generated by the j-1-th synthetic audio feature value into the second layer sub-model to predict the synthetic audio feature value, thereby obtaining the j-th synthetic audio feature value in the synthetic audio feature, wherein the j-th sample intermediate acoustic feature value is obtained by inputting the j-1-th sample synthetic audio feature value into the first layer sub-model for prediction; Repeat the previous step until the number of the outputted synthetic audio feature values ​​is equal to the number of sample intermediate acoustic feature values ​​in the sample intermediate acoustic feature; The j-th synthetic audio feature value is combined with the first j-1 synthetic audio feature values ​​to obtain the synthetic audio feature, where j is a positive integer greater than 1.

14. The method according to any one of claims 9 to 13, It is characterized in that The step of obtaining the sample acoustic features corresponding to the sample reference speech includes: Acquiring the sample reference speech; Extract features from the sample reference speech to obtain the sample acoustic features.

15. The method according to claim 14, It is characterized in that The step of extracting features from the sample reference speech to obtain the sample acoustic features includes: Inputting the sample reference speech into the acoustic feature extraction network to extract features, thereby obtaining initial sample acoustic features, wherein the initial sample acoustic features refer to continuous features extracted from the sample reference speech; The initial sample acoustic feature is quantified to obtain the sample acoustic feature.

16. A speech synthesis device, It is characterized in that The device comprises: An acquisition module, used to acquire semantic features and acoustic features, wherein the semantic features are used to represent features of text information corresponding to the audio to be synthesized, and the acoustic features are features of acoustic information corresponding to a reference voice, wherein the reference voice refers to the voice of an object corresponding to the selected timbre; A feature processing module, used for embedding the semantic feature into the acoustic feature through a first layer sub-model in a speech synthesis model to obtain an intermediate acoustic feature, wherein the first layer sub-model is used for embedding the semantic feature into the acoustic feature, and the intermediate acoustic feature is used for representing a feature obtained after embedding the semantic feature into the acoustic feature; The feature processing module is used to input the intermediate acoustic features into the second layer sub-model in the speech synthesis model to obtain synthesized audio features, the second layer sub-model is used to synthesize the synthesized audio features based on the intermediate acoustic features, and the synthesized audio features are used to represent the features corresponding to the audio to be synthesized; A generating module is used to generate a synthesized audio having the same timbre as the reference speech based on the synthesized audio feature.

17. A training device for a speech synthesis model, It is characterized in that The device comprises: An acquisition module, used to acquire sample semantic features, sample acoustic features and sample audio, wherein the sample semantic features are used to represent features of semantic text information corresponding to the audio to be synthesized, and the sample acoustic features are features of acoustic information corresponding to the sample reference speech, and the sample reference speech refers to the speech of an object corresponding to the selected timbre; A feature processing module, used for embedding the sample semantic feature into the sample acoustic feature through the first layer sub-model in the speech synthesis model to obtain a sample intermediate acoustic feature, wherein the first layer sub-model is used for embedding the sample semantic feature into the sample acoustic feature, and the sample intermediate acoustic feature is used for representing a feature obtained after embedding the sample semantic feature into the sample acoustic feature; The feature processing module is used to input the intermediate acoustic features of the sample into the second layer sub-model in the speech synthesis model to obtain synthesized audio features, the second layer sub-model is used to synthesize the synthesized audio features based on the intermediate acoustic features of the sample, and the synthesized audio features are used to represent the features corresponding to the audio to be synthesized; A generating module, configured to generate a synthesized audio having the same timbre as the sample reference speech based on the synthesized audio feature; A calculation module, used for calculating the training loss of the speech synthesis model based on the sample audio and the synthesized audio; An updating module is used to update the model parameters of the speech synthesis model according to the training loss.

18. A computer device, It is characterized in that The computer device includes: a processor and a memory, wherein at least one computer program is stored in the memory, and at least one computer program is loaded and executed by the processor to implement the speech synthesis method as described in any one of claims 1 to 8, or the training method of the speech synthesis model as described in any one of claims 9 to 15.

19. A computer storage medium, It is characterized in that The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the speech synthesis method as described in any one of claims 1 to 8, or the training method of the speech synthesis model as described in any one of claims 9 to 15.

20. A computer program product, It is characterized in that The computer program product includes a computer program, which is stored in a computer-readable storage medium; the computer program is read and executed from the computer-readable storage medium by a processor of a computer device, so that the computer device executes the speech synthesis method as described in any one of claims 1 to 8, or the training method of the speech synthesis model as described in any one of claims 9 to 15.

Citation Information

Cited By

  • Speech synthesis method, apparatus, device, storage medium, and program product

    EP4773133A1