Speech synthesis method and device, electronic equipment, computer readable storage medium and computer program product

By extracting and adjusting text features, calculating similarity and selecting matching voices in the voice cluster, the problem of style mismatch in the speech synthesis model is solved, and the style consistency and applicability of speech synthesis are improved.

CN120690170APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376873.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the existing technology, when focusing on timbre, the speech synthesis model ignores whether the style of the reference speech matches the style of the text content to be synthesized, resulting in the problem that the synthesized speech and text style are inconsistent or violate regulations.

Method used

By extracting features from the first text and the second text that describes the style of the speech cluster, the attention mechanism is used to adjust the features, the text similarity is calculated, and the speech in the speech cluster that matches the highest similarity is selected for synthesis to ensure that the synthesized speech is consistent with the text style.

Benefits of technology

It achieves style matching in the speech synthesis process, avoids inconsistencies and violations, and improves the naturalness and applicability of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690170A_ABST
    Figure CN120690170A_ABST
Patent Text Reader

Abstract

The invention provides a speech synthesis method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the steps that feature extraction is carried out on a first text to obtain first text features, feature extraction is carried out on a second text to obtain second text features, and the second text is used for describing the voice style of a first voice cluster; performing attention adjustment on the first text feature based on the second text feature to obtain a third text feature, and performing attention adjustment on the second text feature based on the first text feature to obtain a fourth text feature; determining a first similarity between the first text and the second text based on the third text feature and the fourth text feature; and determining a first voice matched with the second text with the highest first similarity in the first voice cluster, and synthesizing a second voice of the first text based on the first voice. Through the method and the device, the voice synthesized for the first text can better conform to the style represented by the first text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to speech synthesis technology, and in particular to a speech synthesis method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] Text-to-speech (TTS) is a technology that converts written text into human-readable speech. Using computer algorithms and language processing techniques, it transforms textual input into natural, fluent speech. TTS technology is widely used in voice assistants, accessibility services, education, entertainment, navigation systems, and other fields. Summary of the Invention

[0003] The embodiments of the present application provide a speech synthesis method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can make the speech synthesized for a first text more consistent with the style represented by the first text.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] The present invention provides a method for speech synthesis, comprising:

[0006] Perform feature extraction on a first text to obtain a first text feature, and perform feature extraction on a second text to obtain a second text feature, wherein the second text is used to describe the voice style of a first voice cluster; perform attention adjustment on the first text feature based on the second text feature to obtain a third text feature, and perform attention adjustment on the second text feature based on the first text feature to obtain a fourth text feature; determine a first similarity between the first text and the second text based on the third text feature and the fourth text feature; determine a first voice in the first voice cluster that matches the second text with the highest first similarity, and synthesize the second voice of the first text based on the first voice.

[0007] The present invention provides a speech synthesis device, comprising:

[0008] a feature extraction module, configured to extract features from the first text to obtain first text features, and to extract features from the second text to obtain second text features, wherein the second text is used to describe the speech style of the first speech cluster;

[0009] a feature adjustment module, configured to perform attention adjustment on the first text feature based on the second text feature to obtain a third text feature, and to perform attention adjustment on the second text feature based on the first text feature to obtain a fourth text feature;

[0010] a similarity determination module, configured to determine a first similarity between the first text and the second text based on the third text feature and the fourth text feature;

[0011] The speech synthesis module is configured to determine a first speech that matches the second text with the highest first similarity in the first speech cluster, and synthesize a second speech of the first text based on the first speech.

[0012] An embodiment of the present application provides an electronic device, comprising:

[0013] a memory for storing computer-executable instructions or computer programs;

[0014] The processor is used to implement the speech synthesis method provided in the embodiment of the present application when executing the computer-executable instructions or computer program stored in the memory.

[0015] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the speech synthesis method provided in the embodiment of the present application when executed by a processor.

[0016] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the speech synthesis method provided in the embodiment of the present application is implemented.

[0017] The embodiments of the present application have the following beneficial effects:

[0018] In the above manner, when performing speech synthesis on a first text, first text features of the first text are extracted, and second text features of a second text used to describe the speech style of the first speech cluster are extracted. Based on the attention mechanism, the attention of the first text features is adjusted using the second text features to obtain third text features, and the attention of the second text features is adjusted using the first text features to obtain fourth text features. In this way, the similarity in style between the first text and the second text can be fully learned through two cross-attention learnings. Then, the first similarity in style between the first text and the second text is calculated based on the third text features and the fourth text features, so that the first speech that matches the second text with the highest first similarity is determined in the first speech cluster, and the second speech of the first text is synthesized based on the first speech. In this way, since the speech style described by the second text with the highest first similarity is most similar to the style expressed by the text content of the first text, and the first speech matches the second text with the highest first similarity, the first speech determined in this manner matches the first text in style. With the first speech as a reference, the second speech synthesized for the first text is more consistent with the style expressed by the first text. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a schematic diagram of the architecture of the speech synthesis system provided in an embodiment of the present application;

[0020] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0021] Figure 3A This is a schematic diagram of the first flow chart of the speech synthesis method provided in an embodiment of the present application;

[0022] Figure 3B is a flow chart of a method for generating a second text provided in an embodiment of the present application;

[0023] Figure 3C is a flowchart of a method for determining a third text feature provided in an embodiment of the present application;

[0024] Figure 3D is a flowchart of a method for determining a fourth text feature provided in an embodiment of the present application;

[0025] Figure 4A This is a first schematic diagram of the style extraction model provided in an embodiment of the present application;

[0026] Figure 4B This is a second schematic diagram of the style extraction model provided in an embodiment of the present application;

[0027] Figure 5is a schematic diagram of the structure of the style matching model provided in an embodiment of the present application;

[0028] Figure 6 It is a schematic diagram of speech synthesis by the large speech synthesis model provided in an embodiment of the present application.

[0029] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0031] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0032] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0033] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0034] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0035] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0036] Text-to-speech synthesis is a technology that converts written text into human-readable speech. Using computer algorithms and language processing techniques, it transforms textual input into natural, fluent speech. Text-to-speech synthesis is widely used in voice assistants, accessibility services, education, entertainment, navigation systems, and other fields.

[0037] The Speech Generation Large Model (SGLM) is the core of text-to-speech synthesis technology. It converts input text into speech signals. Training a SGLM typically requires inputting a text sample to be synthesized and a reference speech. The reference speech is fed into the SGLM using a prompt (also known as a reference or borrowing). Based on the timbre of the reference speech, the SGLM converts the input text sample into speech of the corresponding timbre. The speech content is the same as the text sample.

[0038] Here, the template (prompt) can be a reference, meaning it provides additional conditional variables for the speech synthesis model. This means it provides the model with a case study, letting it know the desired content, style, and timbre of the speech it wants to synthesize. In text-to-speech synthesis, prompts typically take the form of paired reference examples, such as a <text, speech> pair. This tells the speech synthesis model that when synthesizing speech corresponding to the text, the user expects the model to synthesize speech of the same type as the reference example. Prompts serve as a reference or model, containing at least one reference speech, which serves as a target for the speech synthesis model. Based on the reference speech, the model analyzes timbre (i.e., the speaker's intonation, enabling speaker identification based on timbre), prosody (i.e., acoustic characteristics such as speaking rate, pauses, pitch, and volume), emotional information (emotional information such as joy, anger, sadness, and happiness), and personalized information (e.g., leaks, emphasis, catchphrases, and the use of markers).

[0039] When synthesizing speech from text, a reference speech is usually required. However, in related technologies, when providing a reference speech, people often only focus on the timbre of the speech and do not pay attention to whether the style of the reference speech matches the text to be synthesized. This approach may lead to problems such as the style of the synthesized speech being inconsistent with the text.

[0040] In an embodiment of the present application, the information expressed by a voice can be divided into two types of information, namely, timbre information and style information. The style information may include all information except timbre information. For example, the style information may include the above-mentioned rhythmic information, emotional information, personalized information, etc. The style information can reflect the comprehensive auditory psychological feeling given to the listener.

[0041] Large speech synthesis models can be trained using semi-supervised or unsupervised deep learning technologies. They can synthesize various styles of speech that are close to human-like. From the perspective of reasoning, large speech synthesis models include autoregressive (AR) models and non-autoregressive (NAR) models. Both types of large speech synthesis models have good speech synthesis naturalness, anthropomorphism, and generalization capabilities.

[0042] The reason why the large speech synthesis model has a high degree of anthropomorphic ability is largely because it is learned with the help of reference speech provided by the user during the inference stage. The large speech synthesis model can generate a timbre and style that is close to the reference speech.

[0043] However, in related technologies, when providing reference speech for large speech synthesis models, people often only focus on timbre. For example, if a user wants to synthesize speech with the voice of a certain celebrity, the user will usually provide a recording sample of the celebrity as a reference speech, thereby synthesizing speech in a specified text that matches the celebrity's voice.

[0044] However, focusing only on timbre without considering whether the style of the reference speech matches or matches the style of the text to be synthesized often leads to the following problems: First, text-sound disharmony. The ultimate goal of a large-scale speech synthesis model is to synthesize speech for a specified text. If the style of the provided reference speech does not match the desired emotional expression of the text to be synthesized, this can create a dissonant auditory experience. For example, if a sentence "Patience is the key" is synthesized in an intense tone based on the reference speech, the resulting synthesized speech for the specified text will be very discordant. Second, style violations occur. In certain application domains, the rhythm and emotion of certain speech styles are not permitted (or do not conform to the application domain's regulations). For example, in the field of AI customer service, speech with emotions such as anger, pain, and sadness, or with rhythms or characteristics such as laziness, intensity, and ambiguity, should not be synthesized and played to customers.

[0045] Therefore, embodiments of the present application provide a speech synthesis method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can make the speech synthesized for a first text more consistent with the style represented by the first text.

[0046] It should be noted that the speech synthesis method provided in the embodiment of the present application is a text-to-speech synthesis technology that can be used in various application scenarios such as assisted reading, intelligent voice interaction, dubbing, customer service and marketing, public information broadcasting, personalized voice assistants, etc.

[0047] See also Figure 1 , Figure 1 This is an architectural diagram of the speech synthesis system 100 provided in an embodiment of the present application. To support a speech synthesis application, terminals (terminal 400-1 and terminal 400-2 are shown as examples) are connected to the server 200 via a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0048] Here, we will use customer service and marketing scenarios as examples to illustrate.

[0049] In some embodiments, the terminal (e.g., terminal 400-1, terminal 400-2) and server 200 may jointly serve as the execution body to execute the speech synthesis method of the embodiment of the present application. Specifically, the user may ask questions to the intelligent customer service through the terminal (e.g., terminal 400-1, terminal 400-2) by phone or online customer service; the terminal (e.g., terminal 400-1, terminal 400-2) responds to the user's question and generates a response text for the question content based on the large language model, and uses the response text as the first text, and generates a speech synthesis request for the first text, and sends the speech synthesis request to the server 200. The server 200 responds to the speech synthesis request. According to a request, feature extraction is performed on the first text to obtain the first text feature, and feature extraction is performed on the second text to obtain the second text feature, where the second text is used to describe the voice style of the first voice cluster; attention adjustment is performed on the first text feature based on the second text feature to obtain the third text feature, and attention adjustment is performed on the second text feature based on the first text feature to obtain the fourth text feature; based on the third text feature and the fourth text feature, a first similarity between the first text and the second text is determined; a first voice that matches the second text with the highest first similarity in the first voice cluster is determined, and a second voice of the first text is synthesized based on the first voice using a pre-trained voice synthesis model. Subsequently, the server 200 sends the second voice to the terminal (e.g., terminal 400-1, terminal 400-2), so that the terminal (e.g., terminal 400-1, terminal 400-2) outputs the second voice to the user through the intelligent customer service.

[0050] In some embodiments, the speech synthesis method provided in the embodiments of the present application can be independently executed by a terminal (e.g., terminal 400-1, terminal 400-2). Specifically, a user can ask questions to the intelligent customer service through a telephone or online customer service at the terminal (e.g., terminal 400-1, terminal 400-2); the terminal (e.g., terminal 400-1, terminal 400-2) generates a response text for the question content based on a large language model in response to the user's question, and uses the response text as the first text. Then, the terminal (e.g., terminal 400-1, terminal 400-2) performs feature extraction on the first text to obtain the first text feature, and performs feature extraction on the second text to obtain the second text feature, where the second text is used to describe the voice style of the first speech cluster; based on the second text feature, attention is adjusted on the first text feature to obtain the third text feature, and based on the first text feature, attention is adjusted on the second text feature to obtain the fourth text feature; based on the third text feature and the fourth text feature, the first similarity between the first text and the second text is determined; the first speech that matches the second text with the highest first similarity in the first speech cluster is determined, and the second speech of the first text is synthesized based on the first speech using a pre-trained speech synthesis model. The terminal (e.g., terminal 400-1, terminal 400-2) outputs the second speech to the user through intelligent customer service.

[0051] It can be understood that the first speech in the above embodiment is the reference speech used as a reference for the speech synthesis model.

[0052] In some embodiments, the terminal (e.g., terminal 400-1, terminal 400-2) can be implemented as various types of terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, car terminals, etc., and can also be implemented as servers.

[0053] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.

[0054] See also Figure 2 , Figure 2 is a structural diagram of an electronic device 400 provided in an embodiment of the present application, Figure 2The electronic device 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .

[0055] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0056] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0057] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0058] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0059] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0060] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0061] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);

[0062] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0063] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.

[0064] In some embodiments, the speech synthesis device provided in the embodiments of the present application can be implemented in software. Figure 2 A speech synthesis device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a feature extraction module 4551, a feature adjustment module 4552, a similarity determination module 4553, and a speech synthesis module 4554. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0065] In other embodiments, the speech synthesis device provided in the embodiments of the present application can be implemented in hardware. As an example, the speech synthesis device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the speech synthesis method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.

[0066] In some embodiments, the terminal or server can implement the speech synthesis method provided in the embodiment of the present application by running various computer executable instructions or computer programs. For example, the computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP; it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0067] The following describes the speech synthesis method provided by the embodiments of the present application in conjunction with the accompanying drawings. As previously mentioned, the electronic device 400 that implements the speech synthesis method of the embodiments of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.

[0068] The speech synthesis method of the embodiment of the present application is described by taking the execution subject as a terminal as an example. Figure 3A , Figure 3A This is a first flow chart of the speech synthesis method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.

[0069] In step 101 , feature extraction is performed on a first text to obtain first text features, and feature extraction is performed on a second text to obtain second text features.

[0070] Here, the first text is the text used for synthesizing speech. In different application scenarios, the method of obtaining and determining the first text can be different. For example, in an assisted reading scenario, the first text can be at least part of the text in a book, article, webpage content, email, etc., such as the first text can be a paragraph or chapter in an article; in the field of dubbing, the first text can be the lines of a virtual character or actor; in the field of customer service and marketing, the first text can be the reply text of the intelligent customer service to the customer's question. As an example, in the field of customer service and marketing, the first text can be generated by obtaining the customer's question information, inputting the question information into a large language model, generating a reply text for the question information through the large language model, and using the reply text generated by the large language model as the first text for speech synthesis.

[0071] In actual implementation, the second text is used to describe the speech style of the first speech cluster. The second text can be used to refer to speech belonging to a certain speech style. Different second texts can describe different speech styles. The speech styles described by the second texts may include information such as prosody, emotion, and personalized features. A second text is used to describe the speech style of a first speech cluster. Speech within the same first speech cluster is consistent in speech style and conforms to the speech style described by the corresponding second text.

[0072] In some embodiments, Figure 3B This is a flow chart of the method for generating the second text provided in the embodiment of the present application, see Figure 3B Before performing feature extraction on the second text, the second text may be generated through steps 105 to 108.

[0073] In step 105, a speech set is constructed.

[0074] The voice set includes a plurality of first voice groups corresponding to the timbres, and the first voice group includes a plurality of first voice samples matching the timbres.

[0075] In actual implementation, the speech set is used to select the first speech (i.e., the reference speech) for reference when performing speech synthesis on the first text. Therefore, to increase selectivity, a speech set including first speech samples of various timbres can be constructed. The construction of the speech set can be based on the needs of the specific scenario in which the embodiments of this application are applied.

[0076] Below, we will use the construction of a voice collection in the field of customer service and marketing as an example to illustrate. We will obtain the recording files generated by multiple customer service personnel when they communicated with customers by phone during a historical period, and extract the voice audio corresponding to each customer service personnel based on the recording files. Since different customer service personnel correspond to different timbres, a corresponding first voice group can be constructed for each customer service personnel. The first voice group includes the voice audio corresponding to the customer service personnel (i.e., the first voice sample). In this way, multiple first voice groups are obtained, and different first voice groups correspond to different customer service personnel, that is, to different timbres. For example, taking customer service personnel A as an example, the voice audio generated by customer service personnel A communicating with customers by phone is collected, and the voice audio of customer service personnel A is classified as the first voice sample of customer service personnel A into the first voice group. The first voice group corresponds to the timbre of customer service personnel A.

[0077] In actual implementation, in order to facilitate the management of voice sets, the first voice group can be stored in the format of (first voice sample, voice text). For example, the first voice group A can be recorded as S1[(SA1, TA1), (SA2, TA2)…(SAi, TAi)], where S1 represents the first voice group A, TA1 is the voice text corresponding to the first voice sample SA1, TA2 is the voice sample corresponding to the first voice sample SA1, and TAi is the voice sample corresponding to the first voice sample SAi.

[0078] In actual implementation, when constructing the first voice group corresponding to each timbre, the voice audio extracted from the recording file can be screened, and the first voice group can be constructed based on the screened voice audio. Among them, the screening method may include: identifying the voice text corresponding to each voice audio according to voice recognition technology, detecting whether each voice text contains samples of illegal words and sentences, and if so, removing the voice audio containing illegal words and sentences. For the voice audio after removal, retain the voice audio with an audio duration within the specified duration range, for example, retain the voice audio with an audio duration of 3 to 6 seconds. Based on the voice audio screened in the above manner, construct the first voice group of each timbre.

[0079] In step 106 , for the first speech group corresponding to each timbre, style features are extracted for each first speech sample in the first speech group to obtain style features of the first speech sample.

[0080] In actual implementation, when subsequently selecting the first speech for the first text from the speech collection, the focus is on selecting first speech with matching style. Therefore, style features can be extracted from each first speech sample to obtain style features of the first speech sample, thereby determining the speech style embodied by the first speech sample. The style features are used to characterize style-related attribute information such as the speech style or speech type of the first speech sample.

[0081] In some embodiments, step 106 can be implemented by: performing feature fusion on the pitch feature, volume feature and Mel spectrum feature of the first speech sample to obtain a first latent feature of the first speech sample; extracting the voiceprint feature of the first speech sample based on the Mel spectrum feature, and eliminating the voiceprint feature from the first latent feature to obtain a second latent feature; quantizing and encoding the second latent feature to obtain a third latent feature; and predicting the style feature of the first speech sample based on the third latent feature to obtain the style feature of the first speech sample.

[0082] In actual implementation, a style extraction model for extracting style features may be pre-trained, and the style features of each first speech sample may be extracted through the style extraction model.

[0083] As an example, Figure 4AThis is the first schematic diagram of the style extraction model provided in the embodiment of the present application, see Figure 4A The style extraction model includes a feature processing module 110, an encoder 120, a quantization module 130, a decoder 140, and a voiceprint extraction module 150. When the style extraction model is used to extract style features from the first speech sample, the input of the style extraction model includes three basic features of the first speech sample, namely, pitch features, volume features, and mel-spectrogram features. These three basic features are manifested in feature sequences, namely, pitch feature sequences, volume feature sequences, and mel-spectrogram feature sequences, and all three sequences are frame-level feature sequences. In other words, the resolution of the three basic features is consistent, namely, the frame length is consistent, so that the lengths of the pitch feature sequences, volume feature sequences, and mel-spectrogram feature sequences are consistent.

[0084] Among them, the feature processing module 110 is used to normalize the three basic features input, and the voiceprint extraction module 150 is used to extract the voiceprint features of the first speech sample based on the mel-spectrogram features of the first speech sample. Then, the three normalized basic features and the extracted voiceprint features are input into the encoder 120. In the encoder 120, the three normalized basic features are fused, and the voiceprint features are extracted from the first latent features obtained after fusion to obtain the second latent features. Then, the second latent features are quantized and encoded by the quantization module 130. The purpose of quantization encoding is to extract the voice style representation of the first speech sample. The third latent features obtained after quantization encoding are the voice style representation extracted after the voiceprint features are extracted. Finally, the third latent features are decoded and calculated (i.e., predicted) by the decoder 140 to obtain the style features of the first speech sample. When the decoder 140 performs decoding calculations, the extracted voiceprint features can also be added to the third latent features so that the style features decoded by the decoder 140 can also reflect the timbre corresponding to the first speech sample.

[0085] It should be noted that when the decoder 140 of the style extraction model decodes the third latent feature, its goal is to output the Mel-spectrogram feature of the specified dimension. When training the style extraction model, the loss calculation can be performed based on the difference between the output Mel-spectrogram feature of the specified dimension and the actual Mel-spectrogram feature of the first speech sample, so as to adjust the parameters of the style extraction model based on the loss, so that the style extraction model is continuously optimized until the style extraction model converges.

[0086] Figure 4B is a second schematic diagram of the style extraction model provided in an embodiment of the present application, Figure 4B The model structure shown is Figure 4A The detailed structure of each module of the style extraction model shown below is combined with Figure 4B , a detailed description is given of the specific process of extracting the style of the first speech sample through the style extraction model.

[0087] See also Figure 4B The feature processing module 110 in the figure includes a pitch normalization layer, a volume normalization layer, a stacking layer, a linear layer, an activation layer, and a generalization layer. The pitch normalization layer is used to normalize the pitch features, and the volume normalization layer is used to normalize the volume features. The default representation values ​​of the pitch features and the volume features are both floating-point numbers greater than 1. The goal of the normalization layer is to normalize these two features to an approximate normal distribution with a mean of 0 and a deviation of 1, so as to facilitate the subsequent deep learning process of the style extraction model. The input Mel spectrum features are normalized Mel spectrums, so no Mel spectrum normalization layer is specifically configured here. If the Mel spectrum features are unnormalized Mel spectrum values, a normalization layer is also required. Then, these three input vectors (i.e., normalized pitch features, normalized volume features, and normalized Mel spectrum features) are stacked together to form an input feature vector with 22 channels (i.e., 22 dimensions). The stacking method here can be a stack method. This 22-dimensional input feature vector passes through a linear layer (assuming the number of input channels is 22 and the number of output channels is 256), and then undergoes nonlinear activation of the activation layer (also called the ReLU layer), and a generalization layer (also called the Dropout layer) to increase generalization, and finally obtains the first latent feature of the first speech sample. It should be noted that the number of input channels of the linear layer of the style extraction model in the embodiment of the present application is set to 256 in order to align the channel size to 256. This value is not a required value and can be set according to actual conditions. In addition, the parameters of the generalization layer are not required values ​​and can be set based on actual needs.

[0088] Continue to see Figure 4B , feature processing is continued through the encoder 120. The encoder 120 includes an autoencoding layer, a first enhancement layer, a downsampling layer, a second enhancement layer, and an autoencoding layer. Among them, the first enhancement layer includes a normalization layer, a regularization layer, and an activation layer, and the second enhancement layer also includes a normalization layer, a regularization layer, and an activation layer; the downsampling layer can use a Cony1D convolution layer; the autoencoding layer can include a Transform layer, or it can be composed of multiple Transform stacks, which are linked in a residual manner. The Transform layer is used to perform nonlinear transformation and enhancement on the features. After the first latent feature is processed by the autoencoding layer, the first enhancement layer, the downsampling layer, and the second enhancement layer, the voiceprint feature extracted by the voiceprint extraction module 150 is subtracted, and then it is processed by the normalization layer and the autoencoding layer to obtain the second latent feature of the first speech sample.

[0089] Here, in the encoder, the convolution stride is set to 2, so that each time a feature passes through a convolution layer, the length of the feature is compressed to one-half. When multiple downsampling layers are set, the length of the feature is compressed multiple times. Therefore, after the first latent feature passes through the downsampling layer, its length is compressed to half of its original length. In other words, after the first latent feature is downsampled, the feature size becomes a matrix of L / 2*256. The voiceprint feature is usually a single-length vector of 1*256. Before subtracting the first latent feature from the voiceprint feature, the voiceprint feature can be horizontally copied to align its length with the downsampled first latent feature.

[0090] It should be noted that the voiceprint extraction module 150 may directly adopt a pre-trained voiceprint extraction model, which may extract voiceprint features based on the Mel-spectrogram features of the input first speech sample.

[0091] In the above manner, since the voiceprint feature can directly reflect the timbre of the first speech sample, eliminating the voiceprint feature from the first latent feature can make the style extraction model unaffected by the timbre, thereby extracting a purer style feature.

[0092] Continue to see Figure 4B After the encoder 120 normalizes and encodes the first latent feature after removing the voiceprint feature, it obtains the second latent feature of the first speech sample. The second latent feature represents the voice style of the first speech sample. The second latent feature is then input into the quantization module 130, where it is quantized and encoded. The core parameter of the quantization module 130 is the size of the quantization code table, which can be set to 1024. This means that the input 256-dimensional second latent feature will be approximated to an element in a table consisting of 1024 latent features, and this element is also a 256-dimensional latent feature. Here, the goal of quantization encoding is to divide the high-dimensional continuous feature space into several subregions, each represented by a code vector. The input second latent feature is mapped to the nearest code vector, thereby achieving discretization. The quantization module can be used to reduce data volume or improve storage and transmission efficiency. The style extraction model has a code table loss term during training, which is specifically responsible for making the latent features in the code table more representative. After the second latent feature is quantized and encoded by the quantization module 130, a third latent feature of the same dimension is obtained.

[0093] Continue to see Figure 4BThe network structure of the decoder 140 is roughly opposite to the network result of the encoder 120, that is, the third latent feature of the input is first subjected to autoencoding learning of the autoencoding layer. Here, the voiceprint feature that is removed above can be added to the third latent feature to become a third latent feature with timbre information. The third latent feature with timbre information is then subjected to an upsampling layer for a processing process opposite to the downsampling layer in the encoder 120, so that the third latent feature is restored to its original size of L*256. After the size is restored, it is processed by an autoencoding layer, a second enhancement layer, and a linear layer to map the number of channels of the third latent feature from 256 to 20, that is, the value of the Mel spectrum on the 20-frequency band of the original input of the style extraction model is predicted to obtain the style feature of the first speech sample.

[0094] Here, we will further explain the training process of the style extraction model. The loss function used in the training of the style extraction model consists of two parts: one is the quantization loss of the quantization module, and the other is the prediction loss of the decoder module. The weight configuration of the two losses can be set to 1:8, that is, the weight of the quantization loss is 1, and the weight of the prediction loss is 8.

[0095] In the above manner, when the pre-trained style extraction model is used to extract the style features of each first voice sample, by eliminating the voiceprint features, a purer style feature of the first voice sample can be predicted, thereby reducing the influence of timbre.

[0096] In step 107 , based on the style features of the first speech samples, a plurality of first speech samples in the first speech group are clustered to obtain a first speech cluster.

[0097] The purpose of clustering here is to cluster the first speech samples of the same style in the first speech group corresponding to each timbre into the same first speech cluster, that is, to classify the first speech group according to speech style for subsequent use.

[0098] In some embodiments, step 107 can be implemented in the following manner: clustering the first speech samples in the speech set based on the semantic features of each first speech sample in the speech set to obtain a second speech cluster; determining the second similarity between the semantic features corresponding to the cluster center of the second speech cluster and the semantic features of each first speech sample in the first speech group, and screening out the second speech sample with the largest second similarity from the multiple first speech samples in the first speech group; using the screened out second speech sample as the cluster center, clustering the multiple first speech samples in the first speech group based on the style features of each first speech sample in the first speech group to obtain a first speech cluster.

[0099] In actual implementation, to improve the accuracy of clustering based on style features within the first speech group, an embodiment of the present application proposes a method for initializing the clustering algorithm by selecting cluster centers for style clustering based on maximizing semantic differences. Specifically, a first semantic clustering is first performed on all first speech samples in the speech set. The cluster centers used for clustering each first speech group are determined based on the second speech clusters obtained from the first clustering. In this way, when clustering the first speech samples in the first speech group, both semantic and style information can be taken into account, thereby improving the accuracy of clustering.

[0100] The following describes the first semantic clustering of all first speech samples in the speech set. As previously described, each first speech sample in the speech set has a corresponding speech text, which is the textual content represented by the first speech sample. Therefore, semantic features can be extracted from the speech text of each first speech sample to obtain the semantic features of the first speech sample. The semantic features of the first speech samples are then used to calculate the semantic similarity between any two first speech samples. Based on the semantic similarity, all first speech samples in the speech set are clustered to obtain multiple second speech clusters. Each second speech cluster includes multiple first speech samples whose semantic similarity exceeds a similarity threshold.

[0101] In some embodiments, the above-mentioned "clustering the first speech samples in the speech set based on the semantic features of each first speech sample in the speech set to obtain a second speech cluster" can be achieved by: clustering the first speech samples in the speech set based on the fourth similarity (i.e., semantic similarity) between the semantic features of each first speech sample in the speech set to obtain a third speech cluster, and the fourth similarity between the semantic features of the first speech samples in the third speech cluster exceeds the second similarity threshold; and screening out a second speech cluster from the third speech cluster in which the number of first speech samples exceeds the number threshold.

[0102] In actual implementation, after calculating the fourth similarity between the semantic features of each first speech sample, hierarchical clustering can be performed on the first speech samples in the speech set based on the fourth similarity. The hierarchical clustering algorithm uses each sample as the initial cluster center. The clustering process gradually merges samples with closer distances into the same cluster. Cluster centers are then calculated for the resulting clusters. Here, each first speech sample in the speech set can be used as the initial cluster center. Based on the fourth similarity between the first speech samples, hierarchical clustering is performed, so that first speech samples with semantic similarity exceeding the second similarity threshold are clustered into the same cluster, ultimately resulting in multiple third speech clusters. Since the number of first speech samples included in the third speech cluster is typically inconsistent, a quantity threshold can be set to improve accuracy. Second speech clusters whose number of first speech samples exceeds the quantity threshold can be screened out from the third speech cluster.

[0103] Here, it should be noted that hierarchical clustering is to form second speech clusters by constructing a tree structure (such as agglomerative clustering or divisive clustering). Therefore, for each second speech cluster, the mean or median of all first speech samples included in the second speech cluster can be calculated to determine the cluster center of the second speech cluster, or the first speech sample closest to the mean in the second speech cluster can be selected as the cluster center of the second speech cluster.

[0104] In some embodiments, the above-mentioned "clustering the first speech samples in the speech set based on the semantic features of each first speech sample in the speech set to obtain a second speech cluster" can also be achieved in the following way: selecting a third speech sample from each first speech sample in the speech set; using the third speech sample as the cluster center of the second speech cluster, and clustering the first speech samples in the speech set based on the third similarity between the semantic features of the third speech sample and the semantic features of the fourth speech sample in the speech set to obtain a second speech cluster; wherein the third similarity between the semantic features of the first speech samples in the second speech cluster is greater than the first similarity threshold.

[0105] Here, the fourth voice sample is the first voice sample in the voice set that does not include the third voice sample, that is, the fourth voice sample is the first voice sample in the voice set that is different from the third voice sample.

[0106] In actual implementation, in addition to the hierarchical clustering described in the above embodiment, clustering can also be performed by preselecting cluster centers. Specifically, a third speech sample can be randomly selected from the first speech samples in the speech collection. This third speech sample is used as the cluster center of a second speech cluster and clustered once to obtain a second speech cluster. Then, the remaining first speech samples in the speech collection are clustered multiple times using the same method until no second speech clusters exceeding a threshold number can be found, thereby obtaining multiple second speech clusters. In this method, the cluster center of each second speech cluster is determined before clustering, and the number of first speech samples included in the second speech cluster exceeds the threshold number.

[0107] When clustering is performed with the third speech sample as the cluster center, a third similarity (i.e., semantic similarity) between the semantic features of the third speech sample and the semantic features of each fourth speech sample may be calculated. Then, fourth speech samples having a third similarity greater than a first similarity threshold are clustered into the second speech cluster corresponding to the third speech sample as the cluster center. The fourth speech sample is the first speech sample in the speech set excluding the third speech sample.

[0108] Through the above method, two clustering methods are provided, and a suitable clustering method can be selected based on actual clustering requirements to cluster the first speech sample in the speech set.

[0109] It can be understood that the number of second speech clusters obtained by semantically clustering the first speech samples in the speech set can be determined based on a quantity threshold or can be manually specified. For example, N second speech clusters can be screened out from multiple third speech clusters according to the quantity from high to low. N can be set based on actual needs. For example, in a customer service scenario, N can be set to 20. This is because the voice style of customer service itself has a limited range of variation, and 20 styles are sufficient to meet the needs of style selection during subsequent speech synthesis. Other application scenarios need to be adjusted according to specific circumstances.

[0110] Next, step 107 will be explained. After semantically clustering the speech set to obtain multiple second speech clusters, the cluster centers of the multiple second speech clusters can be used as semantic centers. When clustering the various first speech groups subsequently, the cluster center for clustering the first speech groups can be selected based on the multiple semantic centers.

[0111] In actual implementation, for the first speech group corresponding to any timbre, the second similarity (i.e., semantic similarity) between the semantic features of each first speech sample in the first speech group and the semantic features of the cluster center of each second speech cluster is calculated, thereby screening out the second speech sample with the greatest second similarity to the cluster center of each second speech cluster in the first speech group.

[0112] As an example, assuming that the first speech group includes 100 first speech samples and the number of second speech clusters is N=20, it is necessary to determine in the first speech group the 20 second speech samples that are closest to the cluster centers of the 20 second speech clusters. The closest distance indicates the maximum second similarity.

[0113] Then, the screened second speech samples are used as cluster centers for clustering the first speech group. Based on the similarity between the stylistic features of the first speech samples in the first speech group, the multiple first speech samples in the first speech group are clustered to obtain multiple first speech clusters. The number of first speech clusters is the same as the number of second speech clusters. In this way, the first speech samples with the same style in the first speech group can be clustered into the same first speech cluster. In other words, among the multiple first speech clusters corresponding to the first speech group, different first speech clusters correspond to different speech styles, but all correspond to the same timbre, namely the timbre corresponding to the first speech group.

[0114] In step 108 , for each first speech cluster, a second text is generated for describing the speech style of the first speech cluster based on the style features of the first speech samples in the first speech cluster.

[0115] In actual implementation, after obtaining multiple first speech clusters of the first speech group through the above-mentioned clustering process, for each first speech cluster, a second text for describing the speech style of the first speech sample in the first speech cluster can be summarized and generated based on the style characteristics of the first speech sample in the first speech cluster. The second text can be used to refer to the speech style of the first speech sample in the first speech cluster.

[0116] It can be understood that, through the above method, multiple second texts corresponding to the first voice group with different timbres can be obtained, that is, each timbre in the voice set will correspond to multiple second texts.

[0117] In the above manner, the first speech samples of different voice styles under different timbres in the speech set are referred to by the second text in text form. When style matching is subsequently performed for the first text, the second text can be used for matching without directly using the first speech sample in the speech set for matching. In other words, there is no need for cross-modal processing between text and speech, thereby making the style matching for the first text more accurate.

[0118] In actual implementation, the process shown in steps 101 to 103 in the speech synthesis method provided in the embodiment of the present application can be implemented by a pre-trained style matching model. Figure 5 This is a schematic diagram of the structure of the style matching model provided in the embodiment of the present application, see Figure 5 The style matching model consists of two structures: a first text processing module (represented by the left structure) and a second text processing module (represented by the right structure). To comply with the style matching model, the first and second texts can be input into the style matching model in the form of word units (tokens).

[0119] In actual implementation, in order to be able to more closely match the style and emotions expressed by the text content of the first text when performing speech synthesis on the first text, the first encoding layer, when extracting features (i.e., encoding) the first text, focuses more on reflecting the style information of the first text for the first text features extracted from the first text, and position-encodes the tokens of the first text through the position encoding layer, and then adds the position encoding to the extracted first text features. Similarly, in order to enable the style matching model to better understand the style information described by the second text, the second encoding layer, when extracting features from the second text, will also focus more on extracting the style information of the second text, obtain the second text features, and use the position encoding layer to position-encode the tokens of the second text, and add the position encoding to the second text features.

[0120] It can be understood that the first encoding layer and the second encoding layer of the style matching model are intended to capture the language characteristics, sentiment tendencies, emotions and other style information of the respective input texts, and through feature extraction, the style expressed by the text can be quantified.

[0121] The first position-encoded text feature is then processed through a convolutional layer (i.e., a Transform layer) and a linear layer before attention adjustment. Similarly, the second position-encoded text feature is also processed through a convolutional layer (i.e., a Transform layer) and a linear layer before attention adjustment. Here, the convolutional layer is primarily used for multi-layer self-attention learning to obtain a deep encoding representation, while the linear layer is primarily used to perform linear transformations on the features to map them to a new feature space and extract higher-level features.

[0122] Continue to see Figure 3A , continue with the above step 101 for explanation.

[0123] In step 102 , attention adjustment is performed on the first text feature based on the second text feature to obtain a third text feature, and attention adjustment is performed on the second text feature based on the first text feature to obtain a fourth text feature.

[0124] Continue to see Figure 5 After the first text feature is processed by a linear layer and the second text feature is processed by a linear layer, the first text feature and the second text feature can be cross-attention aligned based on the attention mechanism.

[0125] The attention mechanism is a widely used technique in deep learning that allows models to selectively focus on important parts of information while ignoring less important information. This mechanism mimics the human tendency to focus attention on the most critical information while ignoring distracting information when processing a task. Its basic principle is to determine which parts are important by calculating the correlation between query features, key features, and value features.

[0126] The core idea of ​​the attention mechanism is to simulate how humans focus their attention. It calculates attention weights for each input position, enabling the model to dynamically assign weights based on different parts of the input data, allowing the model to dynamically focus on key parts of the input data. When processing data, the attention mechanism allows the model to consider information at every position in the feature sequence, rather than relying solely on input at fixed positions.

[0127] In an embodiment of the present application, two cross-attention learning processes are set up based on the cross-attention mechanism. In the left network branch corresponding to the first text, the first attention alignment learning is performed on the weighted representation of the second text based on the first text. In the right network branch corresponding to the second text, the second attention alignment learning is performed on the weighted representation of the first text based on the second text.

[0128] Below, we introduce two cross-attention learning processes in detail.

[0129] In some embodiments, Figure 3C This is a flow chart of the method for determining the third text feature provided in the embodiment of the present application, see Figure 3C , Figure 3A The step 102 of “adjusting the attention of the first text feature based on the second text feature to obtain the third text feature” can be implemented through steps 1021 to 1023 .

[0130] In step 1021, the first text feature is used as the query feature in the attention mechanism, and the second text feature is used as the key feature and value feature in the attention mechanism.

[0131] See also Figure 5 For the first attention alignment learning, the first text feature after being processed by the linear layer is used as the query feature in the attention mechanism, and the second text feature after being processed by the linear layer is used as the key feature and value feature in the attention mechanism. This can be used to guide the first text feature to extract relevant style information from the second text feature, so that the first text feature focuses on the content related to its own style in the second text feature, and realizes the dynamic interaction of the style information of the two.

[0132] In step 1022 , a first attention weight between the first text feature and the second text feature is determined based on a dot product operation of the query feature and the key feature.

[0133] In actual implementation, a dot product operation is performed on the query feature and the key feature to calculate the style similarity score between the query feature and the key feature, which measures the degree of match between the query feature and the key feature at each position. In addition, the similarity scores at each position are kept at a similar scale before being processed by the multi-class classification function to ensure the convergence of the style matching model, and the dimension of the result of the dot product operation can be scaled. Then, the result of the scaled dot product operation is applied to the multi-class classification function, and the result of the dot product operation is normalized to convert the similarity score into a probability distribution. In other words, the multi-class classification function can be used to convert the similarity scores at all positions into a probability form so that their sum is 1. Here, after converting the similarity score into a probability distribution, the first attention weight between the first text feature and the second text feature can be obtained.

[0134] In step 1023, a dot product operation is performed on the first attention weight and the value feature to obtain a third text feature.

[0135] Since the essence of the cross-attention mechanism is to quantify the similarity between the query feature and the key feature through dot product, and then assign attention weights through a multi-category classification function, and perform dot product operations on the value features based on these weights to form a contextual style representation for each position in the feature sequence of the first text feature. This mechanism dynamically pays attention to different positions in the first text feature through the query feature, and enhances the ability of the style matching model to capture the style relationship between the first text feature and the second text feature through dot product operations. Here, the third text feature is obtained by the dot product operation of the first attention weight and the value feature. The third text feature can represent the information of the first text feature after the style information of the second text feature is integrated through the attention mechanism. This method can improve the style matching model's ability to learn and match the style between the first text and the second text.

[0136] In some embodiments, Figure 3D This is a flow chart of the method for determining the fourth text feature provided in the embodiment of the present application, see Figure 3D , Figure 3A The step 102 of “adjusting the attention of the second text feature based on the first text feature to obtain the fourth text feature” can be implemented through steps 1024 to 1026 .

[0137] In step 1024, the second text feature is used as the query feature in the attention mechanism, and the first text feature is used as the key feature and value feature in the attention mechanism.

[0138] See also Figure 5 The second attention alignment learning is exactly the opposite of the first attention alignment learning process mentioned above. It reflects the dynamic interaction of the style information of the two when the second text feature focuses on the information related to its own style in the first text feature.

[0139] Specifically, the second text feature after linear processing can be used as the query feature in the attention mechanism, and the first text feature after linear layer processing can be used as the key feature and value feature in the attention mechanism, which can be used to guide the second text feature to extract relevant style information from the first text feature.

[0140] In step 1025 , a second attention weight between the second text feature and the first text feature is determined based on a dot product operation of the query feature and the key feature.

[0141] In actual implementation, a dot product operation is performed on the query feature and the key feature to calculate the style similarity score between the query feature and the key feature, which measures the degree of match between the query feature and the key feature at each position. In addition, to maintain the similarity scores at each position at a similar scale before processing through the multi-class classification function to ensure the convergence of the style matching model, the dimension of the dot product operation result can be scaled. Then, the scaled dot product operation result is applied to the multi-class classification function, the dot product operation result is normalized, and the similarity score is converted into a probability distribution. In other words, the multi-class classification function can be used to convert the similarity scores at all positions into a probability form so that their sum is 1. Here, after converting the similarity score into a probability distribution, the second attention weight between the second text feature and the first text feature can be obtained.

[0142] In step 1026, a dot product operation is performed on the second attention weight and the value feature to obtain a fourth text feature.

[0143] Similarly, since the essence of the cross-attention mechanism is to quantify the similarity between the query feature and the key feature through dot product, and then assign attention weights through a multi-category classification function, and perform dot product operations on the value features based on these weights to form a contextual style representation for each position in the feature sequence of the second text feature. This mechanism dynamically pays attention to different positions in the second text feature through the query feature, and enhances the ability of the style matching model to capture the style relationship between the second text feature and the first text feature through dot product operations. Here, the fourth text feature is obtained by the dot product operation of the second attention weight and the value feature. The fourth text feature can represent the information of the second text feature after the style information of the first text feature is fused through the attention mechanism. This method can improve the style matching model's ability to learn and match the style between the first text and the second text.

[0144] Continue to see Figure 3A , continue with the above step 102 for description.

[0145] In step 103 , a first similarity between the first text and the second text is determined based on the third text feature and the fourth text feature.

[0146] In some embodiments, Figure 3A Step 103 shown can be implemented in the following manner: determining the first vector and the second vector of the vertices on the first diagonal in the first vector matrix according to the first vector matrix corresponding to the third text feature; determining the third vector and the fourth vector of the vertices on the second diagonal in the second vector matrix according to the second vector matrix corresponding to the fourth text feature; and determining the first similarity between the first text and the second text based on the first vector, the second vector, the third vector and the fourth vector.

[0147] In actual implementation, after the cross-attention mechanism performs cross-attention learning on the first and second text features, the resulting third and fourth text features are vector matrices with consistent dimensions. That is, the third text feature is represented by the first vector matrix, and the fourth text feature is represented by the second vector matrix, and the dimensions of the first and second vector matrices are consistent. The vectors at each position in the first and second vector matrices reflect the similarity in style between the feature vectors of the first and second text features at corresponding positions. Therefore, the vectors at the vertices of the diagonal lines of the two vector matrices can be selected as elements for calculating the first similarity between the first and second texts.

[0148] In actual implementation, see Figure 5 , the first diagonal of the first vector matrix can be a diagonal formed by the upper left vertex and the lower right vertex. Similarly, the second diagonal of the second vector matrix is ​​also a diagonal formed by the upper left vertex and the lower right vertex. Then, the two vectors at the vertices of the first diagonal in the first vector matrix are used as the first vector and the second vector, and the two vectors at the vertices of the second diagonal in the second vector matrix are used as the third vector and the fourth vector.

[0149] As an example, the sum of the first vector and the second vector can be used as the second similarity between the first text and the second text determined by the first attention cross-learning, and the sum of the third vector and the fourth vector can be used as the third similarity between the first text and the second text determined by the second attention cross-learning. Then, the second similarity and the third similarity are directly added to obtain the first similarity in style between the first text and the second text predicted by the style matching model.

[0150] As an example, corresponding weights may be configured for the second similarity and the third similarity respectively, and the second similarity and the third similarity are weightedly summed to obtain the first similarity in style between the first text and the second text predicted by the style matching model.

[0151] Continue to see Figure 3A , continue with the above step 103 for explanation.

[0152] In step 104 , a first speech that matches the second text with the highest first similarity is determined in the first speech cluster, and the second speech of the first text is synthesized based on the first speech.

[0153] In actual implementation, since there are multiple second texts, it is necessary to perform a first similarity calculation on the first text and each second text separately, thereby determining the second text with the highest first similarity among the multiple second texts. The second text with the highest first similarity describes a speech style that is closest to the style embodied in the textual content of the first text. Therefore, the first speech sample contained in the first speech cluster corresponding to the second text with the highest first similarity in the speech set is also the reference speech that best matches the style of the first text.

[0154] In some embodiments, the first speech that matches the second text with the highest first similarity can be determined in the following manner: determining a second speech group corresponding to the first timbre among multiple first speech groups in the speech set, and determining multiple first speech clusters corresponding to the second speech group, where different first speech clusters have different speech styles; determining the fifth similarity between the speech style described by the second text and the speech style of each first speech cluster in the second speech group, and determining a fourth speech cluster with the greatest fifth similarity among the multiple first speech clusters corresponding to the second speech group; and taking the first speech sample in the fourth speech cluster that is closest to the cluster center of the fourth speech cluster as the first speech.

[0155] In actual implementation, when selecting a reference voice for the first text, two types of information, timbre and voice style, are usually considered. Therefore, when constructing the voice set in the embodiment of the present application, it is first grouped by timbre, that is, the voice set includes first voice groups corresponding to multiple timbres, and then, for the first voice groups of each timbre, the first voice samples contained in the first voice group are grouped according to style, that is, the first voice group includes multiple first voice clusters, and each first voice cluster has a different voice style. Therefore, the second text corresponding to each first voice cluster and used to describe its voice style is different, that is, there are usually differences between the second texts corresponding to different first voice clusters.

[0156] In actual implementation, when selecting a first speech for a first text in a speech collection, a second speech group corresponding to a first timbre can be first selected from multiple first speech groups in the speech collection, where the first timbre is specified based on actual speech synthesis requirements. Then, multiple first speech clusters corresponding to the second speech group are determined, and then, from the multiple first speech clusters corresponding to the second speech group, a fourth speech cluster is determined that matches the speech style described by the second text with the highest first similarity. All first speech samples included in the fourth speech cluster are speech that both matches the style of the text content of the first text and has the first timbre. Therefore, any first speech sample in the fourth speech cluster can be selected as the first speech of the first text, and the first speech is then input into a speech synthesis model as a reference speech for the first text. This allows the speech synthesis model to synthesize the second speech corresponding to the first text using the first speech as a reference.

[0157] As an example, a fourth speech cluster with a speech style described by a second text having the highest first similarity can be determined by determining a fifth similarity between the speech style described by the second text having the highest first similarity and the speech styles described by the second text corresponding to each first speech cluster in the second speech group, and then determining a fourth speech cluster with the highest fifth similarity among the multiple first speech clusters corresponding to the second speech group. Here, when calculating the fifth similarity between speech styles, the semantic similarity between the two can be used as the fifth similarity.

[0158] Here, the second text with the highest first similarity can be used as the third text. The fifth similarity between the speech style described by the second text and the speech style of each first speech cluster in the second speech group is determined, that is, the fifth similarity between the speech style described by the third text and the speech style described by the second text corresponding to each first speech cluster in the second speech group is determined. The first speech cluster with the highest fifth similarity to the speech style described by the third text is used as the fourth speech cluster, that is, the fourth speech cluster that meets the speech style described by the second text (i.e., the third text) with the highest first similarity.

[0159] As an example, Figure 6 This is a schematic diagram of the speech synthesis model provided in the embodiment of the present application, see Figure 6 After the first voice is determined through the fourth voice cluster, a prompt for synthesizing the voice can be constructed for the voice synthesis model based on a preset template (i.e., prompt), and the first voice is used as a case in the prompt, that is, the first voice is used as a reference voice of the voice synthesis model, and the first text and the prompt carrying the first voice are input into the voice synthesis model, so that the voice synthesis model synthesizes the second voice corresponding to the first text based on the instructions given by the prompt and with the timbre and voice style of the first voice as a reference.

[0160] In actual implementation, since the fourth speech cluster is obtained by clustering, the first speech sample in the fourth speech cluster that is closer to the cluster center of the fourth speech cluster is usually more consistent with the speech style described by the second text corresponding to the fourth speech cluster. Therefore, the first speech sample in the fourth speech cluster that is closest to the cluster center of the fourth speech cluster can also be used as the first speech, and the first speech can be used as the reference speech of the first text and input into the speech synthesis model, so that the speech synthesis model uses the first speech as a reference to synthesize the second speech corresponding to the first text.

[0161] In the above manner, when performing speech synthesis on a first text to be synthesized, first text features of the first text are extracted, and second text features of the second text used to describe the speech style are extracted. Based on the attention mechanism, the attention of the first text features is adjusted by the second text features to obtain third text features, and the attention of the second text features is adjusted by the first text features to obtain fourth text features. In this way, the similarity in style between the first text and the second text can be fully learned through two cross-attention learnings. Then, the first similarity in style between the first text and the second text is calculated based on the third text features and the fourth text features, so that the second speech of the first text is synthesized based on the first speech that matches the second text with the highest first similarity. In this way, since the speech style described by the second text with the highest first similarity is most similar to the style expressed by the text content of the first text, and the first speech matches the second text with the highest first similarity, the first speech determined in this way matches the first text in style. With the first speech as a reference, the second speech synthesized for the first text is more consistent with the style expressed by the first text.

[0162] The process of the above embodiment is a process of inference based on a trained style matching model. With the help of the introduction of the style matching model in the above embodiment, the training process of the style matching model is schematically explained.

[0163] In some embodiments, the training process of the style matching model may include: performing feature extraction on the first text sample to obtain a fifth text feature, and performing feature extraction on the second text sample used to describe the speech style to obtain a sixth text feature; performing attention adjustment on the fifth text feature based on the sixth text feature to obtain a seventh text feature, and performing attention adjustment on the sixth text feature based on the fifth text feature to obtain an eighth text feature; determining a first loss based on the seventh text feature and the first label of the first text sample, and determining a third loss based on the eighth text feature and the second label of the second text sample; and updating the model parameters based on the first loss and the second loss.

[0164] The process of extracting features from the first text sample and the second text sample may refer to the feature extraction process for the first text and the second text in the above embodiment, which will not be described in detail here.

[0165] Here, the first label of the first text sample used in calculating the first loss can be the style probability annotated by the feature vector at each position of the fifth text feature of the first text sample. When calculating the loss, the difference between the style probability predicted at each position of the seventh text feature and the style probability of the fifth text feature at the corresponding position is used for calculation. Similarly, the second representation of the second text sample used in calculating the second loss is the style probability annotated by the feature vector at each position of the sixth text feature of the second text sample. When calculating the loss, the difference between the style probability predicted at each position of the eighth text feature and the style probability of the sixth text feature at the corresponding position is used for calculation.

[0166] See also Figure 5 It can be understood that the process of adjusting the attention of the fifth text feature based on the sixth text feature to obtain the seventh text feature is equivalent to the first attention alignment learning process described above, and the process of adjusting the attention of the sixth text feature based on the fifth text feature to obtain the eighth text feature is equivalent to the second attention alignment learning process described above. When training the style matching model, after performing the first attention alignment learning and the second attention alignment learning on the first and second text samples, the cross-entropy loss can be calculated after processing through the normalization layer. That is, the first loss of the first attention alignment learning and the second loss of the second attention alignment learning are calculated, and the sum of the first loss and the second loss is used as the total loss of the style matching model. The parameters of the style matching model are adjusted and optimized in reverse order until the style matching model converges.

[0167] The following continues to describe the exemplary structure of the speech synthesis device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the speech synthesis device 455 of the memory 450 may include:

[0168] The feature extraction module 4551 is configured to perform feature extraction on the first text to obtain first text features, and perform feature extraction on the second text to obtain second text features. The second text is used to describe the speech style of the first speech cluster.

[0169] The feature adjustment module 4552 is configured to adjust the attention of the first text feature based on the second text feature to obtain a third text feature, and to adjust the attention of the second text feature based on the first text feature to obtain a fourth text feature.

[0170] The similarity determination module 4553 is configured to determine a first similarity between the first text and the second text based on the third text feature and the fourth text feature.

[0171] The speech synthesis module 4554 is configured to determine, in the first speech cluster, a first speech that matches the second text with the highest first similarity, and synthesize the second speech of the first text based on the first speech.

[0172] In some embodiments, the feature adjustment module 4552 is also used to use the first text feature as the query feature in the attention mechanism and the second text feature as the key feature and value feature in the attention mechanism; determine the first attention weight between the first text feature and the second text feature based on the dot product operation of the query feature and the key feature; perform a dot product operation on the first attention weight and the value feature to obtain a third text feature.

[0173] In some embodiments, the feature adjustment module 4552 is further used to use the second text feature as the query feature in the attention mechanism and the first text feature as the key feature and value feature in the attention mechanism; determine the second attention weight between the second text feature and the first text feature based on the dot product operation of the query feature and the key feature; perform a dot product operation on the second attention weight and the value feature to obtain a fourth text feature.

[0174] In some embodiments, the speech synthesis device 455 also includes a text generation module for constructing a speech set, the speech set including a first speech group corresponding to a plurality of timbres, the first speech group including a plurality of first speech samples matching the timbres; for the first speech group corresponding to each timbre, style features are extracted for each first speech sample in the first speech group to obtain the style features of the first speech sample; based on the style features of the first speech samples, the plurality of first speech samples in the first speech group are clustered to obtain a first speech cluster; for each first speech cluster, based on the style features of the first speech samples in the first speech cluster, a second text for describing the speech style of the first speech cluster is generated.

[0175] In some embodiments, the text generation module is also used to cluster the first speech samples in the speech set based on the semantic features of each first speech sample in the speech set to obtain a second speech cluster; determine the second similarity between the semantic features corresponding to the cluster center of the second speech cluster and the semantic features of each first speech sample in the first speech group, and screen out the second speech sample with the largest second similarity from the multiple first speech samples in the first speech group; use the screened second speech sample as the cluster center, and cluster the multiple first speech samples in the first speech group based on the style features of each first speech sample in the first speech group to obtain a first speech cluster.

[0176] In some embodiments, the text generation module is further used to select a third speech sample from each first speech sample in the speech set; use the third speech sample as the cluster center of the second speech cluster, and cluster the first speech samples in the speech set based on the third similarity between the semantic features of the third speech sample and the semantic features of the fourth speech sample in the speech set to obtain a second speech cluster; wherein the third similarity between the semantic features of the first speech samples in the second speech cluster is greater than the first similarity threshold.

[0177] In some embodiments, the text generation module is further used to cluster the first speech samples of the speech set based on the fourth similarity between the semantic features of each first speech sample in the speech set to obtain a third speech cluster, where the fourth similarity between the semantic features of the first speech samples in the third speech cluster exceeds the second similarity threshold; and to filter out a second speech cluster from the third speech cluster in which the number of first speech samples exceeds the number threshold.

[0178] In some embodiments, the speech synthesis module 4554 is also used to determine a second speech group corresponding to a first timbre among multiple first speech groups in a speech set, and determine multiple first speech clusters corresponding to the second speech group, where different first speech clusters have different speech styles; determine the fifth similarity between the speech style described by the second text and the speech style of each first speech cluster in the second speech group, and determine a fourth speech cluster with the largest fifth similarity among the multiple first speech clusters corresponding to the second speech group; and use the first speech sample in the fourth speech cluster that is closest to the cluster center of the fourth speech cluster as the first speech.

[0179] In some embodiments, the text generation module is also used to perform feature fusion on the pitch feature, volume feature and Mel spectrum feature of the first speech sample to obtain a first latent feature of the first speech sample; extract the voiceprint feature of the first speech sample based on the Mel spectrum feature, and eliminate the voiceprint feature from the first latent feature to obtain a second latent feature; quantize and encode the second latent feature to obtain a third latent feature; and predict the style feature of the first speech sample based on the third latent feature to obtain the style feature of the first speech sample.

[0180] In some embodiments, the similarity determination module 4553 is further used to determine the first vector and the second vector of the vertices on the first diagonal in the first vector matrix based on the first vector matrix corresponding to the third text feature; determine the third vector and the fourth vector of the vertices on the second diagonal in the second vector matrix based on the second vector matrix corresponding to the fourth text feature; and determine the first similarity between the first text and the second text based on the first vector, the second vector, the third vector and the fourth vector.

[0181] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the speech synthesis method described in the present invention.

[0182] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the speech synthesis method provided by the embodiment of the present application, for example, Figure 3A The speech synthesis method is shown.

[0183] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0184] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0185] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0186] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0187] In summary, according to the embodiments of the present application, when performing speech synthesis on a first text, first text features of the first text are extracted, and second text features of the second text used to describe the speech style of the first speech cluster are extracted. Based on the attention mechanism, the first text features are adjusted by the second text features to obtain third text features, and the second text features are adjusted by the first text features to obtain fourth text features. In this way, the similarity in style between the first text and the second text can be fully learned through two cross-attention learnings. Then, the first similarity in style between the first text and the second text is calculated based on the third text features and the fourth text features, so that the first speech that matches the second text with the highest first similarity is determined in the first speech cluster, and the second speech of the first text is synthesized based on the first speech. In this way, since the speech style described by the second text with the highest first similarity is most similar to the style expressed by the text content of the first text, and the first speech matches the second text with the highest first similarity, the first speech determined in this way matches the first text in style. With the first speech as a reference, the second speech synthesized for the first text is more consistent with the style expressed by the first text.

[0188] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A speech synthesis method, characterized in that: The method comprises: Performing feature extraction on the first text to obtain first text features, and performing feature extraction on the second text to obtain second text features, wherein the second text is used to describe the speech style of the first speech cluster; performing attention adjustment on the first text feature based on the second text feature to obtain a third text feature, and performing attention adjustment on the second text feature based on the first text feature to obtain a fourth text feature; determining a first similarity between the first text and the second text based on the third text feature and the fourth text feature; A first speech that matches a second text with the highest first similarity is determined in the first speech cluster, and a second speech of the first text is synthesized based on the first speech.

2. The method according to claim 1, characterized in that The step of adjusting the attention of the first text feature based on the second text feature to obtain a third text feature includes: Using the first text feature as a query feature in an attention mechanism, and using the second text feature as a key feature and a value feature in the attention mechanism; determining a first attention weight between the first text feature and the second text feature based on a dot product operation of the query feature and the key feature; Perform a dot product operation on the first attention weight and the value feature to obtain a third text feature.

3. The method according to claim 1, characterized in that The step of adjusting the attention of the second text feature based on the first text feature to obtain a fourth text feature includes: Using the second text feature as a query feature in an attention mechanism, and using the first text feature as a key feature and a value feature in the attention mechanism; determining a second attention weight between the second text feature and the first text feature based on a dot product operation of the query feature and the key feature; Perform a dot product operation on the second attention weight and the value feature to obtain a fourth text feature.

4. The method according to claim 1, wherein Before extracting features from the second text, the method further includes: Constructing a speech set, the speech set including first speech groups corresponding to a plurality of timbres, wherein the first speech groups include a plurality of first speech samples matching the timbres; For the first speech groups corresponding to the respective timbres, extracting style features of the first speech samples in the first speech groups to obtain style features of the first speech samples; clustering the plurality of first speech samples in the first speech group based on the style features of the first speech samples to obtain a first speech cluster; For each of the first speech clusters, a second text is generated for describing the speech style of the first speech cluster based on the style features of the first speech samples in the first speech cluster.

5. The method according to claim 4, characterized in that The clustering of the plurality of first speech samples in the first speech group based on the style features of the first speech samples to obtain a first speech cluster includes: clustering the first speech samples in the speech set based on the semantic features of each of the first speech samples in the speech set to obtain second speech clusters; Determining a second similarity between a semantic feature corresponding to a cluster center of the second speech cluster and a semantic feature of each of the first speech samples in the first speech group, and selecting a second speech sample having the greatest second similarity from the plurality of first speech samples in the first speech group; The screened second speech sample is used as a clustering center, and based on the style characteristics of each of the first speech samples in the first speech group, multiple first speech samples in the first speech group are clustered to obtain a first speech cluster.

6. The method according to claim 5, characterized in that The clustering of the first speech samples in the speech set based on the semantic features of each of the first speech samples in the speech set to obtain second speech clusters includes: Selecting a third speech sample from each of the first speech samples in the speech set; Taking the third speech sample as the cluster center of the second speech cluster, clustering the first speech sample in the speech set based on a third similarity between the semantic features of the third speech sample and the semantic features of the fourth speech sample in the speech set to obtain a second speech cluster; The third similarity between the semantic features of the first speech samples in the second speech cluster is greater than the first similarity threshold.

7. The method according to claim 5, characterized in that The clustering of the first speech samples in the speech set based on the semantic features of each of the first speech samples in the speech set to obtain second speech clusters includes: clustering the first speech samples of the speech set based on a fourth similarity between the semantic features of the first speech samples of the speech set to obtain a third speech cluster, wherein the fourth similarity between the semantic features of the first speech samples in the third speech cluster exceeds a second similarity threshold; A second speech cluster is selected from the third speech cluster, wherein the number of the first speech samples exceeds a number threshold.

8. The method according to claim 5, characterized in that The step of determining, in the first speech cluster, a first speech that matches the second text having the highest first similarity, includes: Determining a second voice group corresponding to a first timbre from a plurality of first voice groups in the voice set, and determining a plurality of first voice clusters corresponding to the second voice group, wherein different first voice clusters have different voice styles; determining a fifth similarity between the speech style described by the second text and the speech styles of each of the first speech clusters in the second speech group, and determining a fourth speech cluster having the greatest fifth similarity among the plurality of first speech clusters corresponding to the second speech group; The first speech sample in the fourth speech cluster that is closest to the cluster center of the fourth speech cluster is used as the first speech.

9. The method according to claim 4, characterized in that The extracting style features of each of the first speech samples in the speech group to obtain the style features of the first speech samples includes: Performing feature fusion on the pitch feature, volume feature, and mel-spectrogram feature of the first speech sample to obtain a first latent feature of the first speech sample; Extracting a voiceprint feature of the first speech sample based on the mel-spectrogram feature, and removing the voiceprint feature from the first latent feature to obtain a second latent feature; Quantizing and encoding the second latent feature to obtain a third latent feature; The style feature of the first speech sample is predicted based on the third latent feature to obtain the style feature of the first speech sample.

10. The method according to claim 1, characterized in that The determining, based on the third text feature and the fourth text feature, a first similarity between the first text and the second text includes: Determining, according to the first vector matrix corresponding to the third text feature, first vectors and second vectors of vertices on a first diagonal line in the first vector matrix; Determining, according to the second vector matrix corresponding to the fourth text feature, a third vector and a fourth vector of vertices on a second diagonal line in the second vector matrix; A first similarity between the first text and the second text is determined based on the first vector, the second vector, the third vector, and the fourth vector.

11. A speech synthesis device, characterized in that: The device comprises: a feature extraction module, configured to extract features from the first text to obtain first text features, and to extract features from the second text to obtain second text features, wherein the second text is used to describe the speech style of the first speech cluster; a feature adjustment module, configured to perform attention adjustment on the first text feature based on the second text feature to obtain a third text feature, and to perform attention adjustment on the second text feature based on the first text feature to obtain a fourth text feature; a similarity determination module, configured to determine a first similarity between the first text and the second text based on the third text feature and the fourth text feature; The speech synthesis module is configured to determine a first speech that matches the second text with the highest first similarity in the first speech cluster, and synthesize a second speech of the first text based on the first speech.

12. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the speech synthesis method according to any one of claims 1 to 10 when executing the computer-executable instructions or computer program stored in the memory.

13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the speech synthesis method according to any one of claims 1 to 10 is implemented.

14. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the speech synthesis method according to any one of claims 1 to 10 is implemented.