Speech synthesis method and device, model training method and device, electronic equipment, computer readable storage medium and computer program product
By performing word segmentation and sentence component analysis on text samples and expanding the speech sequence to match the speech bit rate, the stability and effect issues in speech synthesis model training are solved, and more accurate and natural speech synthesis is achieved.
Patent Information
- Application Number
- CN202510681276.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-23
AI Technical Summary
Existing speech synthesis technology suffers from an imbalance in the bit rates of text tokens and voice tokens, resulting in poor model training stability and poor training results. It is also difficult to control the voice style and prone to problems such as phantom reading, repeated reading, and misreading.
By performing word segmentation on text samples, determining the sentence components and word order structure of words, and expanding them to obtain sequences that match the speech unit bit rate of the speech samples, speech synthesis is performed. The bit rates of text and speech are aligned when training the model, which solves the problems of instability and poor results in model training.
Improves the accuracy and naturalness of speech synthesis, ensures that the synthesized speech conforms to the preset timbre style, and enhances personalization and semantic expression.
Smart Images

Figure CN120690175A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to speech synthesis technology, and in particular to a speech synthesis method, model training method, device, electronic device, computer-readable storage medium and computer program product. Background Art
[0002] Text-to-speech synthesis is a technology that converts written text into human-audible speech. Using computer algorithms and language processing techniques, it transforms textual input into natural, fluent speech. Text-to-speech synthesis technology is widely used in voice assistants, accessibility services, education, entertainment, navigation systems, and other fields. Summary of the Invention
[0003] The embodiments of the present application provide a speech synthesis method, a model training method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which can make the speech synthesized for the first text more consistent with a preset timbre style.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present invention provides a method for speech synthesis, comprising:
[0006] Performing word segmentation on the first text to obtain a first sequence, where the first sequence includes a plurality of words obtained by word segmentation of the first text;
[0007] Determining the sentence component of each of the words in the first text based on the part of speech of each of the words, and determining the word order structure of the first text according to the arrangement order of each of the words in the first text;
[0008] Expanding the first sequence based on the sentence components of each of the words in the first text and the word order structure of the first text to obtain a second sequence;
[0009] Speech synthesis is performed based on the second sequence and the first timbre to obtain a first speech of the first text.
[0010] The present invention provides a model training method, which includes:
[0011] Performing word segmentation on the text sample to obtain a third sequence, and performing speech unit division on the speech sample to obtain a fourth sequence, wherein the first sample sequence includes a plurality of words obtained by word segmentation of the text sample, and the fourth sequence includes a plurality of speech units obtained by division of the speech sample;
[0012] Determining the sentence component of each of the words in the text sample based on the part of speech of each of the words, and determining the word order structure of the text sample according to the order of each of the words in the text sample;
[0013] Based on the sentence components of each of the words in the text sample and the word order structure of the first text, the third sequence is expanded to obtain a fifth sequence;
[0014] The fourth sequence and the fifth sequence are spliced together, and a model is trained based on the splicing result.
[0015] The present invention provides a speech synthesis device, comprising:
[0016] A word segmentation module, configured to perform word segmentation on the first text to obtain a first sequence, wherein the first sequence includes a plurality of words obtained by word segmentation of the first text;
[0017] a determination module, configured to determine a sentence component of each of the words in the first text based on the part of speech of each of the words, and determine a word order structure of the first text according to the order of each of the words in the first text;
[0018] an expansion module, configured to expand the first sequence based on the sentence components of each of the words in the first text and the word order structure of the first text to obtain a second sequence;
[0019] A synthesis module is used to perform speech synthesis based on the second sequence and the first timbre to obtain a first speech of the first text.
[0020] The present invention provides a model training device, comprising:
[0021] a sequence extraction module configured to perform word segmentation on the text sample to obtain a third sequence, the third sequence including a plurality of words obtained by word segmentation of the text sample, and to perform speech unit segmentation on the speech sample to obtain a fourth sequence, the fourth sequence including a plurality of speech units obtained by segmenting the speech sample;
[0022] a text analysis module, configured to determine the sentence component of each of the words in the text sample based on the part of speech of each of the words, and to determine the word order structure of the text sample according to the order of each of the words in the text sample;
[0023] a sequence processing module, configured to expand the third sequence to obtain a fifth sequence based on the sentence components of each word in the text sample and the word order structure of the text sample, and to concatenate the fourth sequence and the fifth sequence to obtain a sixth sequence;
[0024] The model training module is used to train the model based on the sixth sequence and the timbre of the speech sample.
[0025] An embodiment of the present application provides an electronic device, comprising:
[0026] a memory for storing computer-executable instructions or computer programs;
[0027] The processor is used to implement the speech synthesis method provided in the embodiment of the present application when executing the computer-executable instructions or computer program stored in the memory.
[0028] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the speech synthesis method provided in the embodiment of the present application when executed by a processor.
[0029] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the speech synthesis method provided in the embodiment of the present application is implemented.
[0030] The embodiments of the present application have the following beneficial effects:
[0031] In the above manner, when performing speech synthesis on a first text to be synthesized, the first text is segmented to obtain a first sequence consisting of multiple words obtained by segmenting the first text; the sentence component of each word in the first text is determined based on the part of speech of each word, and the word order structure of the first text is determined based on the order of each word in the first text; the first sequence is expanded based on the sentence component of each word in the first text and the word order structure of the first text to obtain a second sequence, and speech synthesis is performed based on the second sequence and the first timbre to obtain the first speech of the first text; thus, since determining the sentence component of each word in the first text helps to understand the importance and function of each word in the first text, and determining the word order structure of the first text helps to capture the semantics and emotional expression of the first text, the expansion of the first sequence based on the sentence component and word order structure can enrich the representation of the text. Therefore, when performing speech synthesis on the first text based on the second sequence and the preset first timbre, the rich text representation in the second sequence and the timbre characteristics of the first timbre can be used to generate speech that matches the text content and timbre requirements, thereby improving the accuracy, naturalness, and personalization of the speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a schematic diagram of the structure of the speech synthesis system architecture provided by an embodiment of the present application;
[0033] Figures 2A-2B is a structural diagram of an electronic device provided in an embodiment of the present application;
[0034] Figure 3A This is a schematic diagram of the first flow chart of the speech synthesis method provided in an embodiment of the present application;
[0035] Figure 3B This is a second flow chart of the speech synthesis method provided in an embodiment of the present application;
[0036] Figure 4 This is a data flow diagram of the model training provided in the embodiment of the present application;
[0037] Figure 5 It is a flowchart of the model training method provided in the embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0039] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0040] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0041] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0042] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0043] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0044] Text-to-speech synthesis is a technology that converts written text into human-audible speech. Using computer algorithms and language processing techniques, it transforms textual input into natural, fluent speech. Text-to-speech synthesis technology is widely used in voice assistants, accessibility services, education, entertainment, navigation systems, and other fields.
[0045] Among them, the Speech Generation Large Model (SGLM) is the core of text-to-speech synthesis technology. The Speech Generation Large Model is responsible for converting the input text into speech signals. When training the Speech Generation Large Model, it is usually necessary to input the text sample to be synthesized, the speech sample corresponding to the text sample, and the timbre of the speech sample (used to guide the timbre of the synthesized speech). The timbre of the speech sample can be input into the Speech Generation Large Model through a template (also called reference or borrowing). The Speech Generation Large Model can convert the input text sample into speech with a preset timbre based on the timbre of the speech sample. The content of the speech is the content of the text sample. A loss function is constructed based on the speech output by the Speech Generation Large Model and the speech sample corresponding to the text sample, and the model parameters of the Speech Generation Large Model are updated based on the loss function.
[0046] Here, the term "template" can be used as a reference or borrowed term. It refers to using the template to provide additional conditional variables for the speech synthesis model. This means providing the model with a case study, letting it know the desired content, style, and timbre of the speech. In text-to-speech synthesis, templates typically take the form of paired reference examples, such as a <text, voice> pair. This pair tells the speech synthesis model what the user expects when synthesizing speech corresponding to the text. Templates serve as a reference or model, containing at least one piece of information about the preset timbre, which serves as a target for the speech synthesis model. Based on the preset timbre, the speech synthesis model can analyze timbre information (i.e., the speaker's intonation, enabling speaker identification based on timbre), prosody information (i.e., acoustic characteristics such as speaking rate, pauses, pitch, and volume), emotional information (emotional information such as joy, anger, sadness, and happiness), and personalized information (e.g., leaks, emphasis, catchphrases, and the use of markers).
[0047] Large speech synthesis models can be trained using semi-supervised or unsupervised deep learning techniques and can synthesize speech in various styles with near-human-like speech effects. In terms of inference methods, large speech synthesis models include autoregressive and non-autoregressive models. Because autoregressive models are based on word units (tokens), inference based on these models can only be generated token by token, not parallelized. Therefore, the token bitrate (the number of tokens encoded per second of speech) is crucial to inference efficiency. For example, when the token bitrate of speech is too low, the compression ratio of the token representation to the original speech signal is too high, which loses many speech details and significantly impacts the subsequent training of the large speech synthesis model. Therefore, the token bitrate of mainstream speech is not lower than 20Hz, with a typical token bitrate of 25Hz. The bitrate of text is often significantly lower than the token bitrate of speech. For example, using text phoneme encoding, assuming an average speaking speed of 4.5 words per second, 4.5 Chinese characters contain a total of 9 pinyin initials and finals, with a bitrate of only 9Hz, significantly lower than the speech token bitrate (e.g., 25Hz). When training large speech synthesis models, the density of the concatenated representation of text and speech tokens is approximately three times that of the text representation. This imbalance in bitrates often leads to the following problems: First, during training, it is difficult for the model to learn the alignment between text and phonemes, which limits model performance and makes it prone to phantom reading (i.e., the synthesized speech does not correspond to the input text to be synthesized, resulting in omissions, repetitions, misreadings, and inaccurate pronunciation). Second, it is even more difficult to control the speech style of the synthesized speech, which includes but is not limited to speech rate, emotion, and prosodic characteristics.
[0048] To this end, the embodiments of the present application provide a speech synthesis method, a model training method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product; these methods are capable of aligning the bit rate of text tokens and the bit rate of speech tokens, thereby solving the problems of poor model training stability and poor training effect, and making the speech synthesized for text more consistent with a preset timbre style.
[0049] It should be noted that the speech synthesis method provided in the embodiment of the present application is a text-to-speech synthesis technology that can be used in various application scenarios such as assisted reading, intelligent voice interaction, dubbing, customer service and marketing, public information broadcasting, personalized voice assistants, etc.
[0050] See also Figure 1 , Figure 1This is an architectural diagram of the speech synthesis system 100 provided in an embodiment of the present application. To support a speech synthesis application, terminals (terminal 400-1 and terminal 400-2 are shown as examples) are connected to the server 200 via a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0051] Here, we will use customer service and marketing scenarios as examples to illustrate.
[0052] In some embodiments, the terminal (e.g., terminal 400-1, terminal 400-2) and server 200 may jointly serve as the execution body to execute the speech synthesis method of the embodiment of the present application. Specifically, the user may ask questions to the intelligent customer service through the terminal (e.g., terminal 400-1, terminal 400-2) by phone or online customer service; the terminal (e.g., terminal 400-1, terminal 400-2) responds to the user's question and generates a response text for the question content based on the large language model, and uses the response text as the first text, and generates a speech synthesis request for the first text, and uses the response text as the first text. A speech synthesis request is sent to server 200. In response to the speech synthesis request, server 200 performs word segmentation on the first text to obtain a first sequence, wherein the first sequence includes multiple words obtained by word segmentation of the first text; determines the sentence components of each word in the first text based on the part of speech of each word, and determines the word order structure of the first text based on the order of each word in the first text; expands the first sequence based on the sentence components of each word in the first text and the word order structure of the first text to obtain a second sequence; performs speech synthesis based on the second sequence and the first timbre to obtain the first speech of the first text. Subsequently, server 200 sends the first speech to the terminal (e.g., terminal 400-1, terminal 400-2), so that the terminal (e.g., terminal 400-1, terminal 400-2) outputs the first speech to the user through intelligent customer service.
[0053] In some embodiments, the speech synthesis method provided in the embodiments of the present application can be independently executed by a terminal (e.g., terminal 400-1, terminal 400-2). Specifically, a user can ask a question to the intelligent customer service through a telephone or online customer service at the terminal (e.g., terminal 400-1, terminal 400-2); the terminal (e.g., terminal 400-1, terminal 400-2) generates a response text for the question content based on the large language model in response to the user's question, and uses the response text as the first text. Then, the terminal (e.g., terminal 400-1, terminal 400-2) performs word segmentation on the first text to obtain a first sequence, wherein the first sequence includes multiple words obtained by word segmentation of the first text; determines the sentence component of each word in the first text based on the part of speech of each word, and determines the word order structure of the first text based on the order of each word in the first text; expands the first sequence based on the sentence component of each word in the first text and the word order structure of the first text to obtain a second sequence; performs speech synthesis based on the second sequence and the first timbre to obtain the first speech of the first text. The terminal (eg, terminal 400 - 1 , terminal 400 - 2 ) outputs the first voice to the user through the intelligent customer service.
[0054] In some embodiments, the model training method of the embodiment of the present application can be jointly executed by the terminal (for example, terminal 400-1, terminal 400-2) and the server 200 as the execution body. Specifically, the user can obtain the training samples (including text samples, voice samples corresponding to the text samples, and timbre of the voice samples) constructed in the telephone or online customer service application scenario at the terminal (for example, terminal 400-1, terminal 400-2), and send the model training request to the server 200 when triggering the model training request; when the server 200 trains the large speech synthesis model, it segments the text sample to obtain a third sequence, and divides the voice sample into voice units to obtain a fourth sequence, wherein the third sequence includes multiple words (i.e., text tokens) obtained by segmenting the text sample, and the fourth sequence includes multiple voice units (i.e., voice units) obtained by dividing the voice sample. Based on the part of speech of each word, the sentence component of each word in the text sample (a text token) is determined, and the word order structure of the text sample (also a text token) is determined according to the order of each word in the text sample; based on the sentence component of each word in the text sample and the word order structure of the text sample, the third sequence is expanded to obtain a fifth sequence (i.e., rich text token), so that the bit rate of the rich text token in the fifth sequence is equivalent to the bit rate of the speech token in the fourth sequence corresponding to the speech sample, and then after splicing the fourth and fifth sequences to obtain a sixth sequence, a large speech synthesis model is trained based on the sixth sequence, thereby effectively solving the problem of poor model training stability and poor training effect due to the significant gap between the bit rate of text tokens and the bit rate of speech tokens during the training of the large speech synthesis model.
[0055] In some embodiments, the terminal (e.g., terminal 400-1, terminal 400-2) can be implemented as various types of terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, car terminals, etc., and can also be implemented as servers.
[0056] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.
[0057] See also Figures 2A-2B , Figures 2A-2Bis a structural diagram of an electronic device 400 provided in an embodiment of the present application, Figures 2A-2B The electronic device 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figures 2A-2B Various buses are labeled as bus system 440 .
[0058] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0059] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0060] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0061] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0062] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0063] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0064] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);
[0065] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0066] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.
[0067] In some embodiments, the speech synthesis device provided in the embodiments of the present application can be implemented in software. Figure 2A A speech synthesis device 455A stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a word segmentation module 4551A, a determination module 4552A, an expansion module 4553A, and a synthesis module 4554A. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0068] In some embodiments, the model training device provided in the embodiments of the present application can be implemented in software. Figure 2B Model training device 455B stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: sequence extraction module 4551B, text analysis module 4552B, sequence processing module 4553B, and model training module 4554B. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0069] In other embodiments, the speech synthesis device provided in the embodiments of the present application can be implemented in hardware. As an example, the speech synthesis device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the speech synthesis method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0070] In some embodiments, the terminal or server can implement the speech synthesis method provided in the embodiment of the present application by running various computer executable instructions or computer programs. For example, the computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP; it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.
[0071] The following describes the speech synthesis method provided by the embodiments of the present application in conjunction with the accompanying drawings. As previously mentioned, the electronic device 400 that implements the speech synthesis method of the embodiments of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.
[0072] The speech synthesis method of the embodiment of the present application is described by taking the execution subject as a terminal as an example. Figure 3A , Figure 3A This is a first flow chart of the speech synthesis method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.
[0073] In step 101 , a first text is segmented to obtain a first sequence, wherein the first sequence includes a plurality of words obtained by segmenting the first text.
[0074] Here, the first text is the text used for synthesizing speech. In different application scenarios, the method for obtaining and determining the first text can be different. For example, in an assisted reading scenario, the first text can be at least part of the text in a book, article, webpage content, email, etc., such as a paragraph or chapter in an article; in the field of dubbing, the first text can be the lines of a virtual character or actor; in the field of customer service and marketing, the first text can be the reply text of the intelligent customer service to the customer's question. As an example, in the field of customer service and marketing, the first text can be generated by obtaining the customer's question information, inputting the question information into a large language model, generating a response text for the question information through the large language model, and using the response text generated by the large language model as the first text for speech synthesis.
[0075] When segmenting the first text, the first text may be segmented using a preset segmentation algorithm (e.g., rule-based segmentation, statistical segmentation, machine learning segmentation, and deep learning segmentation). For example, the segmentation algorithm divides the first text into meaningful word units (i.e., words) based on factors such as a vocabulary and contextual information. After the segmentation process is completed, the resulting multiple words are arranged in the order in which they appear in the original text (i.e., the first text) to form a first sequence. The first sequence may be a list, an array, or other structure suitable for storing sequential data.
[0076] Taking the first text "The quick brown fox jumped over the lazy dog" as an example, the first sequence obtained by segmenting the first text is: [quick-brown-fox-jumped-lazy-dog]. After dividing the first text into multiple words, each word after segmentation can be annotated with its part of speech, such as noun, verb, adjective, etc., that is, the part of speech of each word is used as the basic tag of each word. For example, in the above example, the part of speech of each word is annotated as follows: quick / adjective, brown / adjective, fox / noun, skip / verb, jumped / particle, lazy / adjective, dog / noun. These basic tags are the above-mentioned text tokens.
[0077] In step 102 , the sentence component of each word in the first text is determined based on the part of speech of each word, and the word order structure of the first text is determined according to the order of each word in the first text.
[0078] Here, after tokenizing the first text to obtain multiple words, the sentence component of each word in the first text, such as subject, predicate, object, modifier, etc., can be determined based on the part of speech of each word. For example, continuing with the above example, the sentence component tags corresponding to each word are as follows: fast / modifier, brown / modifier, fox / subject, skip / predicate, finished / perfect tense marker, lazy / modifier, dog / object.
[0079] The word order structure can reflect the structure of the first text (such as sequence, inversion, ellipsis, etc.). Therefore, according to the order of each word in the first text or the structure of the first text, it can be determined which word order structure the first text belongs to, such as sequential, reverse, or unknown style. For example, in the above example, since the structure of the first text is "subject + modifier + predicate + object", the word order structure of the first text belongs to the sequential structure.
[0080] Through the above method, through accurate word - part tagging and determination of sentence components, it is convenient for the machine to better understand the structure and meaning of the first text. The determination of sentence components helps the machine to perform a more in - depth semantic analysis. For example, in the above example, knowing that "fox" is the subject, it can be inferred that it is the entity performing the action "jumped over"; knowing that "quick" and "brown" are modifiers describing the "fox", it can be understood what kind of fox it is; knowing that "dog" is the object, it can be understood that the "dog" is the object that the "fox" jumped over, thereby inferring the specific content of the action, which has a significant beneficial effect on improving the accuracy of speech synthesis.
[0081] In step 103, based on the sentence components of each word in the first text and the word order structure of the first text, the first sequence is expanded to obtain the second sequence.
[0082] In some embodiments, the first sequence can be expanded to obtain the second sequence based on the sentence components of each word in the first text and the word order structure of the first text in the following manner: Based on the word part of each word, the sentence components of each word in the first text, and the word order structure of the first text, each word in the first sequence is marked to obtain the second sequence; where the second sequence includes subsequences corresponding to each word, and the subsequence consists of the following word elements: the word and the identification information of the word, where the identification information includes: the word part of the word, the sentence component of the word in the first text, and the word order structure of the first text.
[0083] Here, after determining the sentence components of each word in the first text and the word order structure of the first text, the sentence components and the word order structure can be added as additional marks for each word to the first sequence, and these additional marks are also called the above - mentioned text tokens.
[0084] For example, continuing with the above example, adding the analysis results of the sentence components and word order structure of each word in the first text to the first sequence, the obtained second sequence is: [quick / adjective / modifier / sequence - brown / adjective / modifier / sequence - fox / noun / subject / sequence - jumped over / verb / predicate / sequence - le / auxiliary / completion marker / sequence - lazy / adjective / modifier / sequence - dog / noun / object / sequence].
[0085] It can be seen therefrom that the second sequence is obtained by arranging the subsequences corresponding to each word in the order of each word in the first text. Each subsequence consists of word elements such as each word and the identification information of the word. For example, for the word "fast", its corresponding subsequence is: [fast / adjective / modifier / order]; for the word "brown", its corresponding subsequence is: [brown / adjective / modifier / order]; for the word "fox", its corresponding subsequence is: [fox / noun / subject / order]; for the word "jumped", its corresponding subsequence is: [jumped / verb / predicate / order]; for the word "over", its corresponding subsequence is: [over / particle / completion marker / order]; for the word "lazy", its corresponding subsequence is: [lazy / adjective / modifier / order]; for the word "dog", its corresponding subsequence is: [dog / noun / object / order]. The respective subsequences are arranged in the order of the corresponding words in the original text (i.e., the first text) to form the second sequence.
[0086] It should be noted that in practical applications, the code rate of the tokens of the text corresponding to the expanded second sequence is not the larger the better. Instead, according to the code rate of the tokens of the speech for speech synthesis, by expanding the first sequence into the second sequence, the code rate of the tokens of the text corresponding to the first text in the second sequence is made to be comparable to the code rate of the tokens of the speech.
[0087] In step 104, speech synthesis is performed based on the second sequence and the first voice color to obtain the first voice of the first text.
[0088] In practical applications, after expanding the first sequence to obtain the second sequence, speech synthesis can be performed on the first text based on the second sequence and the first voice color required for speech synthesis to obtain the first voice conforming to the first voice color. Among them, the first voice color (i.e., the preset voice color) is preset, which means a set of pre-defined and designed speech features. These features determine the properties of the synthesized speech such as sound quality, gender, age, accent, emotional state, etc. The preset voice color is usually created by the speech synthesis system provider or developer, and users can select or customize different voice colors according to their needs.
[0089] Preset timbres usually include the following aspects: Voice quality: Voice quality refers to the clarity, purity, and richness of the voice. Different voice qualities can bring different feelings to the audience. For example, clear voice quality is suitable for news broadcasts, while soft voice quality is suitable for storytelling. Gender: The gender attribute of the timbre determines whether the synthesized voice sounds male or female. This is usually achieved by adjusting the fundamental frequency and formant. Age: The age attribute of the timbre determines whether the synthesized voice sounds like a child, teenager, adult, or elderly person. This can be achieved by adjusting the pitch, formant, and pronunciation of the voice. Accent: The accent attribute of the timbre determines the regional or national characteristics of the synthesized voice. For example, you can choose different accents such as American English, British English, Australian English, etc. Emotional state: The emotional state attribute of the timbre determines the emotional color of the synthesized voice, such as happiness, sadness, anger, calmness, etc. This can be achieved by adjusting the tone, rhythm, and stress of the voice. Personality: The personality attributes of the timbre determine the personality of the synthesized speech, such as formal, casual, professional, friendly, etc. This can be achieved by adjusting the speech speed, pronunciation and intonation.
[0090] Preset timbres are typically created through a combination of audio recording and speech synthesis technology. First, one or more voice samples from a professional voice actor are recorded. Then, speech synthesis technology is used to extract speech features from these samples and create timbre models. These models can then be used by the speech synthesis system to generate speech that matches the preset timbres.
[0091] In practical applications, users can select or customize preset timbres to meet different needs, such as customer service, voice navigation, audiobooks, game dubbing, etc. The quality and diversity of preset timbres are crucial to improving the practicality and user experience of speech synthesis systems.
[0092] In some embodiments, see Figure 3B , Figure 3B This is a second flow chart of the speech synthesis method provided in an embodiment of the present application. Figure 3A Step 104 is shown as being able to be Figure 3B Steps 1041 to 1043 are shown to achieve:
[0093] In step 1041 , the second sequence is encoded to obtain semantic features of the second sequence, and the semantic features are decoded to obtain first acoustic features of the second sequence.
[0094] In some embodiments, the second sequence can be encoded to obtain semantic features in the following manner: word embedding is performed on each word in the second sequence to obtain a feature vector of each word; encoding is performed based on the position of each word in the second sequence and the feature vector of each word to obtain the encoding features of each word; mapping is performed on each encoding feature to obtain a mapping result of each word, and biasing is performed on the mapping result of each word to obtain the semantic features of the second sequence.
[0095] In practical applications, when encoding the second sequence, each word (i.e., token) in the second sequence is firstly subjected to word embedding processing to convert each word into a vector of a fixed size; then, encoding processing is performed according to the position of each word in the second sequence and the feature vector of each word to obtain the encoding features of each word; then, the encoding features of each word of the information can be linearly or nonlinearly mapped through multi-layer full connection to obtain the corresponding mapping results, such as fully connecting the encoding features of each word, transmitting them to the hidden layer through the input layer, obtaining the corresponding hidden layer features through the hidden layer, performing feature mapping on the hidden layer features, and obtaining the mapping features of each word as the mapping results; finally, the obtained mapping results are biased through an activation function (such as ReLu) to obtain the semantic features of the second sequence.
[0096] Through the above methods, vector conversion and encoding processing can capture the semantic relationship and contextual information between word units, thereby improving the semantic understanding ability of text; through mapping processing and bias processing, richer and more accurate semantic features can be generated, which can be used to generate more natural and fluent speech.
[0097] In some embodiments, encoding processing can be performed based on the position of each word in the second sequence and the feature vector of each word to obtain the encoding characteristics of each word: encoding processing is performed on the feature vector of each word according to the position of each word in the second sequence to obtain the position code of each word; and the feature vector of each word and the position code of each word are added together to obtain the encoding characteristics of each word.
[0098] Here, subsequent encoding processing is performed on each word in the second virtual sequence as a dimension, which can improve the ability to understand the semantics of the text. Taking into account the differences in the semantic information carried by words appearing at different positions in the second sequence, the feature vectors of words appearing at different positions in the second training are distinguished by adding corresponding position codes, which can better express the true meaning of the first text.
[0099] In some embodiments, the dimension of the position code of each word unit is the same as the dimension of the feature vector of each word unit; accordingly, the feature vector of each word unit can be encoded according to the position of each word unit in the second sequence to obtain the position code of each word unit: when the serial number of the dimension in the position code is an even number, the encoding value of the corresponding dimension in the position code is determined according to the sine function, wherein the sine function takes the position of the word unit in the second sequence and the dimension of the position code as parameters; when the serial number of the dimension in the position code is an odd number, the encoding value of the corresponding dimension in the position code is determined according to the cosine function, wherein the cosine function takes the position of the word unit in the second sequence and the dimension of the position code as parameters.
[0100] As an example, when the ordinal number of the dimension in the position code is an even number, the code value of the corresponding dimension in the position code is determined according to the following sine function (1):
[0101]
[0102] When the ordinal number of the dimension in the position code is an odd number, the code value of the corresponding dimension in the position code is determined according to the following cosine function (2):
[0103]
[0104] Where PE(i) is the encoding value of the i-th dimension in the position encoding, pos is the sorting position of the word in the second sequence, i is the serial number of each dimension in the position encoding, and i is an integer not less than 0, d model Dimension to encode the position.
[0105] Through the above method, the encoding method of trigonometric functions such as sine function or cosine function is used to determine the encoding value of the corresponding dimension in the position encoding, which can not only express the absolute position information of the word unit, but also express the relative position relationship of the word unit. Due to the formula characteristics of the trigonometric function, the position encoding of the next position can be used to represent the position encoding of the previous position, so the relative position relationship between word units can be learned. The sine function is used for encoding in the even-dimensional positions of the position encoding, and the cosine function is used for encoding in the odd-dimensional positions of the position encoding, so that the position encoding is easier to obtain time series information.
[0106] It should be noted that the above is only one implementation method of position coding. In practical applications, position coding is not limited to encoding using trigonometric functions, and this application does not limit the implementation method of position coding.
[0107] In some embodiments, after obtaining the feature vector of each word unit, iterative encoding processing can be performed based on the feature vector of each word unit to obtain the semantic features of the second sequence. Specifically: the input of the nth neural network model is encoded by the nth neural network model in the N cascaded neural network models, and the nth encoding processing result output by the nth neural network model is transmitted to the n+1th neural network model for further encoding; the Nth encoding processing result output by the Nth neural network model is used as the semantic feature corresponding to the second sequence.
[0108] Among them, n is an integer whose value increases from 1, and the value range of n satisfies 1≤n≤N-1, and N is an integer greater than or equal to 2; when n is 1, the input of the nth neural network model is the feature vector of each word unit, and when n is 2≤n≤N-1, the input of the nth neural network model is the encoding processing result of the n-1th neural network model; the cascade number N of the neural network model can be set. For example, if it is set to 3, the output result of the previous neural network model is the input of the next neural network model, and the output of the last neural network model is the encoding processing result of encoding the feature vector of each word unit, and the input of the first neural network model is the feature vector of each word unit.
[0109] In practical applications, each neural network model includes an attention layer, a first normalization layer, a forward transmission layer and a second normalization layer; the above-mentioned encoding processing of the input of the nth neural network model by the nth neural network model in the N cascaded neural network models can be achieved in the following manner: through the attention layer, the input of the nth neural network model is subjected to attention processing to obtain the attention feature corresponding to the input of the nth neural network model; through the first normalization layer, the attention feature and the input of the nth neural network model are subjected to residual connection processing and normalization processing to obtain the normalized feature corresponding to the input of the nth neural network model; through the forward transmission layer, the normalized feature is subjected to linear rectification processing to obtain the linear rectification processing result corresponding to the input of the nth neural network model; through the second normalization layer, the normalized feature of the input of the nth neural network model and the linear rectification processing result are subjected to residual connection processing and normalization processing to obtain the nth encoding processing result output by the nth neural network model.
[0110] As an example, the cascaded neural network model can be obtained by cascading N neural network models, each of which includes an attention layer, a first normalization layer, a forward transmission layer, and a second normalization layer. Next, taking the processing flow of the first neural network model as an example, the feature vector corresponding to each word unit is used as the input of the first neural network model. The self-attention mechanism of the attention layer is used to perform self-attention processing on the feature vectors of each word unit to obtain the attention features of each word. The self-attention mechanism can learn the dependency relationship between each word unit, thereby mining the important features in the first text for subsequent speech synthesis, thereby realizing accurate speech synthesis function.
[0111] Both the first and second normalization layers are used for residual connections and normalization. For example, the first normalization layer transposes the attention features of each word to obtain the transposed features of each word; the transposed features corresponding to each word are summed with the feature vector corresponding to each word to obtain the summed result of each word; and the summed result of each word is normalized to obtain the normalized features corresponding to each word. This is because the depth of the network can help the model extract richer, more abstract, and semantically informative features. Increasing depth cannot be achieved simply by increasing the number of layers. This will not only cause gradients to diffuse or explode, but more seriously, it will cause model degradation. Residual connections are used to address the degradation problem, preserving the original input of the previous layer (i.e., the attention features corresponding to each word) as much as possible. Normalization is the normalization of the feature vector corresponding to each word in this input. The normalization factor is the number of neurons in this layer. Normalization can improve the convergence speed of the model.
[0112] The forward transmission layer includes a two-layer deep neural network (DNN) structure and an activation function layer. The activation function layer performs linear rectification on the input of this layer. Linear rectification can be achieved through an activation function (such as ReLU). The stacking of multiple forward transmission layers can increase the accuracy of the characterization of each word. The linear rectification result obtained by the forward transmission layer is input into the second normalization layer. The second normalization layer performs residual connection and normalization on the linear rectification result corresponding to each word, obtaining the first encoding processing result output by the first neural network model, i.e., the encoding feature of each word. The encoding feature of each word is input into the subsequent cascaded neural network model until the Nth encoding processing result output by the last neural network model in the cascade is obtained as the semantic feature of the second sequence. That is, the feature vector of each word is converted into a context vector. The context vectors of all words constitute the semantic feature of the second sequence. The semantic feature contains the semantic and structural information of the second sequence and is used for subsequent decoding to obtain the first acoustic feature.
[0113] Using this approach, during the encoder stage, each word in the second sequence (including information about the part of speech, sentence composition, and word order structure of each word and its token) is first converted into a fixed-length feature vector using word embedding technology to capture the semantic and grammatical features of the word. To enable the model to understand the order of words in the first text, positional encodings are cleverly added to the feature vectors of each word, using sine and cosine functions to generate a unique encoding for each position. These encoded vectors are then processed through a multi-layer Transformer encoder. Each layer incorporates a self-attention mechanism and a feedforward neural network. The self-attention mechanism enables the model to focus on the relationship between different positions in the input second sequence, while the feedforward neural network applies an independent nonlinear transformation to the feature vector at each position. Residual connections and layer normalization are used to stabilize and optimize the output of each layer. After being refined through all encoder layers, a high-level semantic representation containing rich semantic and syntactic information is ultimately obtained, namely the semantic features of the second sequence, laying a solid foundation for the subsequent speech synthesis process.
[0114] In some embodiments, the decoding process of the semantic features is implemented through i decoding layers, where i is an integer greater than 1; accordingly, the semantic features can be decoded in the following manner to obtain the first acoustic feature: for the i-th decoding layer, the output of the i-1-th decoding layer is used as the query vector, and the semantic feature is used as the key vector and the value vector; based on the dot product operation of the query vector and the key vector, the weight between the query vector and the key vector is determined; based on the weight, the value vector is weighted summed, and the result of the weighted summation is linearly transformed to obtain the first acoustic feature.
[0115] In practical applications, after obtaining the semantic features of the second sequence output by the encoder stage, the semantic features of the second sequence are input into the decoding network of the decoder stage. The decoder network usually includes multiple decoding layers, that is, the semantic features are decoded by multiple decoding layers to obtain the first acoustic features. The input of the first decoding layer in the decoder network is usually the feature vector of the start mark of speech synthesis (such as " <sos>"The corresponding feature vector), the input of the current decoding layer is the output of the previous decoding layer. For the current decoding layer, the input of the current decoding layer can be converted into a query vector, and the output of the encoder (that is, the semantic features of the second sequence) is converted into a key vector and a value vector. For each query vector, the similarity score between it and all key vectors is calculated. This is usually achieved by dividing the dot product of the query vector and the key vector by a scaling factor. These similarity scores are normalized by the softmax function to obtain the attention weight.
[0116] Using the calculated attention weights, we perform a weighted summation of all the value vectors to obtain a weighted value vector. This weighted vector contains information related to the current query vector in the encoder output. This weighted value vector is transformed through a linear layer (fully connected layer) to obtain the final encoder-decoder attention output, i.e., the first acoustic feature.
[0117] That is, the task of the decoder stage is to convert the high-level semantic representation output by the encoder stage (i.e., the semantic features of the second sequence) into acoustic features of speech. This process involves the following key steps, each of which is closely linked to ensure that the final generated speech can accurately reflect the content and style of the input text (i.e., the first text): Initialize the decoder input: The decoder input usually starts with a special start token (such as " <sos>”), the start token indicates the start of speech synthesis, and this start token is converted into an embedding vector as the first input of the decoder. Self-attention mechanism: The decoder first applies the self-attention mechanism to process the current decoder input sequence, which allows the model to focus on the relationship between the generated acoustic features and ensure the coherence of the speech. Encoder-decoder attention: The decoder applies the encoder-decoder attention mechanism to combine the high-level semantic representation (i.e., the semantic features of the second sequence) with the current decoder state. This step ensures that the generated acoustic features can accurately reflect the semantic content of the input text. Feedforward neural network: After the attention mechanism, The decoder applies a feedforward neural network to further process the feature vector at each position. This step usually includes two linear layers and a nonlinear activation function, such as ReLU. Residual connection and layer normalization: Similar to the encoder, the decoder also uses residual connection and layer normalization to stabilize the training process, accelerate convergence, and improve model performance. Multi-layer stacking: The decoder is usually composed of multiple such decoder layers stacked together, and each layer further refines and optimizes the generation of acoustic features. Acoustic feature generation: After processing all decoder layers, the model generates a series of acoustic features, namely the first acoustic features, which describe the pitch, duration, spectrum and other characteristics of the speech.
[0118] Through this approach, first, the model gradually refines and enriches acoustic features through multiple decoding layers in the decoder stage, thereby generating more natural and realistic speech. Second, the self-attention mechanism enables the model to flexibly capture long-range dependencies, which is particularly important for processing complex text structures and semantic relationships. Furthermore, the multi-layer decoder enhances the model's expressive power, making it better suited to different speech synthesis tasks and requirements. Overall, the multi-layer design and self-attention mechanism in the decoder stage provide strong support for generating high-quality acoustic features, thereby improving the overall performance of speech synthesis.
[0119] In step 1042 , the first acoustic feature is adjusted based on the first timbre and the identification information of each word in the second sequence to obtain a second acoustic feature.
[0120] In some embodiments, the first acoustic feature can be adjusted based on the first timbre and the labels of each word in the second sequence to obtain the second acoustic feature in the following manner: feature extraction is performed on the first timbre to obtain the timbre feature of the first timbre; the rhythmic feature of the first text is determined based on the sentence components of each word label in the second sequence and the word order structure of the first text; the first acoustic feature is adjusted based on the timbre feature and the rhythmic feature to obtain the second acoustic feature.
[0121] During the speech synthesis process, in order to generate speech that conforms to the preset first timbre and the rhythm of the first text to be synthesized, it is necessary to adjust the first acoustic feature based on the timbre feature of the first timbre and the rhythm feature of the first text to obtain the second acoustic feature. Specifically, the first text to be synthesized is: "The quick brown fox jumps over the lazy dog.", the second sequence is: "quick / adjective / modifier / sequence-brown / adjective / modifier / sequence-fox / noun / subject / sequence-jump / verb / predicate / sequence-d / particle / perfective marker / sequence-lazy / adjective / modifier / sequence-dog / noun / object / sequence", and the preset first timbre is the timbre of star x as an example:
[0122] First, feature extraction is performed on the preset first timbre (such as the timbre of star x). This usually involves analyzing the voice sample of star x, extracting its unique sound quality, pitch, resonance peak and other features, and forming a timbre feature vector. The timbre feature vector reflects the personality and style of the first timbre (such as the timbre of star x).
[0123] In some embodiments, feature extraction of the first timbre can be performed in the following manner to obtain the timbre features of the first timbre: feature fusion of the pitch features, volume features, sound quality features and Mel spectrum features of the first timbre to obtain the first fusion features of the first timbre; quantization encoding of the first fusion features to obtain the second fusion features; and timbre feature prediction of the first timbre based on the second fusion features to obtain the timbre features of the first timbre.
[0124] In practical applications, when extracting features from a first timbre, it is first necessary to collect multiple speech samples corresponding to the first timbre. For example, if the first timbre is that of celebrity x, multiple speech samples of celebrity x need to be collected. These speech samples should cover celebrity x's different pronunciations, intonations, and emotional states to ensure that the extracted timbre features are representative. Next, pitch features, volume features, sound quality features, and mel-spectrogram features are extracted from these speech samples. Pitch features reflect the pitch variations of the speech samples, volume features reflect the loudness of the speech samples, sound quality features reflect the clarity and richness of the speech samples, and mel-spectrogram features reflect the frequency distribution of the speech samples. The extracted pitch features, volume features, sound quality features, and mel-spectrogram features are then fused to obtain the first fused features of the first timbre. This step can be implemented through various methods, such as simple concatenation, weighted summation, or feature learning using a deep learning model. The first fused feature may be a high-dimensional continuous feature vector. To reduce the feature dimension and improve the generalization and efficiency of the model, the first fused feature can be quantized and encoded to obtain the second fused feature. That is, the purpose of quantization encoding is to convert the continuous feature into a discrete representation, which helps to reduce the feature dimension, improve the robustness of the feature, and speed up subsequent processing. Quantization encoding can discretize the continuous feature space through clustering algorithms (such as K-means) or use deep learning models such as autoencoders for feature compression. Finally, the timbre characteristics of the first timbre are predicted based on the second fused feature to obtain the timbre characteristics of the first timbre. This step usually involves training a prediction model. This prediction model can be a regression model for predicting continuous timbre feature values, or a classification model for predicting discrete timbre feature categories. Through training, the prediction model can learn the mapping relationship between the second fused feature and the timbre characteristics of the first timbre, thereby predicting the timbre characteristics of the first timbre given the second fused feature.
[0125] Through the above process, the timbre features of the first timbre (ie, star x) can be obtained, and these timbre features can be used for speech synthesis so that the generated speech sounds like the voice of star x.
[0126] Next, the prosodic features of the first text are determined based on the sentence components of each word marker in the second sequence and the word order structure of the first text. In some embodiments, the prosodic features of the first text can be determined based on the sentence components of each word marker in the second sequence and the word order structure of the first text in the following manner: determining the stress features of the first text based on the sentence components of each word marker in the second sequence and the word order structure of the first text; determining the intonation features of the first text based on the sentence components of each word marker in the second sequence and the word order structure of the first text; determining the pause features of the first text based on the sentence components of each word marker in the second sequence and the word order structure of the first text; determining the rhythm features of the first text based on the sentence components of each word marker in the second sequence and the word order structure of the first text; and determining the prosodic features of the first text based on at least one of the stress features, intonation features, pause features, and rhythm features.
[0127] In practice, the main components of a sentence, such as the subject, predicate, and object, typically carry stress. Furthermore, emphasis on components, new information, or specific rhetorical effects can also lead to stress. Furthermore, different word order structures can affect the location of stress. For example, in a subject-verb-object sentence structure, the object may be stressed, while in an object-verb-subject sentence structure, the object may be at the beginning of the sentence and, therefore, may also be stressed. Therefore, when determining the stress features of the first text, the stress features can be determined based on the sentence components and position of the word within the sentence, according to predefined rules or through machine learning models.
[0128] The rise and fall of intonation is often related to the sentence type and information structure. For example, declarative sentences often end with a falling intonation, while interrogative sentences end with a rising or level intonation. Word order can also influence intonation patterns. For example, inverted sentences may require different intonation treatments to maintain semantic clarity. Therefore, when determining the intonation characteristics of the first text, it is possible to use intonation rule libraries or analyze large corpora to learn intonation patterns under different word order structures.
[0129] Since pauses typically occur at grammatical boundaries, such as at the end of a sentence or clause, or between parallel elements, and different word order structures may affect the location and length of pauses—for example, complex sentence structures may require more pauses to separate different units of information—when determining the pause characteristics of the first text, we can identify the structural boundaries of sentences based on the results of grammatical analysis and determine the location and length of pauses based on the word order structure.
[0130] Because rhythm is related to word importance, semantic cohesion, and speech rate, important words or closely related semantic units may have a faster rhythm. Different word order structures may also affect rhythmic patterns. For example, a compact word order may require a faster rhythm, while a loose word order may require a slower rhythm. Therefore, when determining the rhythmic characteristics of the first text, the speed and pattern of the rhythm can be determined based on the sentence components of the words and the word order structure of the first text, combined with semantic information.
[0131] After determining at least one of the stress feature, intonation feature, pause feature, and rhythm feature, the determined at least one of the stress feature, intonation feature, pause feature, and rhythm feature may be used as a prosodic feature of the first text.
[0132] By combining the results of grammatical analysis and word order structure analysis and applying knowledge of linguistics and phonetics, the prosodic features of the first text can be extracted, which can provide accurate prosodic features for subsequent speech synthesis and help improve the accuracy of speech synthesis.
[0133] In the above example, the second sequence provides each word's part of speech, sentence element, and order in the sentence. This information helps us determine the prosodic features of each word, such as stress, intonation, pauses, and rhythm. For example, "fast" and "brown," as adjectives and modifiers, might require unstressed pronunciation; "fox," as a subject, might require stress; "skip," as a predicate, might require a rising intonation; "le," as a perfect tense marker, might require a falling intonation; and "lazy" and "dog," as adjectives and objects, might require appropriate stress and rhythm.
[0134] Finally, the first acoustic feature is adjusted based on the timbre and prosodic features to obtain a second acoustic feature. This step involves combining the timbre and prosodic features to adjust the first acoustic feature to ensure that the generated speech matches the timbre of celebrity X and has appropriate prosodic features. For example, the fundamental frequency, formant, and energy in the acoustic feature can be adjusted to match the timbre of celebrity X, and the stress, intonation, and rhythm of the speech can be adjusted based on the prosodic features.
[0135] In step 1043 , speech synthesis is performed based on the second acoustic feature to obtain a first speech of the first text.
[0136] Here, after generating a second acoustic feature that matches the timbre characteristics of the first timbre and possesses appropriate prosodic features, speech synthesis can be performed based on the second acoustic feature to obtain the first speech corresponding to the first text. During speech synthesis, an appropriate vocoder technology can be selected based on actual needs, such as a vocoder based on a physical model (e.g., Mel-frequency cepstral coefficient synthesis) or a vocoder based on a statistical model (e.g., a Gaussian mixture model or a vocoder based on deep learning). After selecting a vocoder, the second acoustic feature can be appropriately preprocessed, including normalization and feature alignment, to meet the input requirements of the selected vocoder. The second acoustic feature is then input into the vocoder, which generates a corresponding speech waveform based on these second acoustic features. Finally, the generated speech waveform is post-processed, such as by removing noise, smoothing the waveform, and adjusting the volume, to improve speech quality and naturalness. The processed audio waveform can be output as a digital audio file in formats such as WAV or MP3, or played directly, thereby outputting the first speech corresponding to the first text.
[0137] Through the above process, the speech generated by the acoustic features adjusted based on the first timbre can maintain a high degree of consistency with the first timbre, and is adjusted based on the labels of each word in the second sequence, allowing fine adjustment of various aspects of the speech, including pitch, volume, speaking speed, etc., thereby achieving highly flexible and customizable speech output; that is, by adjusting the first acoustic features using the preset first timbre and second sequence, a speech that conforms to the preset timbre and text rhythm can be generated, and the synthesized speech not only sounds real, but also can convey appropriate tone and emotion.
[0138] Continuing with the above example, the final output speech sounds like star x naturally saying "The quick brown fox jumps over the lazy dog.", making the synthesized speech sound like star x's voice while accurately conveying the semantics and emotion of the first text.
[0139] The process of the above embodiment is a process of reasoning based on a trained speech synthesis model. Next, the training process of a large speech synthesis model is schematically illustrated.
[0140] See also Figure 4 , Figure 4 This is a data flow diagram for model training provided in an embodiment of the present application. First, for the speech sample corresponding to the text sample (i.e., the speech signal of the text sample), the speech sample is divided into speech units to obtain a fourth sequence, that is, the fourth sequence is a speech token sequence of each speech unit obtained by dividing the speech sample into units. For a text sample, the text sample is first segmented to obtain multiple words. Based on the parts of speech of the words, a third sequence is formed according to their order in the text sample, such as the third sequence is: [1st word / part of speech 1, 2nd word / part of speech 2, ..., Nth word / part of speech N]; then, based on the parts of speech of each word, the sentence components of each word in the text sample are determined, such as the sentence components of the 1st word, the 2nd word, ..., the Nth word are sentence component 1, sentence component 2, ..., sentence component N respectively; the word order structure of the text sample is determined according to the order of each word in the text sample, such as whether it is sequential (word order structure 1) or reverse order (word order structure 0). Assuming that the word order structure of the text sample is word order structure 1 (i.e. sequential), the word order structures corresponding to the 1st word, the 2nd word, ..., the Nth word are all word order structure 1. Then, based on the sentence components of each word in the text sample and the word order structure of the text sample, the third sequence is expanded to obtain a fifth sequence (i.e., a rich text token sequence). For example, the fifth sequence is: [1st word / part of speech 1 / sentence component 1 / word order structure 1, second word / part of speech 2 / sentence component 1 / word order structure 1, ..., Nth word / part of speech N / sentence component N / word order structure 1], so that the bit rate of the rich text tokens in the fifth sequence is equivalent to the bit rate of the speech tokens in the sixth sequence. After obtaining the fifth sequence of text samples, the fifth sequence can be concatenated with the fourth sequence of speech samples to obtain a sixth sequence. Based on the sixth sequence and the timbre of the speech sample (e.g., timbre token), a training sample is constructed as follows: [timbre token][rich text token sequence][SPEECH_BEGIN][speech token sequence][SPEECH_END]. Finally, a large language model for speech synthesis can be trained based on this sample sequence pattern.
[0141] Combine Figure 4 , see Figure 5 , Figure 5 This is a flow chart of the model training method provided in the embodiment of the present application, which will be combined with Figure 5 The steps shown are explained:
[0142] In step 201, the text sample is segmented to obtain a third sequence including multiple words obtained by segmenting the text sample, and the speech sample is divided into speech units to obtain a fourth sequence including multiple speech units obtained by dividing the speech sample.
[0143] In practical applications, speech units refer to basic speech segments that can be identified and analyzed within a speech signal. These segments are the fundamental building blocks of speech signals and typically correspond to specific speech features or articulatory units. In speech processing and speech recognition, speech unit decomposition involves breaking down a continuous speech signal (such as the speech sample described above) into a series of discrete units for further analysis and processing. Speech units can be of different types, including common phonemes, syllables, morphemes, phrases, and sentences. The specific type depends on the speech processing application and the purpose of the analysis.
[0144] Among them, phonemes are the smallest units in speech that can distinguish meanings. They are the basic phonetic elements that make up words. For example, the English word "cat" is composed of three phonemes: / k / , and / t / . A syllable is a speech unit consisting of one or more phonemes, usually containing a vowel nucleus and one or more consonants. For example, the English word "banana" consists of three syllables: A morpheme is the smallest semantically meaningful unit in a language. It can be a word or part of a word. For example, the English word "unhappiness" is composed of three morphemes: the prefix "un-," the root "happy," and the suffix "-ness." A sentence is a phonetic unit composed of multiple phrases, usually expressing a complete semantic unit. For example, in English, "The catis on the mat." is a sentence composed of multiple phrases that expresses a complete semantic meaning.
[0145] Here, for the speech sample corresponding to the text sample (i.e., the speech signal of the text sample), the speech sample is divided into speech units to obtain a fourth sequence. When performing speech unit division, a speech unit detection algorithm, such as detection based on an acoustic model (e.g., a hidden Markov model) or detection based on deep learning (e.g., a convolutional neural network or a recurrent neural network), can be used to identify the speech units in the speech sample. After the speech units in the speech sample are identified, the speech sample can be segmented, such as by segmenting the continuous speech signal into independent speech units based on the detected speech unit boundaries, and the segmented speech units are labeled to clarify the type (e.g., phoneme, syllable, word, etc.) and content of each speech unit to form a speech token sequence (i.e., the fourth sequence).
[0146] For text samples, the text samples are first segmented to obtain multiple words. Based on the parts of speech of the words, a third sequence is formed according to their sorting order in the text samples. For example, the third sequence is: [1st word / part of speech 1, 2nd word / part of speech 2, ..., Nth word / part of speech N].
[0147] In step 202, the sentence component of each word in the text sample is determined based on the part of speech of each word, and the word order structure of the text sample is determined according to the order of each word in the text sample.
[0148] Here, the sentence components of each word in the text sample are determined based on the part of speech of each word, such as the sentence components of the first word, the second word, ..., the Nth word are sentence component 1, sentence component 2, ..., sentence component N respectively; the word order structure of the text sample is determined according to the arrangement order of each word in the text sample, such as whether it is sequential (word order structure 1) or reverse order (word order structure 0). Assuming that the word order structure of the text sample is word order structure 1 (i.e. sequential), the word order structures corresponding to the first word, the second word, ..., the Nth word are all word order structure 1.
[0149] In step 203 , based on the sentence components of each word in the text sample and the word order structure of the text sample, the third sequence is expanded to obtain a fifth sequence, and the fourth sequence and the fifth sequence are concatenated to obtain a sixth sequence.
[0150] Here, based on the sentence components of each word in the text sample and the word order structure of the text sample, the third sequence is expanded to obtain a fifth sequence (i.e., a rich text token sequence). For example, the fifth sequence is: [1st word / part of speech 1 / sentence component 1 / word order structure 1, 2nd word / part of speech 2 / sentence component 1 / word order structure 1, ..., Nth word / part of speech N / sentence component N / word order structure 1], so that the bit rate of the rich text tokens in the fifth sequence is equivalent to the bit rate of the speech tokens in the sixth sequence. After obtaining the fifth sequence of the text sample, the fifth sequence can be spliced with the fourth sequence of the speech sample to obtain the sixth sequence.
[0151] In step 204, the model is trained based on the sixth sequence and the timbre of the speech sample.
[0152] Here, a training sample is constructed based on the sixth sequence and the timbre of the speech sample (e.g., a timbre token) as follows: [timbre token][rich text token sequence][SPEECH_BEGIN][speech token sequence][SPEECH_END]. This allows the training sample to be input into the speech synthesis model to train the speech synthesis language model. During training, the speech synthesis language model performs speech prediction on the training sample, and a loss function is constructed using the prediction results and the speech token sequence in the training sample. The parameters of the speech synthesis model are then updated based on the loss function.
[0153] Through the above method, since the bit rate of the rich text token in the fifth sequence is equivalent to the bit rate of the voice token in the fourth sequence corresponding to the voice sample, after splicing the fourth sequence and the fifth sequence to obtain the sixth sequence, the large speech synthesis model is trained based on the sixth sequence. This can effectively solve the problem of poor model training stability and poor training effect due to the significant gap between the bit rate of text tokens and the bit rate of voice tokens during the training of the large speech synthesis model.
[0154] The following continues to describe the exemplary structure of the speech synthesis device 455A provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2A As shown, the software modules stored in the speech synthesis device 455A of the memory 450 may include: a word segmentation module 4551A, used to perform word segmentation processing on the first text to obtain a first sequence, wherein the first sequence includes multiple words obtained by word segmentation of the first text; a determination module 4552A, used to determine the sentence components of each of the words in the first text based on the part of speech of each of the words, and determine the word order structure of the first text according to the order of each of the words in the first text; an expansion module 4553A, used to expand the first sequence based on the sentence components of each of the words in the first text and the word order structure of the first text to obtain a second sequence; a synthesis module 4554A, used to perform speech synthesis based on the second sequence and the first timbre to obtain the first speech of the first text.
[0155] In some embodiments, the expansion module 4553A is also used to mark each of the words in the first sequence based on the part of speech of each of the words, the sentence component of each of the words in the first text and the word order structure of the first text to obtain a second sequence; wherein, the second sequence includes a subsequence corresponding to each of the words, and the subsequence is composed of the following word elements: the words and the identification information of the words, and the identification information includes: the part of speech of the words, the sentence component of the words in the first text and the word order structure of the first text.
[0156] In some embodiments, the synthesis module 4554A is also used to encode the second sequence to obtain semantic features, and decode the semantic features to obtain first acoustic features; adjust the first acoustic features based on the first timbre and the identification information of each word in the second sequence to obtain second acoustic features; and perform speech synthesis based on the second acoustic features to obtain the first speech of the first text.
[0157] In some embodiments, the synthesis module 4554A is also used to perform word embedding processing on each word in the second sequence to obtain a feature vector of each word; perform encoding processing based on the position of each word in the second sequence and the feature vector of each word to obtain the encoding feature of each word; perform mapping processing on each encoding feature to obtain a mapping result of each word, and perform bias processing on the mapping result of each word to obtain the semantic feature of the second sequence.
[0158] In some embodiments, the synthesis module 4554A is further used to encode the feature vector of each word element according to the position of each word element in the second sequence to obtain the position code of each word element; and add the feature vector of each word element and the position code of each word element to obtain the coding feature of each word element.
[0159] In some embodiments, the dimension of the position code of each word unit is the same as the dimension of the feature vector of each word unit; the synthesis module 4554A is further used to determine the code value corresponding to the dimension in the position code according to the sine function when the serial number of the dimension in the position code is an even number, wherein the sine function takes the position of the word unit in the second sequence and the dimension of the position code as parameters; and to determine the code value corresponding to the dimension in the position code according to the cosine function when the serial number of the dimension in the position code is an odd number, wherein the cosine function takes the position of the word unit in the second sequence and the dimension of the position code as parameters.
[0160] In some embodiments, the decoding processing of the semantic features is implemented through i decoding layers, where i is an integer greater than 1. The synthesis module 4554A is also used to use the output of the i-1 decoding layer as the query vector for the i-th decoding layer, and the semantic features as the key vector and the value vector, wherein the input of the first decoding layer is the feature vector of the start marker of speech synthesis; based on the dot product operation of the query vector and the key vector, the weight between the query vector and the key vector is determined; based on the weight, the value vector is weighted summed, and the result of the weighted summation is linearly transformed to obtain the first acoustic feature.
[0161] In some embodiments, the synthesis module 4554A is also used to extract features of the first timbre to obtain timbre features of the first timbre; determine the rhythmic features of the first text based on the sentence components and word order structures of each of the word markers in the second sequence; and adjust the first acoustic features based on the timbre features and the rhythmic features to obtain second acoustic features.
[0162] In some embodiments, the synthesis module 4554A is also used to perform feature fusion on the pitch features, volume features, sound quality features and Mel-spectrogram features of the first timbre to obtain a first fusion feature of the first timbre; perform quantization encoding on the first fusion feature to obtain a second fusion feature; and perform timbre feature prediction on the first timbre based on the second fusion feature to obtain the timbre feature of the first timbre.
[0163] In some embodiments, the synthesis module 4554A is further used to determine the stress features of the first text based on the sentence components and word order structure of each of the word markers in the second sequence; determine the intonation features of the first text based on the sentence components of each of the word markers in the second sequence and the word order structure of the first text; determine the pause features of the first text based on the sentence components of each of the word markers in the second sequence and the word order structure of the first text; determine the rhythm features of the first text based on the sentence components of each of the word markers in the second sequence and the word order structure of the first text; and determine the prosodic features of the first text based on at least one of the stress features, the intonation features, the pause features and the rhythm features.
[0164] The following continues to describe the exemplary structure of the model training device 455B provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2B As shown, the software modules stored in the model training device 455B of the memory 450 may include: a sequence extraction module 4551B, which is used to perform word segmentation on the text sample to obtain a third sequence, wherein the third sequence includes multiple words obtained by word segmentation of the text sample, and to divide the speech sample into speech units to obtain a fourth sequence, wherein the fourth sequence includes multiple speech units obtained by dividing the speech sample; a text analysis module 4552B, which is used to determine the sentence components of each of the words in the text sample based on the part of speech of each of the words, and to determine the word order structure of the text sample according to the order of each of the words in the text sample; a sequence processing module 4553B, which is used to expand the third sequence to obtain a fifth sequence based on the sentence components of each of the words in the text sample and the word order structure of the text sample, and to splice the fourth sequence and the fifth sequence to obtain a sixth sequence; a model training module 4554B, which is used to train the model based on the sixth sequence and the timbre of the speech sample.
[0165] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the speech synthesis method described in the present invention.
[0166] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the speech synthesis method provided by the embodiment of the present application, for example, Figure 3A The speech synthesis method is shown.
[0167] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0168] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0169] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0170] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0171] In summary, through the embodiments of the present application, when performing speech synthesis on a first text to be synthesized, the first text is segmented to obtain a first sequence consisting of multiple words obtained by segmenting the first text; the sentence component of each word in the first text is determined based on the part of speech of each word, and the word order structure of the first text is determined according to the order of each word in the first text; based on the sentence component of each word in the first text and the word order structure of the first text, the first sequence is expanded to obtain a second sequence, and speech synthesis is performed based on the second sequence and the first timbre to obtain the first speech of the first text; in this way, since determining the sentence component of each word in the first text helps to understand the importance and function of each word in the first text, and determining the word order structure of the first text helps to capture the semantics and emotional expression of the first text, the expansion of the first sequence based on the sentence component and word order structure can enrich the representation of the text. Therefore, when performing speech synthesis on the first text based on the second sequence and the preset first timbre, the rich text representation in the second sequence and the timbre characteristics of the first timbre can be used to generate a speech that matches the text content and timbre requirements, thereby improving the accuracy, naturalness and personalization of the speech synthesis.
[0172] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.< / sos> < / sos>
Claims
1. A speech synthesis method, characterized in that: The method comprises: Performing word segmentation on the first text to obtain a first sequence, where the first sequence includes a plurality of words obtained by word segmentation of the first text; Determining the sentence component of each of the words in the first text based on the part of speech of each of the words, and determining the word order structure of the first text according to the order of each of the words in the first text; Expanding the first sequence based on the sentence components of each of the words in the first text and the word order structure of the first text to obtain a second sequence; Speech synthesis is performed based on the second sequence and the first timbre to obtain a first speech of the first text.
2. The method according to claim 1, characterized in that The step of expanding the first sequence based on the sentence components of each word in the first text and the word order structure of the first text to obtain a second sequence includes: Marking each of the words in the first sequence based on the part of speech of each word, the sentence component of each word in the first text, and the word order structure of the first text to obtain a second sequence; The second sequence includes a subsequence corresponding to each of the words, and the subsequence is composed of the following word elements: the word and the identification information of the word, and the identification information includes: the part of speech of the word, the sentence component of the word in the first text, and the word order structure of the first text.
3. The method according to claim 1, characterized in that The performing speech synthesis based on the second sequence and the first timbre to obtain the first speech of the first text includes: performing encoding processing on the second sequence to obtain semantic features of the second sequence, and performing decoding processing on the semantic features to obtain first acoustic features of the second sequence; adjusting the first acoustic feature based on the first timbre and identification information of each of the words in the second sequence to obtain a second acoustic feature; Speech synthesis is performed based on the second acoustic feature to obtain a first speech of the first text.
4. The method according to claim 3, characterized in that The encoding process of the second sequence to obtain semantic features includes: Performing word embedding processing on each word in the second sequence to obtain a feature vector of each word; Performing encoding processing based on the position of each word in the second sequence and the feature vector of each word to obtain encoding features of each word; Mapping processing is performed on each of the encoding features to obtain a mapping result of each of the word units, and bias processing is performed on the mapping result of each of the word units to obtain a semantic feature of the second sequence.
5. The method according to claim 4, characterized in that The encoding process based on the position of each word in the second sequence and the feature vector of each word to obtain the encoding feature of each word includes: encoding the feature vectors of the word units according to the positions of the word units in the second sequence to obtain position codes of the word units; The feature vector of each word unit and the position code of each word unit are added together to obtain the coding feature of each word unit.
6. The method according to claim 3, characterized in that The decoding process of the semantic features is implemented through i decoding layers, where i is an integer greater than 1. The decoding process of the semantic feature to obtain the first acoustic feature includes: For the i-th decoding layer, the output of the i-1-th decoding layer is used as the query vector, and the semantic features are used as the key vector and the value vector, wherein the input of the first decoding layer is the feature vector of the start marker of speech synthesis; determining a weight between the query vector and the key vector based on a dot product operation of the query vector and the key vector; The value vectors are weightedly summed based on the weights, and a linear transformation is performed on the result of the weighted summation to obtain the first acoustic feature.
7. The method according to claim 3, characterized in that The adjusting the first acoustic feature based on the first timbre and the identification information of each word in the second sequence to obtain the second acoustic feature includes: Extracting features of the first timbre to obtain timbre features of the first timbre; determining a prosodic feature of the first text based on the sentence components of each of the words in the second sequence and the word order structure of the first text; Based on the timbre feature and the rhythm feature, the first acoustic feature is adjusted to obtain a second acoustic feature.
8. The method according to claim 7, characterized in that The extracting features of the first timbre to obtain the timbre features of the first timbre includes: Performing feature fusion on the pitch feature, volume feature, sound quality feature, and mel spectrum feature of the first timbre to obtain a first fused feature of the first timbre; quantize and encode the first fused features to obtain second fused features; The timbre feature of the first timbre is predicted based on the second fusion feature to obtain the timbre feature of the first timbre.
9. The method according to claim 7, characterized in that The determining of the prosodic features of the first text based on the sentence components of each of the word markers in the second sequence and the word order structure of the first text includes: determining a stress feature of the first text based on the sentence components of each of the word markers in the second sequence and the word order structure of the first text; determining the intonation features of the first text based on the sentence components of each of the word markers in the second sequence and the word order structure of the first text; determining pause features of the first text based on the sentence components of each of the word markers in the second sequence and the word order structure of the first text; determining a rhythmic feature of the first text based on the sentence components of each of the word markers in the second sequence and the word order structure of the first text; A prosodic feature of the first text is determined based on at least one of the stress feature, the intonation feature, the pause feature, and the rhythm feature.
10. A model training method, characterized in that: The method comprises: Performing word segmentation on the text sample to obtain a third sequence, wherein the third sequence includes a plurality of words obtained by word segmentation of the text sample; Dividing the speech sample into speech units to obtain a fourth sequence, wherein the fourth sequence includes a plurality of speech units obtained by dividing the speech sample; Determining the sentence component of each of the words in the text sample based on the part of speech of each of the words, and determining the word order structure of the text sample according to the arrangement order of each of the words in the text sample; Based on the sentence components of each of the words in the text sample and the word order structure of the text sample, the third sequence is expanded to obtain a fifth sequence, and the fourth sequence and the fifth sequence are concatenated to obtain a sixth sequence; The model is trained based on the sixth sequence and the timbre of the speech sample.
11. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; A processor, configured to implement the method according to any one of claims 1 to 10 when executing computer-executable instructions or computer programs stored in the memory.
12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the method according to any one of claims 1 to 10 is implemented.
13. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the method according to any one of claims 1 to 10 is implemented.