Speech synthesis method and apparatus, electronic device, and computer readable medium

By using base language and target language models combined with timbre conversion technology in the speech synthesis system, the problem of large speech differences in multilingual communication in existing systems has been solved, and the consistency of timbre and the auditory effect of multilingual speech synthesis have been improved.

CN116229935BActive Publication Date: 2026-03-31VOICEAI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing speech synthesis systems are mainly developed for a single language, making it difficult to address the significant differences in synthesized speech between different languages ​​in bilingual or even multilingual communication scenarios. This is especially true in business settings, where it is difficult to find recording personnel proficient in multiple languages, resulting in unsatisfactory speech conversion results.

Method used

By using pre-acquired base language and target language speech synthesis models, a first synthesized speech and a second synthesized speech are obtained. Based on the training speech in the base language, a timbre conversion is performed to generate a third synthesized speech. Finally, the first and third synthesized speech are combined to generate the target synthesized speech, thereby improving the similarity and timbre consistency of synthesized speech in different languages.

Benefits of technology

It achieves timbre consistency in synthesized speech in bilingual or even multilingual communication scenarios, improves the auditory effect of multilingual speech synthesis, and ensures the naturalness and consistency of speech conversion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229935B_ABST
    Figure CN116229935B_ABST
Patent Text Reader

Abstract

The application discloses a speech synthesis method and device, electronic equipment and a computer readable medium, and relates to the technical field of speech synthesis. The method comprises the following steps: based on input text, obtaining a first synthesized speech according to a pre-acquired basic language speech synthesis model, and obtaining a second synthesized speech according to a pre-acquired target language speech synthesis model, wherein the similarity of the training speech of the target language speech synthesis model to the training speech of the basic language speech synthesis model is higher than a preset value; performing speech conversion on the second synthesized speech based on pre-acquired basic language training speech to obtain third synthesized speech; and obtaining target synthesized speech based on the first synthesized speech and the third synthesized speech. Therefore, the similarity of synthesized speech of different languages is further improved, and the target synthesized speech including bilingual or even multilingual speech has high timbre consistency, thereby improving the hearing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and more specifically, to a speech synthesis method, apparatus, electronic device, and computer-readable medium. Background Technology

[0002] Speech synthesis is a crucial aspect of human-computer interaction, and most speech synthesis systems are developed for a single language. However, in real-life situations, especially in business settings, bilingual or even multilingual communication is common. When a speech synthesis system designed for a single language is applied to bilingual or multilingual communication scenarios, significant differences in the synthesized speech between different languages ​​can easily occur. Summary of the Invention

[0003] This application proposes a speech synthesis method, apparatus, electronic device, and computer-readable medium to improve upon the aforementioned deficiencies.

[0004] In a first aspect, embodiments of this application provide a speech synthesis method, the method comprising: obtaining a first synthesized speech based on input text and a pre-acquired basic language speech synthesis model; obtaining a second synthesized speech based on input text and a pre-acquired target language speech synthesis model, wherein the similarity between the training speech of the target language speech synthesis model and the training speech of the basic language speech synthesis model is higher than a preset value; performing speech conversion on the second synthesized speech based on the pre-acquired basic language training speech to obtain a third synthesized speech, wherein the similarity between the third synthesized speech and the basic language training speech is higher than the similarity between the second synthesized speech and the basic language training speech; and obtaining a target synthesized speech based on the first synthesized speech and the third synthesized speech.

[0005] Secondly, embodiments of this application also provide a voiceprint recognition device, the device comprising: a synthesized speech acquisition unit, a speech conversion unit, and a speech synthesis unit. The synthesized speech acquisition unit is used to acquire a first synthesized speech based on input text and a pre-acquired basic language speech synthesis model, and to acquire a second synthesized speech based on a pre-acquired target language speech synthesis model; the speech conversion unit is used to perform speech conversion on the timbre of the second synthesized speech based on pre-acquired basic language training speech, to acquire a third synthesized speech; the speech synthesis unit is used to acquire a target synthesized speech based on the first synthesized speech and the third synthesized speech.

[0006] Thirdly, embodiments of this application also provide an electronic device, including: one or more processors; a memory; one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to perform the methods described above.

[0007] Fourthly, embodiments of this application also provide a computer-readable medium storing processor-executable program code, which, when executed by the processor, causes the processor to perform the above-described method.

[0008] The speech synthesis method, apparatus, electronic device, and computer-readable medium provided in this application include: obtaining a first synthesized speech based on input text and a pre-acquired basic language speech synthesis model; obtaining a second synthesized speech based on input text and a pre-acquired target language speech synthesis model, wherein the similarity between the training speech of the target language speech synthesis model and the training speech of the basic language speech synthesis model is higher than a preset value; then, performing speech conversion on the timbre of the second synthesized speech based on the pre-acquired basic language training speech to obtain a third synthesized speech; and finally, obtaining a target synthesized speech based on the first synthesized speech and the third synthesized speech. Therefore, when the input text is in the target language, this method performs speech synthesis on the input text based on a pre-acquired target language synthesis model to obtain a second synthesized speech. When the input text is in the base language, it performs speech synthesis on the input text based on a pre-acquired base language synthesis model to obtain a first synthesized speech. Since the similarity between the training speech of the target language speech synthesis model and the training speech of the base language speech synthesis model is higher than a preset value, the second synthesized speech has a high similarity with the first synthesized speech in terms of acoustic features. Then, the second synthesized speech is subjected to timbre-based speech conversion based on the pre-acquired base language training speech to obtain a third synthesized speech. Then, based on the first synthesized speech and the third synthesized speech, the target synthesized speech is obtained, further improving the similarity between synthesized speech in different languages. As a result, synthesized speech, including bilingual and even multilingual speech, has a high degree of timbre consistency, improving the auditory effect of multilingual speech synthesis.

[0009] Other features and advantages of the embodiments of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the embodiments of this application. The objects and other advantages of the embodiments of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart of a speech synthesis method provided in an embodiment of this application is shown.

[0012] Figure 2 A flowchart of a speech synthesis method provided in another embodiment of this application is shown.

[0013] Figure 3 A flowchart illustrating the method for obtaining training speech in a target language provided in an embodiment of this application is shown.

[0014] Figure 4 A flowchart illustrating the method for obtaining training speech in a target language provided in an embodiment of this application is shown.

[0015] Figure 5 A flowchart of a speech synthesis method provided in another embodiment of this application is shown.

[0016] Figure 6 A flowchart of a speech synthesis method provided in another embodiment of this application is shown.

[0017] Figure 7 A block diagram of a speech synthesis apparatus provided in one embodiment of this application is shown.

[0018] Figure 8 A schematic diagram of an electronic device provided according to an embodiment of this application is shown.

[0019] Figure 9 A schematic diagram of a storage medium according to an embodiment of this application is shown. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. The components of the embodiments of the present application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without inventive effort are within the scope of protection of the present application.

[0021] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0022] Speech synthesis, also known as text-to-speech technology, involves multiple disciplines such as acoustics, linguistics, digital signal processing, and computer science. It can convert any text into highly natural speech at any time using a computer. The main problem it solves is how to convert text information into audible sound information.

[0023] The theoretical basis of speech synthesis is the mathematical model of speech generation. Early speech synthesis mainly used parametric synthesis techniques based on formant models. Specifically, formants refer to the poles in the frequency response of the vocal tract. The distribution characteristics of the formant frequencies of speech determine the timbre of the speech. Therefore, based on different formant patterns of speech with different timbres, the formant frequencies and bandwidth are used as parameters for further processing to obtain synthesized speech. With the development of technology, current speech synthesis mainly uses waveform splicing methods. Unlike parametric synthesis techniques based on formant models, waveform splicing synthesis is based on splicing the waveforms of recorded synthesis primitives, which improves the naturalness of the synthesized speech. Furthermore, waveform synthesis techniques include linear predictive coding (LPC) and pitch-synchronous overlay (PSOLA), etc.

[0024] However, the inventors discovered during application that most speech synthesis systems are developed for a single language. In real life, especially in business settings, bilingual or even multilingual communication is common. In such cases, the inventors hoped that synthesized speech containing different languages ​​could possess similar timbre, making it less jarring for the listener. Existing speech synthesis systems use deep learning to train speech synthesis models based on extensive voice data from recording personnel. However, in bilingual or multilingual situations, it is difficult to find recording personnel fluent in multiple languages; in other words, it is difficult to obtain a multilingual speech synthesis model based on a single recording person.

[0025] In existing technologies, bilingual speech synthesis typically involves one language that the recording engineer is proficient in, such as Chinese. For this language, a large amount of speech data from the recording engineer is available. Furthermore, if the recording engineer, who is proficient in another language, acquires speech in that language, they can use voice conversion to perform timbre transformation on the acquired speech, thereby obtaining English speech similar to the speaking style of the Chinese recording engineer. However, the result of voice conversion technology is affected by the timbre of the speech being converted. When the timbre differences between the two recording engineers are significant, the speech conversion result is usually unsatisfactory.

[0026] Therefore, to overcome the above-mentioned defects, embodiments of this application provide a speech synthesis method, apparatus, electronic device, and computer-readable medium. When the input text is in a target language, speech synthesis is performed on the input text based on a pre-acquired target language synthesis model to obtain a second synthesized speech. When the input text is in a base language, speech synthesis is performed on the input text based on a pre-acquired base language synthesis model to obtain a first synthesized speech. Since the similarity between the training speech of the target language speech synthesis model and the training speech of the base language speech synthesis model is higher than a preset value, the second synthesized speech has a high similarity with the first synthesized speech at the acoustic feature level. Then, the second synthesized speech is subjected to timbre-based speech conversion based on the pre-acquired base language training speech to obtain a third synthesized speech. Then, based on the first synthesized speech and the third synthesized speech, a target synthesized speech is obtained, further improving the similarity of synthesized speech in different languages. Thus, synthesized speech, including bilingual or even multilingual speech, has a high degree of timbre consistency, improving the auditory effect of multilingual speech synthesis.

[0027] Please see Figure 1 , Figure 1 An embodiment of this application provides a speech synthesis method, specifically comprising S101 to S104.

[0028] S101: Based on the input text, obtain the first synthesized speech according to the pre-acquired basic language speech synthesis model.

[0029] In one implementation, the input text can be text including a base language and a target language. For example, the base language can be Chinese, and the target language can be English. The input text is based on the base language text within the input text. In another implementation, the base language speech synthesis model is a pre-built model using speech data from training speakers whose native language is the base language. Specifically, the model can be a Gaussian Mixture Model (GMM) or a Hidden Markov Model (HMM), etc. In yet another implementation, the method for obtaining the first synthesized speech can be to generate synthesized speech based on speech parameters in the synthesis model, wherein the speech parameters include fundamental spectrum parameters and spectral parameters.

[0030] S102: Based on the input text, obtain the second synthesized speech according to the pre-acquired target language speech synthesis model, wherein the similarity between the training speech of the target language speech synthesis model and the training speech of the basic language speech synthesis model is higher than a preset value.

[0031] In one implementation, the target language can be a variety of languages ​​different from the base language. That is, the input text can also be multilingual text, including the base language and multiple languages. The input text is based on the target language text within the input text. In another implementation, the target language speech synthesis model is a pre-built model using speech data from training participants whose native language is the target language. Furthermore, the training speech for the target language speech synthesis model is the speech data from training participants whose native language is the target language, and the training speech for the base language speech synthesis model is the speech data from training participants whose native language is the base language. Further, the similarity between the training speech for the target language speech synthesis model and the training speech for the base language speech synthesis model is higher than a preset value. This means that the similarity between the two speech data is higher than a preset value. Further, this means that the similarity of the acoustic features of the two speech data is higher than a preset value. The preset value can be pre-set by the model trainer according to requirements. In another implementation, the method for obtaining the second synthesized speech can refer to the above embodiments.

[0032] S103: Based on the pre-acquired training speech in the basic language, the second synthesized speech is converted into speech to obtain a third synthesized speech, wherein the similarity between the third synthesized speech and the training speech in the basic language is higher than the similarity between the second synthesized speech and the training speech in the basic language.

[0033] Specifically, the training speech in the basic language is the speech data of a training speaker whose native language is the basic language. The similarity includes the similarity of speech features such as timbre, speaking style, pronunciation style, and pauses. Furthermore, the fact that the third synthesized speech has a higher similarity to the training speech in the basic language than the second synthesized speech can mean that the comparative similarity of the third synthesized speech and the training speech in the basic language on the aforementioned acoustic features is higher than the comparative similarity of the second synthesized speech and the training speech in the basic language on the aforementioned acoustic features.

[0034] As one implementation method, the method for obtaining the third synthesized speech can be to adjust the timbre, speaking style, pronunciation style, and pause patterns of the second synthesized speech based on pre-acquired training speech in a basic language, thereby converting the second synthesized speech into a third synthesized speech that is closer to the training speech in the basic language. It is understood that the resulting third synthesized speech is closer to the training speech in the basic language in terms of timbre, speaking style, pronunciation style, pause patterns, and other speech features.

[0035] S104: Based on the first synthesized speech and the third synthesized speech, obtain the target synthesized speech.

[0036] As one implementation method, the method of obtaining the target synthesized speech can be to combine the first synthesized speech and the third synthesized speech in the text order based on the arrangement order of the base language text and the target language text in the input text to obtain the target synthesized speech. Specifically, the splicing units of different language speech can be calculated and the optimal splicing point can be calculated based on the recognized strings of different language texts, and the first synthesized speech and the third synthesized speech can be synthesized based on the optimal splicing point.

[0037] Therefore, the speech synthesis method provided in this application embodiment obtains a first synthesized speech based on the input text and a pre-acquired basic language speech synthesis model; obtains a second synthesized speech based on the input text and a pre-acquired target language speech synthesis model, wherein the similarity between the training speech of the target language speech synthesis model and the training speech of the basic language speech synthesis model is higher than a preset value; then, based on the pre-acquired basic language training speech, the timbre of the second synthesized speech is converted to obtain a third synthesized speech; and finally, based on the first synthesized speech and the third synthesized speech, a target synthesized speech is obtained. Therefore, when the input text is in the target language, this method performs speech synthesis on the input text based on a pre-acquired target language synthesis model to obtain a second synthesized speech. When the input text is in the base language, it performs speech synthesis on the input text based on a pre-acquired base language synthesis model to obtain a first synthesized speech. Since the similarity between the training speech of the target language speech synthesis model and the training speech of the base language speech synthesis model is higher than a preset value, the second synthesized speech has a high similarity with the first synthesized speech in terms of acoustic features. Then, the second synthesized speech is subjected to timbre-based speech conversion based on the pre-acquired base language training speech to obtain a third synthesized speech. Then, based on the first synthesized speech and the third synthesized speech, the target synthesized speech is obtained, further improving the similarity between synthesized speech in different languages. As a result, synthesized speech, including bilingual and even multilingual speech, has a high degree of timbre consistency, improving the auditory effect of multilingual speech synthesis.

[0038] Please see Figure 2 , Figure 2 An embodiment of this application provides a speech synthesis method, specifically comprising steps S201 to S208.

[0039] S201: Obtain training audio in the basic language and audio in multiple target languages.

[0040] In one implementation, the training audio for the basic language is collected from a recorder whose native language is the basic language type, and the audio for the multiple target languages ​​is collected from multiple recorders whose native language is the target language type.

[0041] S202: Based on the training speech of the basic language, obtain the speech synthesis model of the basic language.

[0042] As one implementation method, speech can be trained based on the basic language to obtain the corresponding fundamental frequency parameters and spectral parameters, establish a fundamental frequency parameter model and a spectral parameter model, and then combine the fundamental frequency parameter model and the spectral parameter model to obtain the speech synthesis model of the basic language.

[0043] S203: Based on the training speech in the basic language, select a training speech in the target language from among the multiple training speech in the target language, wherein the training speech in the target language is the training speech in the target language with a similarity to the training speech in the basic language that is higher than a preset value.

[0044] As one implementation method, the acquisition of the target language training speech can be achieved by selecting target language speech from the plurality of target language speech that has a similarity higher than a preset value with the base language training speech as the target language training speech. Specifically, the similarity can refer to the similarity of acoustic features between the target language training speech and the base language training speech, and the preset value can be set in advance by the model trainer according to the needs.

[0045] S204: Based on the target language training speech, obtain the target language speech synthesis model.

[0046] S205: Based on the input text, obtain the first synthesized speech according to the pre-acquired basic language speech synthesis model.

[0047] S206: Based on the input text, obtain the second synthesized speech according to the pre-acquired target language speech synthesis model, wherein the similarity between the training speech of the target language speech synthesis model and the training speech of the basic language speech synthesis model is higher than a preset value.

[0048] S207: Based on the pre-acquired training speech in the basic language, the second synthesized speech is converted into speech to obtain a third synthesized speech, wherein the similarity between the third synthesized speech and the training speech in the basic language is higher than the similarity between the second synthesized speech and the training speech in the basic language.

[0049] S208: Based on the first synthesized speech and the third synthesized speech, obtain the target synthesized speech.

[0050] The implementation methods of steps S204 to S208 can be referred to the foregoing embodiments, and will not be repeated here.

[0051] As one implementation method, please refer to Figure 3 , Figure 3This application provides an embodiment of a method for selecting and obtaining training speech in a target language in step S203. Specifically, the method may include steps S301 to S302.

[0052] S301: Compare the multiple target language speech samples with the basic language training speech samples one by one to obtain the similarity between each target language speech sample and the basic language training speech sample.

[0053] S302: Select the target language speech with a similarity higher than a preset value as the target language training speech.

[0054] As one implementation method, please refer to Figure 4 , Figure 4 This application provides an embodiment of a method for selecting and obtaining training speech in a target language in step S203. Specifically, the method may include steps S401 to S406.

[0055] S401: Extract the voiceprint features of each of the target language speech words as the comparison voiceprint features;

[0056] S402: Extract the voiceprint features of the training speech in the basic language as the basic voiceprint features.

[0057] S403: Perform voiceprint comparison on multiple of the comparison voiceprint features and the basic voiceprint features to obtain a voiceprint comparison score;

[0058] S404: Based on the voiceprint comparison score, obtain the similarity between each target language speech and the basic language training speech.

[0059] As one implementation method, the compared voiceprint features can be fed into a pre-obtained voiceprint feature model for scoring and judgment to obtain a judgment score. The voiceprint feature model can be a random model, which uses a probability density function to simulate the user. The training process involves inputting multiple voice segments provided by the user into this probability density function to predict the function's parameters, thereby obtaining the user's personalized voiceprint feature model. Further, the random model can be a Gaussian Mixture Model (GMM) or a Hidden Markov Model (HMM), etc. Furthermore, the voiceprint feature model uses basic voiceprint features to simulate a recorder in the basic language. It is understood that the target language training speech obtained in this way is closer to the basic language training speech in terms of timbre.

[0060] S405: Select the target language speech with a similarity higher than a preset value as the target language training speech.

[0061] The implementation method of step S405 can be referred to the foregoing embodiments, and will not be repeated here.

[0062] Please see Figure 5 , Figure 5 The illustration shows a speech synthesis method provided by an embodiment of this application. Specifically, the method includes: S501 to S506.

[0063] S501: Based on the input text, obtain the first synthesized speech according to the pre-acquired basic language speech synthesis model.

[0064] S502: Based on the input text, obtain the second synthesized speech according to the pre-acquired target language speech synthesis model, wherein the similarity between the training speech of the target language speech synthesis model and the training speech of the basic language speech synthesis model is higher than a preset value.

[0065] S503: Obtain a speech conversion model based on pre-acquired training speech in the basic language;

[0066] S504: Input the second synthesized speech into the speech conversion model to obtain the third converted speech.

[0067] As one implementation method, the speech conversion model is obtained through pre-training based on the audio data of the training speech in the basic language. The training method can adopt the methods in related technologies, and this embodiment does not limit it.

[0068] In one implementation, the speech conversion model includes a pre-trained text encoder, which can generate text encoding vectors from audio data and extract text information from the audio data. When the second synthesized speech is input into the speech conversion model, the pre-trained text encoder generates a target language text encoding vector, and then the speech conversion model uses the converted target language text encoding vector to further generate the converted third synthesized speech.

[0069] S505: Based on the first synthesized speech and the third synthesized speech, obtain the target synthesized speech.

[0070] Please see Figure 6 , Figure 6 An embodiment of this application provides a speech synthesis method, specifically comprising: S601 to S604.

[0071] S601: Determine the language type of the input text.

[0072] As one implementation method, the method for determining the language type of the input text can be to perform language identification on each character in the text to be processed. Specifically, a text language identification method based on Unicode can be used. The type of a character is determined by judging the range of Unicode codes for different languages.

[0073] Specifically, for example: the following are the Unicode encoding ranges for Chinese characters, numbers, uppercase and lowercase letters, and common punctuation marks:

[0074] Basic Chinese characters: [0x4e00,0x9fa5] (or decimal [19968,40869]);

[0075] Number: [0x 0030, 0x0039] (or decimal [48, 57]);

[0076] Lowercase letters: [0x0061,0x007a] (or decimal [97,122]);

[0077] Uppercase letters: [0x0041,0x005a] (or decimal [65,90]);

[0078] Commonly used punctuation marks: 2000-206F.

[0079] S602: If the input text is a basic language text, input the basic language text into the basic language speech synthesis model to obtain the first synthesized speech.

[0080] S603: If the input text is in the target language, input the target language text into the target language speech synthesis model to obtain the second synthesized speech.

[0081] S604: Based on the pre-acquired training speech in the basic language, perform speech conversion on the second synthesized speech to obtain a third synthesized speech, wherein the similarity between the third synthesized speech and the training speech in the basic language is higher than the similarity between the second synthesized speech and the training speech in the basic language.

[0082] S605: Based on the first synthesized speech and the third synthesized speech, obtain the target synthesized speech.

[0083] The implementation methods of steps S602 and S605 can be referred to the foregoing embodiments, and will not be repeated here.

[0084] Please see Figure 7 The diagram shows a structural block of a speech synthesis device 700 provided in an embodiment of this application. The device may include a synthesized speech acquisition unit 701, a speech conversion unit 702, and a speech synthesis unit 703.

[0085] The synthesized speech acquisition unit 701 is used to acquire first synthesized speech based on input text and a pre-acquired basic language speech synthesis model, and to acquire second synthesized speech based on a pre-acquired target language speech synthesis model.

[0086] The speech conversion unit 702 is used to convert the timbre of the second synthesized speech based on the pre-acquired basic language training speech to obtain the third synthesized speech;

[0087] The speech synthesis unit 703 is used to obtain target synthesized speech based on the first synthesized speech and the third synthesized speech.

[0088] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0089] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.

[0090] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0091] Please refer to Figure 8 This diagram illustrates a structural block diagram of an electronic device provided in an embodiment of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The electronic device 800 in this application may include one or more components: a processor 810, a memory 820, and one or more application programs, wherein the one or more application programs may be stored in the memory 820 and configured to be executed by the one or more processors 810, and the one or more programs are configured to perform the methods described in the foregoing method embodiments.

[0092] Processor 810 may include one or more processing cores. Processor 810 connects to various parts within the wearable device 800 using various interfaces and lines, and performs various functions and processes data of the wearable device 800 by running or executing instructions, programs, code sets, or instruction sets stored in memory 820, and by calling data stored in memory 820. Optionally, processor 810 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 110 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and Modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the aforementioned modem may not be integrated into the processor 810, but may be implemented using a separate communication chip. The memory 820 may include random access memory (RAM) or read-only memory (ROM). The memory 820 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 820 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the terminal 800 during use (such as phonebook data, audio and video data, chat log data, etc.).

[0093] Please refer to Figure 9 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 900 stores program code that can be called by a processor to execute the methods described in the above method embodiments.

[0094] The computer-readable storage medium 900 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 900 includes a non-volatile computer-readable storage medium. The computer-readable storage medium 900 has storage space for program code 910 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 910 may, for example, be compressed in a suitable form.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A speech synthesis method characterized by, The method comprises: Based on the base language text in the input text, a first synthesized speech is obtained according to a pre-acquired base language speech synthesis model; Based on the target language text in the input text, a second synthesized speech is obtained according to a pre-acquired target language speech synthesis model, wherein the training speech of the target language speech synthesis model is a target language speech selected from a plurality of target language speeches, and the similarity of the training speech of the base language speech synthesis model is higher than a preset value, and the similarity includes the similarity of tone, speaking style, pronunciation style and pause; Based on the pre-acquired base language training speech, the second synthesized speech is subjected to speech conversion to obtain a third synthesized speech, and the similarity of the third synthesized speech and the base language training speech is higher than the similarity of the second synthesized speech and the base language training speech; Based on the first synthesized speech and the third synthesized speech, a target synthesized speech is obtained.

2. The method of claim 1, wherein, Before the method based on the input text, a second synthesized speech is obtained according to a pre-acquired target language speech synthesis model, the method further comprises: Obtain a base language training speech and a plurality of target language speeches; Based on the base language training speech, a target language training speech is selected from a plurality of target language speeches, and the target language training speech is a target language speech with a similarity higher than a preset value to the base language training speech; Based on the target language training speech, the target language speech synthesis model is obtained.

3. The method of claim 2, wherein, Based on the base language training speech, a target language training speech is selected from a plurality of target language speeches, comprising: Compare each of the target language speeches with the base language training speech one by one to obtain the similarity of each of the target language speeches with the base language training speech; Select the target language speech with a similarity higher than a preset value as a target language training speech.

4. The method of claim 3, wherein, The method comprises: Extract the voiceprint feature of each of the target language speeches as a comparison voiceprint feature; Extract the voiceprint feature of the base language training speech as a base voiceprint feature; Voiceprint comparison is performed on a plurality of comparison voiceprint features and the base voiceprint feature to obtain a voiceprint comparison score; Based on the voiceprint comparison score, the similarity of each of the target language speeches with the base language training speech is obtained.

5. The method of claim 1, wherein, Before the method based on the input text, a first synthesized speech is obtained according to a pre-acquired base language speech synthesis model, the method further comprises: Obtain a base language training speech; Based on the base language training speech, the base language speech synthesis model is obtained.

6. The method of claim 1, wherein, In the method based on the input text, a first synthesized speech is obtained according to a pre-acquired base language speech synthesis model, and a second synthesized speech is obtained according to a pre-acquired target language speech synthesis model, comprising: Determine the language type of the input text; If the input text is the base language text, input the base language text into a pre-acquired base language speech synthesis model to obtain a first synthesized speech; If the input text is the target language text, input the target language text into a pre-acquired target language speech synthesis model to obtain a second synthesized speech.

7. The method of claim 1, wherein, The voice conversion of the tone of the second synthesized speech based on the pre-acquired base language training speech to obtain a third synthesized speech, comprising: Acquiring a voice conversion model based on the pre-acquired base language training speech; Inputting the second synthesized speech into the voice conversion model to obtain a third voice.

8. A speech synthesis apparatus characterized by comprising: The device comprises: A synthesized speech acquisition unit configured to acquire a first synthesized speech based on a base language text in an input text according to a pre-acquired base language speech synthesis model, and acquire a second synthesized speech based on a target language text in the input text according to a pre-acquired target language speech synthesis model, wherein the training speech of the target language speech synthesis model is a target language speech selected from a plurality of target language speeches and having a similarity higher than a preset value to the training speech of the base language speech synthesis model, and the similarity includes similarities in tone, speaking style, pronunciation style and pause; A voice conversion unit configured to perform voice conversion on the tone of the second synthesized speech based on the pre-acquired base language training speech to obtain a third synthesized speech, and the similarity between the third synthesized speech and the base language training speech is higher than the similarity between the second synthesized speech and the base language training speech; A speech synthesis unit configured to acquire a target synthesized speech based on the first synthesized speech and the third synthesized speech.

9. An electronic device, comprising: Comprise: One or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to execute the method of any one of claims 1-7.

10. A computer readable medium characterized by The computer readable medium stores processor executable program code, and the program code is executed by the processor to make the processor execute the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Multilingual speech synthesis method and device thereof

    CN111667814A

  • Speech synthesis method for specific field

    CN115565517A

  • Device and method for text speech synthesis

    JP2006030384A