Speech synthesis method and device, electronic equipment and storage medium

By determining the target audio characteristics and replacing game audio with a pre-trained speech synthesis model, the problem of game audio cannot be personalized is solved, and the user's sense of participation and experience is enhanced.

CN120388558APending Publication Date: 2025-07-29GUANGZHOU SANQI DREAM NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510407448.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing game audio cannot meet users' personalized needs, resulting in a decrease in user participation and experience.

Method used

By responding to the operation of the game user, the audio characteristics of the target audio are determined, and the pre-trained voice synthesis model is used to replace the original game audio with synthetic voice to meet the user's personalized needs.

Benefits of technology

Without changing the original game audio text, replace the game audio with a user-defined sound to improve the user's sense of participation and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388558A_ABST
    Figure CN120388558A_ABST
Patent Text Reader

Abstract

The invention discloses a speech synthesis method and device, electronic equipment and a storage medium, and can respond to the operation of a game user, determine a target audio uploaded by the game user, and determine the audio characteristics of the target audio. And according to the determined audio features and the text corresponding to the original game audio, determining synthetic speech of the text corresponding to the original game audio through a pre-trained speech synthesis model. And finally, replacing the original game audio with the synthetic voice. According to the method, under the condition that the text corresponding to the original game audio is not changed, the sound feature of the original game audio in the game is replaced by the sound feature wanted by the game user, the individual requirement of the user is met, and the participation sense and the experience sense of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and particularly to a method, apparatus, electronic device, and storage medium for speech synthesis. Background Art

[0002] With the rapid development of Internet technology, various game software has been developed, enabling users to experience games by downloading game software on mobile terminals.

[0003] Currently, for most games, it is usually necessary to produce game audio related to the game for playing during the game process. For example, in a Role-Playing Game (RPG), the game audio can be the lines of game characters and the conversations between game characters.

[0004] However, since the game audio in current games is usually recorded by professional voice actors and preset in the game, users cannot replace or customize the game audio themselves. Therefore, it cannot meet the personalized needs of the majority of users and reduces the sense of participation of users. It can be seen that how to make the game audio in games meet the personalized needs of users, improve the sense of participation and experience of users is an important issue.

[0005] Based on this, this specification provides a method for speech synthesis. Summary of the Invention

[0006] This specification provides a method, apparatus, electronic device, and storage medium for speech synthesis to partially solve the above problems existing in the prior art.

[0007] This specification adopts the following technical solutions:

[0008] This specification provides a method for speech synthesis, and the method includes:

[0009] In response to an operation of a game user, determine the target audio uploaded by the game user;

[0010] Determine the audio features of the target audio;

[0011] According to the audio features and the text corresponding to the original game audio, through a pre-trained speech synthesis model, determine the synthetic speech of the text corresponding to the original game audio;

[0012] Replace the original game audio with the synthetic speech.

[0013] Optionally, the audio features include timbre features, pitch features, and emotion features.

[0014] Optionally, determining the audio features of the target audio specifically includes:

[0015] Determine the audio features of the target audio through a pre-trained audio feature extraction model.

[0016] Optionally, the pre-trained speech synthesis model includes: a feature extraction network, a feature alignment network, and an audio synthesis network;

[0017] Determine the synthetic speech of the text corresponding to the original game audio through a pre-trained speech synthesis model, specifically including:

[0018] Input the text corresponding to the original game audio into the pre-trained speech synthesis model, and through the feature extraction network, obtain the original features of the text corresponding to the original game audio output by the feature extraction network;

[0019] Input the audio features into the pre-trained speech synthesis model, and through the feature alignment network, align the original features of the text corresponding to the original game audio with the audio features to obtain aligned features;

[0020] Input the aligned features into the pre-trained speech synthesis model, and through the audio synthesis network, obtain the synthetic speech of the text corresponding to the original game audio output by the audio synthesis network.

[0021] Optionally, the speech synthesis model includes: a feature extraction network, a feature alignment network, and an audio synthesis network;

[0022] The speech synthesis model to be trained is trained by the following method:

[0023] Obtain sample texts and the labeled audio corresponding to the sample texts; and obtain the audio features corresponding to the sample texts;

[0024] Input the sample texts into the feature extraction network to obtain the sample features corresponding to the sample texts;

[0025] Input the sample features and the audio features corresponding to the sample texts into the feature alignment network so that the feature alignment network aligns the sample features with the audio features corresponding to the sample texts to obtain the aligned features corresponding to the sample texts;

[0026] Input the aligned features corresponding to the sample texts into the audio synthesis network to obtain the predicted audio output by the audio synthesis network;

[0027] Train the speech synthesis model to be trained according to the predicted audio and the labeled audio.

[0028] Optionally, the feature extraction network is a Mel spectrum conversion network;

[0029] Inputting the sample text into the feature extraction network to obtain the sample features corresponding to the sample text specifically includes:

[0030] Determining the phoneme sequence of the sample text;

[0031] Inputting the phoneme sequence into the Mel spectrum conversion network to obtain the first Mel spectrum output by the Mel spectrum conversion network;

[0032] Inputting the sample features and the audio features corresponding to the sample text into the feature alignment network specifically includes:

[0033] Determining the second Mel spectrum of the audio features corresponding to the sample text;

[0034] Inputting the first Mel spectrum into the feature alignment network so that the feature alignment network aligns the first Mel spectrum with the second Mel spectrum to obtain the aligned Mel spectrum;

[0035] Obtaining the alignment features corresponding to the sample text according to the obtained aligned Mel spectrum.

[0036] Optionally, training the speech synthesis model according to the predicted audio and the labeled audio specifically includes:

[0037] Determining a loss according to the difference between the predicted audio and the labeled audio;

[0038] Training the speech synthesis model to be trained according to the loss to obtain a trained speech synthesis model.

[0039] This specification provides a device for speech synthesis, and the device includes:

[0040] An audio determination module, configured to determine a target audio uploaded by the game user in response to an operation of the game user;

[0041] A feature determination module, configured to determine the audio features of the target audio;

[0042] A speech synthesis module, configured to determine a synthesized speech of the text corresponding to the original game audio according to the audio features and the text corresponding to the original game audio through a pre-trained speech synthesis model;

[0043] An audio replacement module, configured to replace the original game audio with the synthesized speech.

[0044] Optionally, the audio features include timbre features, pitch features, and emotional features.

[0045] Optionally, the feature determination module is specifically configured to determine the audio features of the target audio through a pre-trained audio feature extraction model.

[0046] Optionally, the pre-trained speech synthesis model includes: a feature extraction network, a feature alignment network, and an audio synthesis network;

[0047] The speech synthesis module is specifically configured to input the text corresponding to the original game audio into the pre-trained speech synthesis model, and obtain the original features of the text corresponding to the original game audio output by the feature extraction network through the feature extraction network; input the audio features into the pre-trained speech synthesis model, and align the original features of the text corresponding to the original game audio with the audio features through the feature alignment network to obtain alignment features; input the alignment features into the pre-trained speech synthesis model, and obtain the synthesized speech of the text corresponding to the original game audio output by the audio synthesis network through the audio synthesis network.

[0048] Optionally, the speech synthesis model includes: a feature extraction network, a feature alignment network, and an audio synthesis network;

[0049] The device further includes a training module;

[0050] The training module is specifically configured to obtain a sample text and a labeled audio corresponding to the sample text; and obtain the audio features corresponding to the sample text; input the sample text into the feature extraction network to obtain sample features corresponding to the sample text; input the sample features and the audio features corresponding to the sample text into the feature alignment network so that the feature alignment network aligns the sample features with the audio features corresponding to the sample text to obtain alignment features corresponding to the sample text; input the alignment features corresponding to the sample text into the audio synthesis network to obtain a predicted audio output by the audio synthesis network; and train the speech synthesis model to be trained according to the predicted audio and the labeled audio.

[0051] Optionally, the feature extraction network is a Mel spectrum conversion network;

[0052] The training module is specifically configured to determine the phoneme sequence of the sample text; input the phoneme sequence into the Mel spectrum conversion network to obtain a first Mel spectrum output by the Mel spectrum conversion network;

[0053] Specifically, the training module is configured to determine the second Mel spectrogram of the audio features corresponding to the sample text; input the first Mel spectrogram into the feature alignment network, so that the feature alignment network aligns the first Mel spectrogram with the second Mel spectrogram to obtain the aligned Mel spectrogram; and obtain the aligned features corresponding to the sample text according to the obtained aligned Mel spectrogram.

[0054] Optionally, the training module is specifically configured to determine a loss according to the difference between the predicted audio and the labeled audio; and train the speech synthesis model to be trained according to the loss to obtain a trained speech synthesis model.

[0055] This specification provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above method for speech synthesis.

[0056] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above method for speech synthesis is implemented.

[0057] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:

[0058] In the method for speech synthesis provided in this specification, in response to an operation of a game user, a target audio uploaded by the game user can be determined, and the audio features of the target audio can be determined. According to the determined audio features and the text corresponding to the original game audio, a synthesized speech of the text corresponding to the original game audio can be determined through a pre-trained speech synthesis model. Finally, the original game audio is replaced with the synthesized speech.

[0059] This method can replace the sound of the original game audio in the game with the sound that the game user himself / herself wants without changing the text corresponding to the original game audio, meeting the personalized needs of the user and improving the user's sense of participation and experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the attached

[0061] In the figure:

[0062] Figure 1 is a schematic flowchart of a method for speech synthesis in this specification;

[0063] Figure 2It is a schematic flowchart of a method for speech synthesis in this specification;

[0064] Figure 3 It is a schematic diagram of the structure of a speech synthesis model in this specification;

[0065] Figure 4 It is a schematic diagram of a device for speech synthesis provided in this specification;

[0066] Figure 5 It corresponds to in this specification Figure 1 Schematic diagram of the electronic device. Specific embodiments

[0067] To make the objectives, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments and corresponding drawings of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of this specification.

[0068] In addition, it should be noted that all actions of obtaining signals, information or data in this specification are carried out on the premise of complying with the corresponding data protection regulations and policies of the location and with the authorization given by the owner of the corresponding device.

[0069] It should be noted that, without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0070] The following will, in conjunction with the drawings, detail the technical solutions provided by each embodiment of this specification.

[0071] Figure 1 It is a schematic flowchart of a method for speech synthesis provided in this specification.

[0072] S100: In response to the operation of the game user, determine the target audio uploaded by the game user.

[0073] In this specification, the execution subject of the method for speech synthesis provided in this specification is the game server. Of course, it can also be the game client, and this specification does not make specific limitations. For the convenience of description, the game server will be taken as an example below to illustrate the technical solutions of this specification.

[0074] As described in the background art, currently, game audio in games such as RPG games is pre-recorded and built into the game. Regarding game audio related to game plots, dialogue audio between game characters, and line audio of game characters, etc., game users cannot actively select or update the sound characteristics. That is to say, game users cannot use their own voices, and game users cannot select the game audio with the sound characteristics they want. Therefore, this specification provides a voice synthesis method that can replace the sound of the original game audio in the game with the sound that the game user wants without changing the text corresponding to the original game audio, meeting the personalized needs of users and improving the user's sense of participation and experience.

[0075] In this specification, the game server can, in response to the operation of the game user, determine the target audio uploaded by the game user. Among them, the target audio has the sound characteristics that the game user wants and is the audio selected by the game user to replace the original game audio.

[0076] In one or more embodiments of this specification, the target audio can be the audio currently recorded by the game user or the audio selected by the game user from their own terminal device, and this specification does not make specific limitations.

[0077] S102: Determine the audio characteristics of the target audio.

[0078] The game server can determine the audio characteristics of the target audio uploaded by the game user. When determining the audio characteristics of the target audio, the game server can use methods such as the Mel-scale Frequency Cepstral Coefficients (MFCC) algorithm, the Fliter Banks feature extraction algorithm, and wavelet transform. The above methods are all relatively mature methods for extracting audio characteristics at present, and will not be elaborated in this specification.

[0079] In addition, in one or more embodiments of this specification, the game server can also determine the audio characteristics of the target audio through a pre-trained audio feature extraction model. Specifically, the game server can input the target audio into the pre-trained audio feature extraction model, and the trained audio feature extraction model can output the audio characteristics of the target audio. Among them, the audio feature extraction model can be VGGish, PANNs (Pre-trained Audio Neural Networks), etc.

[0080] In this specification, the audio characteristics can include timbre characteristics, pitch characteristics, and emotional characteristics.

[0081] S104: According to the audio features and the text corresponding to the original game audio, determine the synthesized speech of the text corresponding to the original game audio through a pre-trained speech synthesis model.

[0082] S106: Replace the original game audio with the synthesized speech.

[0083] In this specification, the game server can input the audio features and the text corresponding to the original game audio into a pre-trained speech synthesis model, and the pre-trained speech synthesis model can output the synthesized speech of the text corresponding to the original game audio.

[0084] As Figure 2 shown, it is a schematic flowchart of a speech synthesis provided in this specification. It can be seen that after steps S100 to S104, the game server can replace the original game audio with the synthesized speech, so that game users can use game voices with the sound characteristics they select in the game, rather than the pre-recorded original game voices.

[0085] In addition, this specification does not limit the method of replacing the original game audio with the synthesized speech. For example, game developers can pre-develop an API or SDK for replacing the game audio corresponding to the game server, so as to implement replacing the original game audio with the synthesized speech by calling the API or SDK. Another example is that the game server can use an audio editing tool to replace the original game audio with the synthesized speech. Still another example is that the developers of the game server can write a script for replacing game audio, etc., to replace the original game audio with the synthesized speech.

[0086] Based on the above Figure 1 shown speech synthesis method, the game server can replace the sound of the original game audio in the game with the sound that the game user himself wants without changing the text corresponding to the original game audio, meeting the personalized needs of users and improving the user's sense of participation and experience.

[0087] In one or more embodiments of this specification, the pre-trained speech synthesis model includes: a feature extraction network, a feature alignment network, and an audio synthesis network, as Figure 3As shown, it is a schematic structural diagram of a speech synthesis model provided in this specification. In the above step S104, when inputting the text corresponding to the original game audio into the pre-trained speech synthesis model, the game server can input the text corresponding to the original game audio into the pre-trained speech synthesis model, so as to obtain the original features of the text corresponding to the original game audio output by the feature extraction network through this feature extraction network. Then, input the audio features into the pre-trained speech synthesis model, and align the original features of the text corresponding to the original game audio with the audio features through this feature alignment network to obtain aligned features. Finally, input the aligned features into the pre-trained speech synthesis model, and obtain the synthetic speech of the text corresponding to the original game audio output by this audio synthesis network through the audio synthesis network.

[0088] Among them, the feature extraction network can be used to extract semantic features and pronunciation features from the text corresponding to the original game audio. The semantic features and pronunciation features include phoneme sequences, lexical information, grammatical structures, etc., to be used to guide the subsequent network layers to generate synthetic speech. The feature alignment network can be used to align the original features of the text corresponding to the original game audio with the audio features to ensure that the finally output synthetic speech can not only accurately reflect the content of the text corresponding to the original game speech, but also be consistent with the target audio selected by the game user, that is, the audio desired by the game user, in terms of sound characteristics such as timbre features, pitch features, and emotional features. The audio synthesis network can be used to generate the final synthetic speech according to the aligned features output by the feature alignment network, that is, the audio waveform or spectrogram.

[0089] In one or more embodiments of this specification, the feature extraction network can be a Mel spectrum conversion network, the feature alignment network can be: a long short-term memory network and a convolutional neural network using an attention mechanism, and the audio synthesis network can be WaveNet, Tacotron, etc.

[0090] This specification provides a method for training a speech synthesis model. Specifically, the game server can obtain sample texts and the labeled audio corresponding to the sample texts, and obtain the audio features corresponding to the sample texts. Then, the sample texts can be input into the feature extraction network to obtain the sample features corresponding to the sample texts. Next, the sample features and the audio features corresponding to the sample texts are input into the feature alignment network, so that the feature alignment network aligns the sample features with the audio features corresponding to the sample texts to obtain the alignment features corresponding to the sample texts. Furthermore, the game server can input the alignment features corresponding to the sample texts into the audio synthesis network to obtain the predicted audio output by the audio synthesis network. Finally, the speech synthesis model to be trained can be trained according to the predicted audio and the labeled audio. Specifically, when training the speech synthesis model according to the predicted audio and the labeled audio, the game server can determine the loss according to the difference between the predicted audio and the labeled audio, and train the speech synthesis model to be trained according to this loss, so as to obtain a trained speech synthesis model.

[0091] When the feature extraction network is a Mel spectrum conversion network, when inputting the sample text into the feature extraction network to obtain the sample features corresponding to the sample text, the game server can first determine the phoneme sequence of the sample text, and input the phoneme sequence of the sample text into the Mel spectrum conversion network to obtain the first Mel spectrum output by the Mel spectrum conversion network.

[0092] When inputting the sample features and the audio features corresponding to the sample text into the feature alignment network, the second Mel spectrum of the audio features corresponding to the sample text can be determined, and the first Mel spectrum is input into the feature alignment network, so that the feature alignment network aligns the first Mel spectrum with the second Mel spectrum to obtain the aligned Mel spectrum, and thus the alignment features corresponding to the sample text are obtained according to the obtained aligned Mel spectrum.

[0093] Of course, in the case where the feature extraction network is a Mel spectrum conversion network, when using this trained speech synthesis model to determine the synthesized speech of the text corresponding to the original game audio according to the audio features and the text corresponding to the original game audio, the game server can first determine the phoneme sequence of the text corresponding to the original game audio, and then input the phoneme sequence of the text corresponding to the original game audio into the Mel spectrum conversion network to obtain the Mel spectrum of the text corresponding to the original game audio output by the Mel spectrum conversion network through this Mel spectrum conversion network, and thus obtain the original features of the text corresponding to the original game audio based on the Mel spectrum of the text corresponding to the original game audio.

[0094] The above is the method for speech synthesis provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding speech synthesis device, as Figure 4 shown.

[0095] Figure 4 A schematic diagram of a voice synthesis device provided for this specification. The device includes:

[0096] An audio determination module 400, configured to determine a target audio uploaded by the game user in response to an operation of the game user.

[0097] A feature determination module 402, configured to determine the audio features of the target audio.

[0098] A voice synthesis module 404, configured to determine a synthesized voice of the text corresponding to the original game audio through a pre-trained voice synthesis model according to the audio features and the text corresponding to the original game audio.

[0099] An audio replacement module 406, configured to replace the original game audio with the synthesized voice.

[0100] Optionally, the audio features include timbre features, pitch features, and emotion features.

[0101] Optionally, the feature determination module 402 is specifically configured to determine the audio features of the target audio through a pre-trained audio feature extraction model.

[0102] Optionally, the pre-trained voice synthesis model includes: a feature extraction network, a feature alignment network, and an audio synthesis network.

[0103] The voice synthesis module 404 is specifically configured to input the text corresponding to the original game audio into the pre-trained voice synthesis model, obtain the original features of the text corresponding to the original game audio output by the feature extraction network through the feature extraction network; input the audio features into the pre-trained voice synthesis model, align the original features of the text corresponding to the original game audio with the audio features through the feature alignment network to obtain aligned features; input the aligned features into the pre-trained voice synthesis model, and obtain the synthesized voice of the text corresponding to the original game audio output by the audio synthesis network through the audio synthesis network.

[0104] Optionally, the voice synthesis model includes: a feature extraction network, a feature alignment network, and an audio synthesis network.

[0105] The device further includes a training module 408.

[0106] The training module 408 is specifically configured to obtain sample texts and the labeled audio corresponding to the sample texts; obtain the audio features corresponding to the sample texts; input the sample texts into the feature extraction network to obtain the sample features corresponding to the sample texts; input the sample features and the audio features corresponding to the sample texts into the feature alignment network, so that the feature alignment network aligns the sample features with the audio features corresponding to the sample texts to obtain the aligned features corresponding to the sample texts; input the aligned features corresponding to the sample texts into the audio synthesis network to obtain the predicted audio output by the audio synthesis network; and train the speech synthesis model to be trained according to the predicted audio and the labeled audio.

[0107] Optionally, the feature extraction network is a Mel spectrum conversion network;

[0108] The training module 408 is specifically configured to determine the phoneme sequence of the sample text; input the phoneme sequence into the Mel spectrum conversion network to obtain the first Mel spectrum output by the Mel spectrum conversion network;

[0109] The training module 408 is specifically configured to determine the second Mel spectrum of the audio features corresponding to the sample text; input the first Mel spectrum into the feature alignment network, so that the feature alignment network aligns the first Mel spectrum with the second Mel spectrum to obtain the aligned Mel spectrum; and obtain the aligned features corresponding to the sample text according to the obtained aligned Mel spectrum.

[0110] Optionally, the training module 408 is specifically configured to determine a loss according to the difference between the predicted audio and the labeled audio; and train the speech synthesis model to be trained according to the loss to obtain a trained speech synthesis model.

[0111] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above Figure 1 shown method for speech synthesis.

[0112] This specification also provides Figure 5 a schematic structural diagram of the electronic device shown. As Figure 5 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 shown method for speech synthesis.

[0113] Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or logical devices.

[0114] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0115] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0116] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0117] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0118] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0119] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0120] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0122] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0123] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0124] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0125] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0126] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, system, or computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0128] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the corresponding description in the method embodiment.

[0129] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A method for speech synthesis, characterized in that, The method includes: In response to an operation of a game user, determining a target audio uploaded by the game user; Determining audio features of the target audio; According to the audio features and the text corresponding to the original game audio, determining a synthesized speech of the text corresponding to the original game audio through a pre-trained speech synthesis model; Replacing the original game audio with the synthesized speech.

2. The method according to claim 1, characterized in that, The audio features include timbre features, pitch features, and emotion features.

3. The method according to claim 1, wherein Determining the audio features of the target audio specifically includes: Determining the audio features of the target audio through a pre-trained audio feature extraction model.

4. The method according to claim 1, characterized in that The pre-trained speech synthesis model includes: a feature extraction network, a feature alignment network, and an audio synthesis network; Determining the synthesized speech of the text corresponding to the original game audio through a pre-trained speech synthesis model specifically includes: Inputting the text corresponding to the original game audio into the pre-trained speech synthesis model, and obtaining original features of the text corresponding to the original game audio output by the feature extraction network through the feature extraction network; Inputting the audio features into the pre-trained speech synthesis model, and aligning the original features of the text corresponding to the original game audio with the audio features through the feature alignment network to obtain aligned features; Inputting the aligned features into the pre-trained speech synthesis model, and obtaining a synthesized speech of the text corresponding to the original game audio output by the audio synthesis network through the audio synthesis network.

5. The method according to claim 1, characterized in that, The speech synthesis model includes: a feature extraction network, a feature alignment network, and an audio synthesis network; The speech synthesis model to be trained is trained by the following method: Obtaining a sample text and a labeled audio corresponding to the sample text; and obtaining audio features corresponding to the sample text; Inputting the sample text into the feature extraction network to obtain sample features corresponding to the sample text; Inputting the sample features and the audio features corresponding to the sample text into the feature alignment network, so that the feature alignment network aligns the sample features with the audio features corresponding to the sample text to obtain aligned features corresponding to the sample text; Inputting the aligned features corresponding to the sample text into the audio synthesis network to obtain a predicted audio output by the audio synthesis network; Training the speech synthesis model to be trained according to the predicted audio and the labeled audio.

6. The method according to claim 5, wherein The feature extraction network is a Mel spectrum conversion network; Inputting the sample text into the feature extraction network to obtain sample features corresponding to the sample text specifically includes: Determining a phoneme sequence of the sample text; Inputting the phoneme sequence into the Mel spectrum conversion network to obtain a first Mel spectrum output by the Mel spectrum conversion network; Inputting the sample features and the audio features corresponding to the sample text into the feature alignment network specifically includes: Determining a second Mel spectrum of the audio features corresponding to the sample text; Input the first Mel spectrogram into the feature alignment network, so that the feature alignment network aligns the first Mel spectrogram with the second Mel spectrogram to obtain the aligned Mel spectrogram; Obtain the aligned features corresponding to the sample text according to the obtained aligned Mel spectrogram.

7. The method according to claim 5, wherein Train the speech synthesis model according to the predicted audio and the labeled audio, which specifically includes: Determine the loss according to the difference between the predicted audio and the labeled audio; Train the speech synthesis model to be trained according to the loss to obtain a trained speech synthesis model.

8. An apparatus for speech synthesis, characterized in that, The device includes: An audio determination module, configured to determine the target audio uploaded by the game user in response to an operation of the game user; A feature determination module, configured to determine the audio features of the target audio; A speech synthesis module, configured to determine the synthesized speech of the text corresponding to the original game audio through a pre-trained speech synthesis model according to the audio features and the text corresponding to the original game audio; An audio replacement module, configured to replace the original game audio with the synthesized speech.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 above is implemented.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 7 above is implemented.