Cloned sound generation method, cloned sound application method and device
The cloned sound generation is achieved through a three-terminal interaction method, allowing users to generate cloned sounds themselves. This solves the problem of complex customized cloned sound processes and realizes the automation and flexible application of cloned sounds.
Patent Information
- Application Number
- CN202411237858.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-09-04
AI Technical Summary
The current process for customizing cloned audio is complex, making it difficult for users to generate the cloned audio they require on their own.
A method for generating cloned sounds is provided, which realizes sound cloning through three-terminal interaction. The client collects sound samples and sends them to the server for training. The application client requests the server to generate cloned sound synthesis data and applies the cloned sound in various scenarios.
Users can clone their own voices, achieving automation and flexibility in the acquisition of cloned voices. It integrates voice acquisition, training, and application, simplifying the process of customizing cloned voices.
Smart Images

Figure CN119132319B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speech processing technology, and more particularly to the field of voice cloning technology. Background Technology
[0002] Currently, voice cloning is used in various scenarios, such as live streaming, social media, and voice navigation. Among these, some users have a need for customized cloned voices. However, the current process for customizing cloned voices is complex, requiring technicians to record and clone the user's voice, making it difficult for users to generate the cloned voice they require themselves. Summary of the Invention
[0003] This disclosure provides a method for generating cloned sounds, a method for applying cloned sounds, and an apparatus for solving at least one of the above-mentioned technical problems.
[0004] According to a first aspect of this disclosure, a method for generating cloned sounds is provided, wherein the method is applied on a server side, the method comprising:
[0005] In response to a cloned sound synthesis request sent by the application client, the cloned sound identifier and the content to be synthesized carried in the cloned sound synthesis request are obtained, and the corresponding sound clone model is determined according to the cloned sound identifier.
[0006] Based on the content to be synthesized and the sound cloning model, cloned sound synthesis data corresponding to the content to be synthesized is obtained, and the cloned sound synthesis data is returned to the application client, wherein the application client is configured to receive and apply the cloned sound synthesis data;
[0007] The sound cloning model is obtained in the following way:
[0008] The user's voice samples collected and sent by the acquisition client are used as training data and input into the initial voice cloning model for training to obtain a trained voice cloning model corresponding to the user's voice samples. The voice cloning model is used to output cloned sound synthesis data with a similar timbre to the user's voice samples.
[0009] According to a second aspect of this disclosure, a method for applying cloned sounds is provided, wherein the application is performed on an application client, the method comprising:
[0010] Based on the content to be synthesized and the clone sound identifier of the sound clone model corresponding to the required clone sound, a clone sound synthesis request is generated and sent to the server. The server is configured to, in response to the application of the clone sound synthesis request, obtain the clone sound identifier and the content to be synthesized carried in the clone sound synthesis request, determine the corresponding sound clone model according to the clone sound identifier, obtain the clone sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound clone model, and return the clone sound synthesis data.
[0011] Receive the cloned sound synthesis data returned by the server and apply the cloned sound synthesis data.
[0012] According to a third aspect of this disclosure, a method for generating cloned sounds is provided, wherein the method is applied in a data acquisition client, the method comprising:
[0013] Collect user voice samples and send them to the server;
[0014] The server processes the sound sample according to the method described above.
[0015] According to a fourth aspect of this disclosure, a cloned sound generation apparatus is provided, wherein the apparatus comprises:
[0016] The model determination module is used to respond to the cloned sound synthesis request sent by the application client, obtain the cloned sound identifier and the content to be synthesized carried in the cloned sound synthesis request, and determine the corresponding sound cloned model according to the cloned sound identifier.
[0017] A sound synthesis module is used to obtain cloned sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound cloning model, and return the cloned sound synthesis data to the application client, wherein the application client is configured to receive and apply the cloned sound synthesis data;
[0018] The device further includes a model acquisition module, used to obtain the sound clone model in the following manner:
[0019] The user's voice samples collected and sent by the acquisition client are used as training data and input into the initial voice cloning model for training to obtain a trained voice cloning model corresponding to the user's voice samples. The voice cloning model is used to output cloned sound synthesis data with a similar timbre to the user's voice samples.
[0020] According to a fifth aspect of this disclosure, a cloned sound application apparatus is provided, wherein the apparatus includes:
[0021] The request module is used to generate a clone sound synthesis request based on the content to be synthesized and the clone sound identifier of the sound clone model corresponding to the required clone sound, and send it to the server. The server is configured to, in response to the application of the clone sound synthesis request, obtain the clone sound identifier and the content to be synthesized carried in the clone sound synthesis request, determine the corresponding sound clone model according to the clone sound identifier, obtain the clone sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound clone model, and return the clone sound synthesis data.
[0022] The application module is used to receive the cloned sound synthesis data returned by the server and apply the cloned sound synthesis data.
[0023] According to the sixth aspect of this disclosure, a data acquisition module is provided for acquiring user voice samples and sending them to a server;
[0024] The server processes the sound samples according to the method described above.
[0025] According to a seventh aspect of this disclosure, an electronic device is provided, comprising:
[0026] At least one processor; and
[0027] A memory communicatively connected to the at least one processor; wherein,
[0028] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the above-described method.
[0029] According to an eighth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method described above.
[0030] According to a ninth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described above.
[0031] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0032] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0033] Figure 1 This is a flowchart illustrating a method for generating cloned sounds according to the first embodiment of this disclosure;
[0034] Figure 2 This is a flowchart illustrating the method for obtaining a sound cloning model provided in the first embodiment of this disclosure;
[0035] Figure 3 This is a flowchart illustrating a method for applying cloned sounds according to a second embodiment of this disclosure;
[0036] Figure 4 This is a flowchart illustrating a method for generating cloned sounds according to a third embodiment of this disclosure;
[0037] Figure 5 This is a flowchart illustrating another method for generating cloned sounds provided in the third embodiment of this disclosure;
[0038] Figure 6 This is a schematic diagram of the structure of a cloned sound generation device provided in the fourth embodiment of this disclosure;
[0039] Figure 7 This is a schematic diagram of the structure of a cloned sound application device provided in the fifth embodiment of this disclosure;
[0040] Figure 8 This is a schematic diagram of the structure of a cloned sound generation device provided in the sixth embodiment of this disclosure;
[0041] Figure 9 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0042] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0043] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0044] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0045] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0046] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0047] The cloned voice generation and application methods disclosed herein can be executed by electronic devices such as terminal devices or servers. Terminal devices can be in-vehicle devices, user equipment (UE), mobile devices, user terminals, terminals, cellular phones, cordless phones, personal digital assistants (PDAs), handheld devices, computing devices, in-vehicle devices, wearable devices, etc. The methods can be implemented by a processor calling computer-readable program instructions stored in memory. Alternatively, the cloned voice generation and application methods provided herein can be executed by a server.
[0048] The method disclosed herein can be applied in a three-terminal interaction scenario, comprising a server, an application client, and a data acquisition client. The data acquisition client is used to collect user voice samples; the server is used to train a voice cloning model based on the voice samples for speech synthesis; and the application client is the client that ultimately applies the technology and plays the cloned audio.
[0049] In some scenarios, the acquisition client is a client using a microservice architecture, in other words, a mini-program client; the application client is a client using a global wide area network, in other words, a web client. In some cases, the audio recording devices of web clients vary greatly, and the web mechanism is difficult to use for audio acquisition, while mobile devices have more uniform audio recording devices, a better audio recording environment, and mini-programs are easier to encode to achieve audio acquisition functionality. Therefore, in this disclosure, a mini-program is used as the acquisition client to collect audio samples, which are then input to the server for training to obtain an audio cloning model. The web client, as the application client, communicates directly with the server to request the server to synthesize cloned audio, thereby combining the characteristics of different clients to achieve convenient and accurate cloned audio generation and application.
[0050] In the first disclosed embodiment, see Figure 1 , Figure 1 This diagram illustrates a flowchart of a method for generating cloned sounds according to a first embodiment of this disclosure. The method can be applied to a server and includes:
[0051] S101. In response to the cloned sound synthesis request sent by the application client, obtain the cloned sound identifier and the content to be synthesized carried in the cloned sound synthesis request, and determine the corresponding sound clone model according to the cloned sound identifier.
[0052] S102. Based on the content to be synthesized and the sound cloning model, obtain the cloned sound synthesis data corresponding to the content to be synthesized, and return the cloned sound synthesis data to the application client.
[0053] The application client is configured to receive and apply cloned sound synthesis data.
[0054] Among them, see Figure 2 , Figure 2 The flowchart illustrates the method for obtaining a sound clone model, which is obtained in the following way:
[0055] S100. Use the user's voice samples collected and sent by the acquisition client as training data, input them into the initial voice cloning model for training, and obtain the trained voice cloning model corresponding to the user's voice samples.
[0056] Among them, the sound cloning model is used to output cloned sound synthesis data that is similar in timbre to the user's sound sample.
[0057] This disclosure provides a method that allows users to perform sound cloning themselves, obtaining cloned audio without the need for external personnel, thus automating the process of acquiring cloned audio. Furthermore, it employs a three-terminal interactive approach for sound cloning and application. The acquisition client collects sound samples and transmits them to the server for training, resulting in a sound sample cloning model. Based on this, when the cloned audio needs to be applied, the application client requests the server to generate cloned audio data corresponding to the content to be synthesized, which is then returned to the application client. The application client can then apply the cloned audio in various scenarios. This approach integrates sound acquisition, training, and application, making sound cloning more flexible.
[0058] The sound cloning model or initial sound cloning model can be implemented in various ways. For example, it can include a large model, or it can use network structures such as recurrent neural networks (RNNs) or encoders (Transformers), etc., without any limitation.
[0059] In some examples, the cloned sound identifier is obtained in the following way:
[0060] S100' After obtaining the trained sound cloning model, generate a clone sound identifier for the sound cloning model, and send the clone sound identifier and the mapping relationship between the clone sound identifier and the sound sample to the application client.
[0061] In other words, the server stores the mapping relationship between clone sound identifier (mid), sound clone model, and sound sample. After training the sound clone model corresponding to a sound sample, the server generates a clone sound identifier for the sound clone model and sends the clone sound identifier and mapping relationship to the application client. The application client saves the clone sound identifier and mapping relationship. When the application needs to use the sound clone model of a certain sound sample, it configures the clone sound identifier and sends it to the server. The server calls the corresponding sound clone model according to the clone sound identifier to output the clone sound that expresses the content to be synthesized.
[0062] In some examples, S100 includes:
[0063] S1001. Use the user's voice samples collected and sent by the acquisition client as training data and input them into the initial voice cloning model for training.
[0064] S1002. During the training process, sound samples are identified to obtain sound adjustment parameters. The model parameters of the initial sound cloning model are adjusted according to the sound adjustment parameters to obtain a trained sound cloning model corresponding to the user's sound sample.
[0065] The voice adjustment parameters can represent voice features, including at least one of the following: keyword parameters of the voice sample, emotional feature parameters of the voice sample, accent feature parameters of the voice sample, and of course, other voice adjustment parameters, which are not limited here.
[0066] The keyword parameters for the sound samples can be obtained by converting the sound samples into text and extracting the keywords as keyword parameters. Emotional feature parameters can be obtained by identifying emotional features based on the sound samples and using these identified emotional features as emotional feature parameters. The identification of emotional features can be achieved through an emotional feature extraction model. Accent feature parameters can also be obtained by identifying accent features based on the sound samples and using these identified accent features as accent feature parameters. During training, sound adjustment parameters can be extracted from the user's sound samples, and then the model parameters can be automatically adjusted based on these sound adjustment parameters to ensure that the cloned sound output by the sound cloning model adapts to the user's sound characteristics, thereby improving the realism of the generated cloned sound.
[0067] In some examples, S100 also includes:
[0068] S1001. Use the user's voice samples collected and sent by the acquisition client as training data and input them into the initial voice cloning model for training.
[0069] S1002. During the training process, sound samples are identified to obtain sound adjustment parameters. The model parameters of the initial sound cloning model are adjusted according to the sound adjustment parameters to obtain a trained sound cloning model corresponding to the user's sound sample.
[0070] S1003. Extract any audio segment from the audio sample to obtain the test content text.
[0071] S1004. Input the test content text into the sound cloning model to obtain the output cloned sound synthesis data corresponding to the test content text.
[0072] S1005. Based on the cloned sound synthesis data corresponding to the test content text and the sound sample corresponding to the test content text, a similarity verification is performed. If the verification passes, the sound clone model is determined to be compliant.
[0073] Specifically, a text segment can be extracted from the sound sample as the test content text, which is then input into the sound cloning model. The output is synthesized data of cloned voice reading the test content text through text-to-speech (TTS) technology. The similarity between the synthesized data of cloned voice of the test content text and the corresponding source (i.e., the sound sample) of the test content text is then compared to perform similarity compliance verification, so as to ensure the accuracy of the sound cloning model and the cloning quality.
[0074] For example, in some embodiments, the content of the sound sample is "It is now xx year xx month xx day xx hour, I am recording a sound sample". The sound segment "I am recording a sound sample" is extracted and converted into test content text. The test content text "I am recording a sound sample" is then input into the trained sound cloning model corresponding to the sound sample. Through TTS technology, the output is synthesized data of cloned sound reading "I am recording a sound sample" in the cloned sound of the sound sample. Then, the similarity between the sound segment of the sound sample and the obtained cloned sound synthesized data is compared. If the similarity is higher than a preset threshold, the sound cloning model is determined to be compliant.
[0075] In some examples, the similarity comparison can be multi-dimensional, including, for example, the similarity of voice timbre, the similarity of keyword pronunciation, the similarity of intonation, and the similarity of emotion, etc., without limitation here. In other words, in S1005, the similarity qualification check is performed based on the cloned sound synthesis data corresponding to the test content text and the sound sample corresponding to the test content text, including at least one of the following:
[0076] Based on the cloned sound synthesis data corresponding to the test content text, and the sound samples corresponding to the test content text, a timbre similarity verification is performed.
[0077] Based on the cloned sound synthesis data corresponding to the test content text and the sound samples corresponding to the test content text, the similarity of the keyword pronunciation is verified.
[0078] The intonation similarity is verified based on the cloned sound synthesis data corresponding to the test content text and the sound samples corresponding to the test content text.
[0079] Based on the cloned voice synthesis data corresponding to the test content text, and the voice samples corresponding to the test content text, the emotional similarity is verified.
[0080] Finally, the final similarity score is determined by combining the similarity scores from multiple similarity verifications. The final score can be the average of the similarity scores from multiple similarity verifications, or different weights can be assigned to each similarity verification and then a weighted calculation can be performed, with the weighted result being the final score. There is no limitation on this.
[0081] In some examples, the content to be synthesized can be in the form of speech or in the form of text. In the case of text content, the text content is directly input into the voice cloning model, and the output is synthesized data of cloned voice reading out the content to be synthesized through TTS technology.
[0082] When the content to be synthesized is in speech form, in step S102, the cloned speech synthesis data corresponding to the content to be synthesized is obtained based on the content to be synthesized and the sound cloning model, including:
[0083] Step 1: Convert the speech content to be synthesized into text content using speech-to-text conversion.
[0084] Step 2: Input the text content to be synthesized into the sound cloning model, and obtain the cloned sound synthesis data corresponding to the content to be synthesized by text-to-speech conversion.
[0085] In this scenario, the server first converts the speech to text and then inputs it into the voice cloning model. For the user's application client, the user directly inputs the original audio and receives a cloned audio with the same content, thus improving the user experience.
[0086] In some examples, the cloned sound synthesis data is in streaming data form or in audio packet form. That is, this disclosure supports both synthesis methods.
[0087] In the first approach, the cloned audio synthesis data is streamed. In this scenario, the user's application client sends a cloned audio synthesis request, including a cloned audio identifier and the content to be synthesized, to the server in real time. The server calls the audio clone model corresponding to the cloned audio identifier, inputs the content to be synthesized, and performs cloned audio synthesis, outputting multiple sub-cloned audio synthesis data segments, which are then sent to the application client in real time. These multiple sub-cloned audio synthesis data segments form a complete cloned audio synthesis data segment. Because this approach requires high real-time performance of the synthesized cloned audio, it uses a streaming method to return the streaming sub-cloned audio synthesis data to the application client. This means the completed synthesis task is broken down into multiple sub-tasks. After completing the cloned audio synthesis for one sub-task, the server returns the resulting sub-cloned audio synthesis data to the application client. This allows the application client to implement a cloned audio application that synthesizes and plays simultaneously, meeting the needs of real-time dialogue scenarios.
[0088] In the second method, the cloned sound synthesis data is in the form of an audio packet. In this case, the user's application client generates a cloned sound synthesis request based on the content to be synthesized and the cloned sound identifier and sends it to the server through an interface. The server calls the sound cloning model corresponding to the cloned sound identifier to synthesize the content to be synthesized completely into cloned sound synthesis data, and then packages the cloned sound synthesis data into an audio packet and sends the audio packet to the application client. The application client parses the audio packet to obtain the cloned sound synthesis data.
[0089] The two forms mentioned above can be switched and combined at will to increase the flexibility of application scenarios.
[0090] In some examples, S102 specifically includes:
[0091] Step 1: Obtain the cloned sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound cloning model.
[0092] Step 2: Obtain the sound effect parameters based on the content to be synthesized, and adjust the cloned sound synthesis data based on the sound effect parameters.
[0093] Step 3: Return the cloned sound synthesis data to the application client.
[0094] Among them, the sound effect parameters are parameters that can add sound effects to the cloned voice. The sound effect parameters can be matched according to the emotional characteristics and keywords of the content to be synthesized. As an example, the sound effect parameters include at least one of the following: sound effect parameters for adjusting intonation; sound effect parameters for adjusting accent; sound effect parameters for adjusting the sound's auditory space; sound effect parameters for adjusting the voice's age; sound effect parameters for adjusting the voice's gender, etc., without limitation.
[0095] For example, in some embodiments, the sound effect parameter is a sound effect parameter for adjusting the voice age; based on step two, the content to be synthesized is analyzed, and if the analysis results determine that the voice age is too low, then a sound effect parameter for reducing the voice age is generated, and the cloned voice synthesis data is adjusted according to the sound effect parameter. Then, the adjusted cloned voice synthesis data with the reduced voice age is sent to the client. In this way, the generated cloned voice can be intelligently adjusted.
[0096] In the second disclosed embodiment, see Figure 3 , Figure 3 This diagram illustrates a flowchart of a method for applying cloned sounds according to a second embodiment of the present disclosure. The method can be applied to an application client and includes:
[0097] S201. Generate a clone sound synthesis request based on the clone sound identifier of the sound clone model corresponding to the content to be synthesized and the required clone sound, and send it to the server.
[0098] The server is configured to respond to the application's cloned sound synthesis request, obtain the cloned sound identifier and the content to be synthesized carried in the cloned sound synthesis request, determine the corresponding sound clone model based on the cloned sound identifier, obtain the cloned sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound clone model, and return the cloned sound synthesis data.
[0099] S202. Receive the cloned sound synthesis data returned by the server and apply the cloned sound synthesis data.
[0100] In some examples, the application client is a client used for live streaming; S202 applies cloned audio synthesis data, including:
[0101] The cloned audio synthesis data is converted into live stream data, and then applied to the live stream.
[0102] In some cases, cloned audio synthesis data output by the audio cloning model can be applied in live streaming. In such cases, a script can be configured directly in the live streaming application client. The script contains the cloned audio identifier and the content to be synthesized. When the script is run, it requests the server to obtain the cloned audio synthesis data. The application client then converts the cloned audio synthesis data into live streaming data and applies it in the live stream, thereby achieving the effect of applying audio cloning to digital human live streaming in real time.
[0103] In some examples, prior to S201, the method also includes:
[0104] S200: Receive and save the cloned sound identifier sent by the server, as well as the mapping relationship between the cloned sound identifier and the sound sample.
[0105] The cloned sound identifier is generated according to S100' of the first disclosed embodiment.
[0106] The server stores a mapping relationship between cloned sound identifier (mid), sound clone model, and sound sample. After training the sound clone model corresponding to a sound sample, the server generates a cloned sound identifier for the sound clone model and sends the cloned sound identifier and mapping relationship to the application client. The application client saves the cloned sound identifier and mapping relationship. When the application needs to use the sound clone model of a certain sound sample, it configures the cloned sound identifier and sends it to the server. The server calls the corresponding sound clone model according to the cloned sound identifier to output the cloned sound that expresses the content to be synthesized.
[0107] In the second disclosed embodiment, see Figure 4 , Figure 4 This diagram illustrates a flowchart of a method for generating cloned sounds according to a third embodiment of this disclosure. This method can be applied to a data acquisition client and includes:
[0108] S301. Collect user's voice samples and send them to the server.
[0109] The server processes the sound samples according to the method of the first publicly disclosed embodiment.
[0110] In some scenarios, such as when the application client is a web client, the sound collection effect of the web client is poor. Therefore, the sound can be collected by the mobile device's audio equipment, separating the collection and application steps. The sound samples collected by the client are sent to the server for training.
[0111] See in some examples Figure 5 , Figure 5 A flowchart illustrating another method for generating cloned sounds provided in the third embodiment is shown, which includes:
[0112] S401. Noise detection is performed on the user's sound acquisition environment.
[0113] S402. Based on the noise detection results, generate a prompt for user voice collection.
[0114] Noise detection is performed to ensure the quality of the data collection environment. Sound collection prompts can take many forms, such as text or voice prompts to the user indicating a poor sound environment, or text or voice prompts to the user indicating that data collection can begin. No specific form is specified here.
[0115] S403. Collect voiceprint samples of the user reading the standard text according to the preset standard text.
[0116] S404. Collect other sound samples from the user according to the preset duration.
[0117] The audio samples include voiceprint samples and other audio samples. Voiceprint samples serve as reference data for model training, representing the user's standard voice. Other audio samples enhance the comprehensiveness of the audio sample set, providing sufficient and diverse samples for training to ensure effective model training. Therefore, each user's voiceprint sample can use standard text with fixed content to ensure consistency across users. Other audio samples can be freely read aloud by the user or provided with optional text, collected over a sufficient period; thus, a preset duration is set to collect enough samples.
[0118] S403 and S404 are implementation methods for collecting user voice samples in S301.
[0119] S405. Perform quality processing on the sound sample to obtain the processed sound sample.
[0120] The quality processing includes at least one of the following: noise reduction, volume normalization, and text alignment. This quality processing ensures the quality of the audio samples, thereby improving the quality of model training.
[0121] S406. Send the sound sample to the server.
[0122] Based on and Figure 1 The same principle, Figure 6 This invention discloses a cloned sound generation apparatus 60 according to a fourth embodiment of the present invention, the apparatus comprising:
[0123] The model determination module 601 is used to respond to the cloned sound synthesis request sent by the application client, obtain the cloned sound identifier and the content to be synthesized carried in the cloned sound synthesis request, and determine the corresponding sound cloned model according to the cloned sound identifier.
[0124] The sound synthesis module 602 is used to obtain the cloned sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound cloning model, and return the cloned sound synthesis data to the application client, wherein the application client is configured to receive and apply the cloned sound synthesis data.
[0125] The device also includes a model acquisition module 600, used to obtain a sound clone model in the following manner:
[0126] The user's voice samples collected and sent by the acquisition client are used as training data and input into the initial voice cloning model for training. This results in a trained voice cloning model that corresponds to the user's voice samples. The voice cloning model is used to output cloned sound synthesis data that has a similar timbre to the user's voice samples.
[0127] In some examples, the apparatus further includes: an identifier generation module for obtaining cloned sound identifiers in the following manner:
[0128] After obtaining the trained sound cloning model, a clone sound identifier is generated for the sound cloning model, and the clone sound identifier and the mapping relationship between the clone sound identifier and the sound sample are sent to the application client.
[0129] In some examples, the model acquisition module is specifically used for:
[0130] The user's voice samples collected and sent by the client are used as training data and input into the initial voice cloning model for training.
[0131] During training, sound samples are identified to obtain sound adjustment parameters. The model parameters of the initial sound clone model are adjusted according to the sound adjustment parameters to obtain a trained sound clone model corresponding to the user's sound sample.
[0132] In some examples, the voice adjustment parameters include at least one of the following: keyword parameters of the voice sample, emotion feature parameters of the voice sample, and accent feature parameters of the voice sample.
[0133] In some examples, the model acquisition module is also used for:
[0134] Extract any segment of the audio sample to obtain the test content text;
[0135] Input the test content text into the sound cloning model to obtain the output cloned sound synthesis data corresponding to the test content text;
[0136] Based on the cloned audio synthesis data corresponding to the test content text, and the audio samples corresponding to the test content text, a similarity verification is performed. If the verification passes, the audio cloning model is determined to be compliant.
[0137] In some examples, the content to be synthesized is in the form of speech.
[0138] The sound synthesis module is specifically used for:
[0139] The speech-to-text method converts the content to be synthesized from speech to text.
[0140] The text content to be synthesized is input into the sound cloning model, and the cloned sound synthesis data corresponding to the content to be synthesized is obtained by text-to-speech conversion.
[0141] In some examples, the cloned sound synthesis data is in streaming data form or in audio packet form.
[0142] In some examples, the device also includes:
[0143] The effects adjustment module is used to obtain sound effect parameters based on the content to be synthesized, and to adjust the cloned sound synthesis data based on the sound effect parameters.
[0144] Based on and Figure 3 The same principle, Figure 7 This invention discloses a cloned sound application device 70 according to a fifth embodiment of the present disclosure. The device includes:
[0145] The request module 701 is used to generate a clone sound synthesis request based on the content to be synthesized and the clone sound identifier of the sound clone model corresponding to the required clone sound, and send it to the server. The server is configured to respond to the application clone sound synthesis request, obtain the clone sound identifier and the content to be synthesized carried in the clone sound synthesis request, determine the corresponding sound clone model according to the clone sound identifier, obtain the clone sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound clone model, and return the clone sound synthesis data.
[0146] Application module 702 is used to receive cloned sound synthesis data returned by the server and apply the cloned sound synthesis data.
[0147] In some examples, the device also includes:
[0148] The identifier storage module is used to receive and store the cloned sound identifiers sent by the server, as well as the mapping relationship between the cloned sound identifiers and the sound samples;
[0149] The cloned sound identifier is generated in the manner described in the disclosed embodiment, S100'.
[0150] Based on and Figure 4 The same principle, Figure 8 This invention discloses a cloned sound generation apparatus 80 according to a sixth embodiment of the present disclosure, the apparatus comprising:
[0151] Acquisition module 801 is used to acquire user voice samples and send them to the server;
[0152] The server processes the sound samples according to the method of the first publicly disclosed embodiment.
[0153] In some examples, the device also includes:
[0154] The environmental monitoring module is used for:
[0155] Noise detection is performed on the user's sound collection environment;
[0156] Based on the noise detection results, a prompt for user voice collection is generated.
[0157] In some examples, the acquisition module is specifically used for:
[0158] Voiceprint samples and other sound samples;
[0159] Collect user voice samples, including:
[0160] Collect voiceprint samples of the user reading the standard text aloud, according to the preset standard text.
[0161] Collect other sound samples from the user according to the preset duration.
[0162] In some examples, the device also includes:
[0163] The sound processing module is used to: perform quality processing on sound samples to obtain processed sound samples;
[0164] Quality processing includes at least one of the following: noise reduction, volume normalization, and text alignment.
[0165] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0166] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, a computer program product, and an autonomous vehicle.
[0167] The computer instructions stored in a non-transitory computer-readable storage medium are used to cause the computer to perform the method described above.
[0168] Computer program products include computer programs that, when executed by a processor, implement the methods described above.
[0169] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0170] like Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 902 or a computer program loaded from a storage unit into random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. An input / output (I / O) interface 905 is also connected to bus 904.
[0171] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0172] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the cloned sound generation method and the cloned sound application method. For example, in some embodiments, the cloned sound generation method and the cloned sound application method can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the cloned sound generation method and the cloned sound application method described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured in any other suitable manner (e.g., by means of firmware) as a cloned sound generation method and a cloned sound application method.
[0173] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0174] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0175] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0176] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0177] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0178] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0179] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0180] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating cloned sounds, wherein, When applied on the server side, the method includes: In response to a cloned sound synthesis request sent by the application client, the cloned sound identifier and the content to be synthesized carried in the cloned sound synthesis request are obtained, and the corresponding sound clone model is determined according to the cloned sound identifier. Based on the content to be synthesized and the sound cloning model, cloned sound synthesis data corresponding to the content to be synthesized is obtained, and the cloned sound synthesis data is returned to the application client, wherein the application client is configured to receive and apply the cloned sound synthesis data; The sound cloning model is obtained in the following way: The user's voice samples collected and sent by the acquisition client are used as training data and input into the initial voice cloning model for training to obtain a trained voice cloning model corresponding to the user's voice samples. The voice cloning model is used to output cloned sound synthesis data with a similar timbre to the user's voice samples.
2. The method according to claim 1, wherein, The cloned sound identifier is obtained in the following way: After obtaining the trained sound cloning model, a clone sound identifier is generated for the sound cloning model, and the clone sound identifier and the mapping relationship between the clone sound identifier and the sound sample are sent to the application client.
3. The method according to claim 1, wherein, The step of using user voice samples collected and sent by the client as training data and inputting them into the initial voice cloning model for training to obtain a trained voice cloning model corresponding to the user's voice samples includes: The user's voice samples collected and sent by the client are used as training data and input into the initial voice cloning model for training. During training, the sound samples are identified to obtain sound adjustment parameters. The model parameters of the initial sound clone model are adjusted according to the sound adjustment parameters to obtain a trained sound clone model corresponding to the user's sound sample.
4. The method according to claim 3, wherein, The voice adjustment parameters include at least one of the following: keyword parameters of the voice sample, emotional feature parameters of the voice sample, and accent feature parameters of the voice sample.
5. The method according to claim 3 or 4, wherein, During the training process, after identifying the sound samples to obtain sound association parameters, adjusting the model parameters of the initial sound cloning model based on the sound association parameters, and obtaining a trained sound cloning model corresponding to the user's sound samples, the method further includes using the user's sound samples collected and sent by the client as training data and inputting them into the initial sound cloning model for training to obtain a trained sound cloning model corresponding to the user's sound samples. Extract any segment of the sound sample to obtain the test content text; The test content text is input into the sound cloning model to obtain the output cloned sound synthesis data corresponding to the test content text; Based on the cloned sound synthesis data corresponding to the test content text and the sound sample corresponding to the test content text, a similarity verification is performed. If the verification passes, the sound clone model is determined to be compliant.
6. The method according to any one of claims 1-4, wherein, The content to be synthesized is in the form of speech. The process of obtaining cloned sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound cloning model includes: The speech-to-text method converts the content to be synthesized from speech to text. The text-to-speech content is input into the sound cloning model, and the cloned sound synthesis data corresponding to the content to be synthesized is obtained by text-to-speech conversion.
7. The method according to any one of claims 1-4, wherein, The cloned sound synthesis data is in streaming data format or audio packet format.
8. The method according to any one of claims 1-4, wherein, After obtaining the cloned sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound cloning model, and before returning the cloned sound synthesis data to the application client, the method further includes: The sound effect parameters are obtained based on the content to be synthesized, and the cloned sound synthesis data is adjusted based on the sound effect parameters.
9. The method according to any one of claims 1-4, wherein, The data collection client is a client that adopts a microservice architecture; the application client is a client that adopts a global wide area network.
10. A method for applying cloned sounds, wherein, When applied to an application client, the method includes: Based on the content to be synthesized and the clone sound identifier of the sound clone model corresponding to the required clone sound, a clone sound synthesis request is generated and sent to the server. The server is configured to, in response to the application of the clone sound synthesis request, obtain the clone sound identifier and the content to be synthesized carried in the clone sound synthesis request, determine the corresponding sound clone model according to the clone sound identifier, obtain the clone sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound clone model, and return the clone sound synthesis data. Receive the cloned sound synthesis data returned by the server, and apply the cloned sound synthesis data; The sound cloning model is obtained in the following way: The user's voice samples collected and sent by the acquisition client are used as training data and input into the initial voice cloning model for training to obtain a trained voice cloning model corresponding to the user's voice samples. The voice cloning model is used to output cloned sound synthesis data with a similar timbre to the user's voice samples.
11. The method according to claim 10, wherein, Before generating a clone sound synthesis request based on the clone sound identifier of the sound clone model corresponding to the content to be synthesized and the required clone sound, and sending it to the server, the method further includes: Receive and save the cloned sound identifier sent by the server, as well as the mapping relationship between the cloned sound identifier and the sound sample; The cloned sound identifier is generated in accordance with the manner described in claim 2.
12. The method according to claim 10 or 11, wherein, The application client is a client used for live streaming; The application of the cloned sound synthesis data includes: The cloned audio synthesis data is converted into live stream data and then applied during the live stream.
13. A method for generating cloned sounds, wherein, When applied to a data acquisition client, the method includes: Collect user voice samples and send them to the server; The server processes the sound sample according to any one of claims 1-9.
14. The method according to claim 13, wherein, Before collecting user voice samples and sending them to the server, the method further includes: Noise detection is performed on the user's sound acquisition environment; Based on the noise detection results, a prompt for user voice collection is generated.
15. The method according to claim 13, wherein, The sound samples include: voiceprint samples and other sound samples; The collection of user voice samples includes: According to the preset standard text, collect voiceprint samples of the user reading the standard text aloud; Collect other sound samples from the user according to the preset duration.
16. The method according to claim 13, wherein, After collecting the user's voice samples and before sending them to the server, the method further includes: The sound sample is subjected to quality processing to obtain the processed sound sample; The quality processing includes at least one of the following: noise reduction processing, volume normalization processing, and text alignment processing.
17. A device for generating cloned sounds, wherein, The device includes: The model determination module is used to respond to the cloned sound synthesis request sent by the application client, obtain the cloned sound identifier and the content to be synthesized carried in the cloned sound synthesis request, and determine the corresponding sound cloned model according to the cloned sound identifier. A sound synthesis module is used to obtain cloned sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound cloning model, and return the cloned sound synthesis data to the application client, wherein the application client is configured to receive and apply the cloned sound synthesis data; The device further includes a model acquisition module, used to obtain the sound clone model in the following manner: The user's voice samples collected and sent by the acquisition client are used as training data and input into the initial voice cloning model for training to obtain a trained voice cloning model corresponding to the user's voice samples. The voice cloning model is used to output cloned sound synthesis data with a similar timbre to the user's voice samples.
18. The apparatus according to claim 17, wherein, The device further includes: an identifier generation module, configured to obtain the cloned sound identifier according to the following method: After obtaining the trained sound cloning model, a clone sound identifier is generated for the sound cloning model, and the clone sound identifier and the mapping relationship between the clone sound identifier and the sound sample are sent to the application client.
19. The apparatus according to claim 17, wherein, The model acquisition module is specifically used for: The user's voice samples collected and sent by the client are used as training data and input into the initial voice cloning model for training. During training, the sound samples are identified to obtain sound adjustment parameters. The model parameters of the initial sound clone model are adjusted according to the sound adjustment parameters to obtain a trained sound clone model corresponding to the user's sound sample.
20. The apparatus according to claim 19, wherein, The voice adjustment parameters include at least one of the following: keyword parameters of the voice sample, emotional feature parameters of the voice sample, and accent feature parameters of the voice sample.
21. The apparatus according to claim 19 or 20, wherein, The model acquisition module is also used for: Extract any segment of the sound sample to obtain the test content text; The test content text is input into the sound cloning model to obtain the output cloned sound synthesis data corresponding to the test content text; Based on the cloned sound synthesis data corresponding to the test content text and the sound sample corresponding to the test content text, a similarity verification is performed. If the verification passes, the sound clone model is determined to be compliant.
22. The apparatus according to any one of claims 17-20, wherein, The content to be synthesized is in the form of speech. The sound synthesis module is specifically used for: The speech-to-text method converts the content to be synthesized from speech to text. The text-to-speech content is input into the sound cloning model, and the cloned sound synthesis data corresponding to the content to be synthesized is obtained by text-to-speech conversion.
23. The apparatus according to any one of claims 17-20, wherein, The cloned sound synthesis data is in streaming data format or audio packet format.
24. The apparatus according to any one of claims 17-20, wherein, The device further includes: The effect adjustment module is used to obtain sound effect parameters based on the content to be synthesized, and to adjust the cloned sound synthesis data based on the sound effect parameters.
25. A device for applying cloned sounds, wherein, The device includes: The request module is used to generate a clone sound synthesis request based on the content to be synthesized and the clone sound identifier of the sound clone model corresponding to the required clone sound, and send it to the server. The server is configured to, in response to the application of the clone sound synthesis request, obtain the clone sound identifier and the content to be synthesized carried in the clone sound synthesis request, determine the corresponding sound clone model according to the clone sound identifier, obtain the clone sound synthesis data corresponding to the content to be synthesized based on the content to be synthesized and the sound clone model, and return the clone sound synthesis data. An application module is used to receive the cloned sound synthesis data returned by the server and apply the cloned sound synthesis data. The sound cloning model is obtained in the following way: The user's voice samples collected and sent by the acquisition client are used as training data and input into the initial voice cloning model for training to obtain a trained voice cloning model corresponding to the user's voice samples. The voice cloning model is used to output cloned sound synthesis data with a similar timbre to the user's voice samples.
26. The apparatus according to claim 25, wherein, The device further includes: The identifier storage module is used to receive and store the cloned sound identifier sent by the server and the mapping relationship between the cloned sound identifier and the sound sample; The cloned sound identifier is generated in accordance with the manner described in claim 2.
27. A device for generating cloned sounds, wherein, The device includes: The acquisition module is used to collect user voice samples and send them to the server. The server processes the sound sample according to any one of claims 1-9.
28. The apparatus according to claim 27, wherein, The device also includes: The environmental monitoring module is used for: Noise detection is performed on the user's sound acquisition environment; Based on the noise detection results, a prompt for user voice collection is generated.
29. The apparatus according to claim 27, wherein, The acquisition module is specifically used for: Voiceprint samples and other sound samples; The collection of user voice samples includes: According to the preset standard text, collect voiceprint samples of the user reading the standard text aloud; Collect other sound samples from the user according to the preset duration.
30. The apparatus according to claim 27, wherein, The device further includes: The sound processing module is used to: perform quality processing on the sound sample to obtain the processed sound sample; The quality processing includes at least one of the following: noise reduction processing, volume normalization processing, and text alignment processing.
31. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-9, and / or the method of any one of claims 10-12, and / or the method of any one of claims 13-16.
32. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-9, and / or the method according to any one of claims 10-12, and / or the method according to any one of claims 13-16.
33. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-9, and / or the method according to any one of claims 10-12, and / or the method according to any one of claims 13-16.
Citation Information
Patent Citations
Voice cloning model generation method and device and electronic equipment
CN115831088A
KR20210087098A