Speech Generation Method, Apparatus, Electronic Device, and Storage Medium
Through the combination of voiceprint feature model and vocoder, personalized voice signals are generated, which solves the problems of complex recording and error risks in the existing technology, and realizes efficient personalized voice synthesis and real-time processing.
Patent Information
- Application Number
- CN202210060611.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-01-19
AI Technical Summary
Existing personalized text-to-speech synthesis techniques require a large amount of training data and high-quality recording, resulting in complex recording processes and risk of recording errors.
The vocalprint feature model is used to extract the reference audio signal features, combine the similarity and naturalness parameters of user requirements, and generate target vocalprint feature vectors, and process the text feature vectors through the vocoder to generate personalized voice signals.
It reduces the difficulty of voice synthesis, avoids recording errors, can generate personalized voice with similar styles, meets user needs, and reduces the cost of deploying models. It is suitable for real-time voice processing and AI follow-up scenarios.
Smart Images

Figure CN114387945B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a voice generation method, apparatus, electronic device, and storage medium. Background Art
[0002] Currently, personalized text-to-speech (TTS) synthesis relies on a large amount of training data, approximately more than 50 sentences, and has relatively high requirements for data quality. It is necessary to record voice according to the specified text, and the recorded voice must be exactly the same as the text without errors, resulting in repeated modifications during recording sometimes, increasing the difficulty of voice synthesis and posing a risk of recording mistakes. Summary of the Invention
[0003] Embodiments of this application provide a voice generation method, apparatus, electronic device, and storage medium. Using the voice generation method provided by the embodiments of this application is conducive to reducing the difficulty of voice synthesis and there is no risk of audio recording mistakes.
[0004] In a first aspect, embodiments of this application provide a voice generation method, including:
[0005] Obtain text data, a reference audio signal, a first parameter, and a second parameter input by a user, where the first parameter is used to characterize the similarity of the user's requirements, and the second parameter is used to characterize the naturalness of the user's requirements;
[0006] Use a voiceprint feature model to extract features from the reference audio signal to obtain a voiceprint feature vector of the reference audio signal;
[0007] Obtain a target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal;
[0008] Extract features from the text data to obtain a text feature vector;
[0009] Obtain a voice spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector;
[0010] Use a vocoder to process the voice spectrum to obtain a target audio signal, and the text data represented by the target audio signal is the text data input by the user.
[0011] In a second aspect, embodiments of this application provide a voice generation apparatus, including: an acquisition unit, a feature extraction unit, a determination unit, and a processing unit;
[0012] The acquisition unit acquires text data, a reference audio signal, a first parameter, and a second parameter input by a user, where the first parameter is used to characterize the similarity of the user's requirements, and the second parameter is used to characterize the naturalness of the user's requirements;
[0013] A feature extraction unit, configured to extract features from a reference audio signal by using a voiceprint feature model to obtain a voiceprint feature vector of the reference audio signal;
[0014] A determination unit, configured to obtain a target voiceprint feature vector according to a first parameter, a second parameter, and the voiceprint feature vector of the reference audio signal;
[0015] The feature extraction unit is further configured to extract features from text data to obtain a text feature vector;
[0016] The determination unit is further configured to obtain a speech spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector;
[0017] A processing unit, configured to process the speech spectrum by using a vocoder to obtain a target audio signal, and the text data represented by the target audio signal is the text data input by the user.
[0018] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, the processor is connected to a memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in the first aspect.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and the computer program enables a computer to execute the method described in the first aspect.
[0020] In a fifth aspect, an embodiment of the present application provides a computer program product, the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer is operable to enable the computer to execute the method described in the first aspect.
[0021] Implementing the embodiments of the present application has the following beneficial effects:
[0022] After obtaining the user's similarity requirement and naturalness requirement for the finally obtained target audio signal, the voiceprint feature model is used to extract features from the reference audio signal to obtain the voiceprint feature vector of the reference audio signal; the voiceprint feature vector of the reference audio signal is processed according to the user's similarity requirement and naturalness requirement for the finally obtained target audio signal to obtain the target voiceprint feature vector; the speech spectrum corresponding to the text data is obtained according to the text feature vector of the input text data and the target voiceprint feature vector; the speech spectrum is processed by a vocoder to obtain the target audio signal. It can be seen that the solution of this application can generate voices with similar styles based on the user's similarity requirement and naturalness requirement, meet the user's personalized needs, and can achieve a good compromise between naturalness and similarity; and adopting the solution of this application does not require training a separate model for each user, and based on a model of this application (this model includes a voiceprint feature model, a speech synthesis model for synthesizing text feature vectors and voiceprint feature vectors, and a vocoder), personalized audio signals can be generated according to the needs of different users, reducing the cost of deploying the model; and because the difficulty of deploying the model is low, real-time speech processing can be realized; it can be directly applied to AI follow-up shooting or other scenarios, which can greatly reduce the time cost of user recording and video production. All in all, adopting the solution of this application is beneficial to reducing the difficulty of speech synthesis, and because there is no need to repeatedly record audio, there is also no risk of audio recording errors. Description of the Drawings
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 Schematic diagram of a speech generation system provided by an embodiment of the present application;
[0025] Figure 2 Flow chart of a speech generation method provided by an embodiment of the present application;
[0026] Figure 3 Flow chart of another speech generation method provided by an embodiment of the present application;
[0027] Figure 4 Functional unit composition block diagram of a speech generation device provided by an embodiment of the present application;
[0028] Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0029] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0030] The terms "first", "second", "third", "fourth", etc. in the specification and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0031] Referring to "embodiments" herein means that specific features, results, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0032] The embodiments of the present application can acquire, extract features, and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0033] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0034] The method of the present application can be applied to terminal devices, such as smart phones, tablets, smart bracelets, etc., and can also be applied to a Figure 1 voice generation system as shown. As Figure 1 shown, the voice generation system includes a terminal device 102 and a voice generation server 101;
[0035] The terminal device 102 sends a voice generation request to the voice generation server 101. The voice generation request carries the text data, the first parameter, and the second parameter input by the user. Optionally, the voice generation request also carries a reference voice signal. In one example, the voice generation server 101 pre-stores the reference voice signal. After receiving the voice generation request, the voice generation server 101 obtains the target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal. Optionally, the voice generation server 101 also extracts features from the reference audio signal by using a voiceprint feature model to obtain the voiceprint feature vector of the reference audio signal. The voice generation server 101 extracts features from the text data to obtain a text feature vector. The voice spectrum corresponding to the text data is obtained according to the text feature vector and the voiceprint feature vector of the reference audio signal. The voice spectrum is processed by using a vocoder to obtain a target audio signal, and the text data identified by the target audio signal is the text data input by the user. The voice generation server 101 sends a response message for responding to the voice generation request to the terminal device 102, and the response message carries the above target audio signal.
[0036] It can be seen that by adopting the solution of the present application, it is not necessary to train a separate model for each user. The present application uses a model that can be applied to multiple people, and different personalized audio signals can be obtained according to the needs of different users, reducing the cost of deploying the model. And because the difficulty of deploying the model is low, real-time voice processing can be realized. Through the method of the present application, the conversion of the user's personalized style voice can be realized with one sentence, and the voice similar to the user's style can be generated, which can meet the user's personalized needs, and can achieve a good compromise in terms of naturalness and similarity, and can be directly applied to AI follow-up shooting or other scenarios, which can greatly reduce the time cost of user recording and video production.
[0037] See Figure 2 , Figure 2 is a schematic flowchart of a voice generation method provided by an embodiment of the present application. This method is applied to a voice generation device, and the voice generation device can be the above terminal device or Figure 1 the voice generation server 101 shown in
[0038] 201: The voice generation device obtains the text data, the reference audio signal, the first parameter, and the second parameter input by the user.
[0039] Among them, the first parameter is used to characterize the similarity of the user's needs, and the second parameter is used to characterize the naturalness of the user's needs.
[0040] It should be noted that the similarity of user requirements refers to the similarity between the voice signal that the user hopes to generate using the method of the present application and the reference audio signal, and the naturalness of user requirements refers to the naturalness of the voice signal that the user hopes to generate using the method of the present application.
[0041] Optionally, the reference audio signal can be collected by the voice generation device, and can be an audio signal pre-stored in the voice generation device; it can also be an audio signal obtained by the voice generation device from other devices after being collected by other devices.
[0042] Optionally, the reference voice signal can be a Chinese voice signal, an English voice signal, a French voice signal or a voice signal of other languages.
[0043] Optionally, the above reference audio signal is of the above specified speaker, and the specified speaker can be the above user or other users, such as the user of the voice generation device.
[0044] Among them, the duration of the reference voice signal is a preset duration; optionally, the preset duration can be 2s, 5s, 8s or other durations.
[0045] 202: The voice generation device extracts features from the reference audio signal using the voiceprint feature model to obtain the voiceprint feature vector of the reference audio signal.
[0046] In an embodiment of the present application, before extracting features from the reference audio signal using the voiceprint feature model, the voiceprint feature model is trained based on a training set.
[0047] Among them, the training set includes audio signals of multiple sample users, and the audio signal of each sample user is an audio signal of a certain duration; that is, the audio signal of each sample user is the audio signal collected for the sample user when speaking at least one sentence; the audio signal of each sample user includes audio signals corresponding to at least one sentence respectively. For example, the training set includes audio signals of 1000 sample users, and for each sample user, the audio signal is an audio signal with a duration of more than 1000 hours.
[0048] Exemplarily, the audio signals in the above training set include foreign language audio signals, such as English audio signals, French audio signals, etc., and can also be Chinese audio signals, which are not limited herein.
[0049] Exemplarily, the audio signal of each sample user can be collected in a quiet environment or a noisy environment for the sample user, that is, the audio signal of each sample user can be a noisy audio signal or an audio signal without noise.
[0050] It should be noted here that the above training set is an unlabeled training set. The advantage of using an unlabeled training set is that unlabeled data is easier to obtain than labeled data, and for the audio signal of each sample user, the feature vector extracted based on the audio signal of the sample user can better represent the features of the sample user.
[0051] 203: The voice generation device obtains the target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal.
[0052] In an embodiment of the present application, obtaining the target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal includes:
[0053] When the first parameter indicates that the similarity of the user's demand is higher than the preset similarity, and the second parameter characterizes that the naturalness of the user's demand is lower than the preset naturalness, that is, the user hopes that the similarity between the voice signal generated by the method of the present application and the reference audio signal is high, but the user has low requirements for the naturalness of the voice signal generated by the method of the present application. At this time, the voiceprint feature vector of the reference audio signal can be directly determined as the target voiceprint feature vector.
[0054] In an embodiment of the present application, obtaining the target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal includes:
[0055] When the first parameter indicates that the similarity of the user's demand is lower than the preset similarity, and the second parameter characterizes that the naturalness of the user's demand is higher than the preset naturalness, the average voiceprint feature vector of each sample user among the M sample users is obtained according to the audio data of the M sample users in the training set for training the voiceprint feature model; M is an integer greater than 1; the similarity between the voiceprint feature vector of the reference audio signal and the average voiceprint feature vector of each sample user among the M sample users is calculated; the voiceprint feature vector of the target sample user is determined as the target voiceprint feature vector, and the target sample user is the sample user among the M sample users with the highest similarity between the voiceprint feature vector and the voiceprint feature vector of the reference audio signal.
[0056] Optionally, the audio data of the above M sample users can be the audio data of all sample users in the above training set. At this time, the average voiceprint feature vector closest to the voiceprint feature vector of the reference audio signal can be found in the training set; in order to avoid traversing in the training set and improve the calculation efficiency, the audio data of the above M sample users can also be the audio data of some sample users in the above training set.
[0057] Further, the audio signals of the M sample users include the audio signals of at least one sentence of each of the M sample users. The average voiceprint feature vectors of each of the M sample users are obtained according to the audio data of the M sample users in the training set for training the voiceprint feature model, including:
[0058] Feature extraction is respectively performed on the audio signals of each sentence in at least one sentence of each sample user to obtain at least one voiceprint feature vector respectively corresponding to at least one sentence of each sample user; the at least one voiceprint feature vector respectively corresponding to at least one sentence of each sample user is averaged to obtain the average voiceprint feature vector of each sample user.
[0059] Further, the voiceprint feature vector of the target sample user is the average voiceprint feature vector of the target sample user, or
[0060] The method of this application further includes:
[0061] Calculate the similarity between the voiceprint feature vector of the reference audio signal and the voiceprint feature vectors corresponding to each sentence in at least one sentence of the target sample user, and determine the voiceprint feature vector of the sentence with the highest similarity to the voiceprint feature vector of the reference audio signal as the voiceprint feature vector of the target sample user.
[0062] Specifically, the audio data of the M sample users includes the audio data of at least one sentence of each of the M sample users; for the audio data of each of the M sample users, the following operations are performed:
[0063] Feature extraction is performed on the audio data corresponding to each sentence in at least one sentence of the sample user to obtain at least one voiceprint feature vector respectively corresponding to at least one sentence; the at least one voiceprint feature vector is processed, such as averaging or weighted averaging, to obtain the average voiceprint feature vector of the sample user.
[0064] According to the above method, the average voiceprint feature vectors of each of the M sample users can be obtained; then calculate the similarity between the voiceprint feature vector of the reference audio signal and the average voiceprint feature vectors of each of the M sample users; then determine the sample user with the highest similarity between the average voiceprint feature vector and the voiceprint feature vector of the reference audio signal among the M sample users. This sample user is the sample user among the M sample users with the most similar voiceprint to the specified speaker, and this user is the above target sample user; the average voiceprint feature vector of the target sample user can be determined as the target voiceprint feature vector, or the above target voiceprint feature vector can be obtained according to the following method:
[0065] Calculate the similarity between the voiceprint feature vector of the reference audio signal and the voiceprint feature vectors corresponding to each sentence in at least one sentence of the target sample user; the voiceprint feature vector of the sentence with the highest similarity to the voiceprint feature vector of the reference audio signal is determined as the voiceprint feature vector of the target sample user. Using this method to obtain the voiceprint feature vector of the target sample user can be understood as selecting the best from the best, and finally, the sample user with the voiceprint feature vector closest to the specified speaker can be selected from M sample users.
[0066] For example, assume that the M sample users include sample user A, sample user B, and sample user C. The audio data of sample user A includes the audio data of 3 sentences, the audio data of sample user B includes the audio data of 4 sentences, and the audio data of sample user C includes the audio data of 5 sentences; perform feature extraction on the audio data of each of the 3 sentences of sample user A to obtain 3 voiceprint feature vectors corresponding to the 3 sentences respectively; perform an averaging operation on the 3 voiceprint feature vectors to obtain the average voiceprint feature vector of sample user A; perform feature extraction on the audio data of each of the 4 sentences of sample user B to obtain 4 voiceprint feature vectors corresponding to the 4 sentences respectively; perform an averaging operation on the 4 voiceprint feature vectors to obtain the average voiceprint feature vector of sample user B; perform feature extraction on the audio data of each of the 5 sentences of sample user C to obtain 5 voiceprint feature vectors corresponding to the 5 sentences respectively; perform an averaging operation on the 5 voiceprint feature vectors to obtain the average voiceprint feature vector of sample user C; calculate the similarity between the voiceprint feature vector of the reference audio signal and the average voiceprint feature vector of each sample user among the 3 sample users (i.e., sample user A, sample user B, and sample user C); assume that the average voiceprint feature vector of sample user C has the highest similarity to the voiceprint feature vector of the reference audio signal, and the average voiceprint feature vector of sample user C can be confirmed as the above-mentioned target voiceprint feature vector; or calculate the similarity between the voiceprint feature vector of the reference audio signal and the voiceprint feature vectors corresponding to each of the 5 sentences of sample user C; assume that the voiceprint feature vector of the 3rd sentence among the 5 sentences has a higher similarity to the voiceprint feature vector of the reference audio signal, and the voiceprint feature vector of the 3rd sentence among the 5 sentences of sample user C is determined as the voiceprint feature vector of the target sample user.
[0067] It should be pointed out here that calculating the similarity between two voiceprint feature vectors in this application specifically refers to calculating the Euclidean distance between the two voiceprint feature vectors. Among them, the smaller the Euclidean distance, the higher the similarity between the two voiceprint feature vectors; the larger the Euclidean distance, the lower the similarity between the two voiceprint feature vectors.
[0068] 204: The voice generation device performs feature extraction on the text data to obtain a text feature vector.
[0069] Specifically, the speech generation device performs a word segmentation operation on the text data to obtain multiple phrases. Specifically, the word segmentation operation on the text data can be performed by a dictionary-based word segmentation algorithm or a statistical machine learning algorithm. Among them, the dictionary-based word segmentation algorithms include the forward maximum matching method, the backward maximum matching method, and the bidirectional matching word segmentation method, etc. The statistical machine learning algorithms include the hidden Markov model (HMM) algorithm, the conditional random fields (CRF) algorithm, and the support vector machine (SVM) algorithm, etc. Each phrase among the multiple phrases is encoded to obtain a feature vector for each phrase. The multiple feature vectors corresponding to the multiple phrases are fused to obtain a text feature vector. Or,
[0070] After obtaining the multiple phrases, determine the part of speech of each phrase among the multiple phrases. Each phrase among the multiple phrases and its part of speech are encoded to obtain a feature vector for each phrase. The multiple feature vectors corresponding to the multiple phrases are fused to obtain a text feature vector.
[0071] 205: The speech generation device obtains the speech spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector.
[0072] In an embodiment of the present application, obtaining the speech spectrum signal corresponding to the text data according to the text feature vector and the target voiceprint feature vector includes:
[0073] The text feature vector and the target voiceprint feature vector are concatenated to obtain a concatenated feature vector. The speech spectrum corresponding to the text data is obtained according to the concatenated feature vector.
[0074] Specifically, the number of dimensions of the text feature vector is the same as the number of dimensions of the target voiceprint feature vector. The text feature vector and the target voiceprint feature vector are concatenated to obtain a concatenated feature vector. The number of dimensions of the concatenated feature vector is the sum of the number of dimensions of the text feature vector and the number of dimensions of the target voiceprint feature vector. For example, both the text feature vector and the voiceprint feature vector are 256-dimensional. The vector obtained by concatenating the text feature vector and the target voiceprint feature vector is 512-dimensional. Then this 512-dimensional vector is processed through a decoding model to obtain the speech spectrum corresponding to the above text data.
[0075] Among them, the above process can be regarded as a speech synthesis process and can be implemented by a speech synthesis model. Optionally, the speech synthesis model can be implemented by a tacotron2 network.
[0076] 206: The voice generation device processes the voice spectrum corresponding to the text data using a vocoder to obtain a target audio signal.
[0077] Among them, the text data represented by the target audio signal is the text data input by the user.
[0078] It should be noted here that before using the voice synthesis model and the vocoder, it is necessary to train the voice synthesis model and the vocoder, and the training of the voice synthesis model and the vocoder can be carried out in a joint training manner. The training data is a training set of labeled data. Specifically, the training data includes the audio data of multiple sample users and the text data corresponding to the audio data.
[0079] The solution of this application can be applied to the AI follow-up scenario or the video recording scenario; after obtaining the video, it is necessary to synthesize an audio signal for the video, and this audio signal can be obtained by using the solution of this application, without the need for professional recording equipment, which can greatly reduce the time cost of users' recording and video production.
[0080] It can be seen that after obtaining the similarity requirement and naturalness requirement of the user for the finally obtained target audio signal, the voiceprint feature model is used to extract the features of the reference audio signal to obtain the voiceprint feature vector of the reference audio signal; the voiceprint feature vector of the reference audio signal is processed according to the similarity requirement and naturalness requirement of the user for the finally obtained target audio signal to obtain the target voiceprint feature vector; the voice spectrum corresponding to the text data is obtained according to the text feature vector of the input text data and the target voiceprint feature vector; the voice spectrum is processed using a vocoder to obtain a target audio signal. It can be seen that by adopting the solution of this application, voice for similar styles can be generated based on the similarity requirement and naturalness requirement of the user, meeting the personalized needs of the user, and being able to achieve a good compromise in terms of naturalness and similarity; and by adopting the solution of this application, it is not necessary to train a separate model for each user, and based on a model of this application (this model includes a voiceprint feature model, a voice synthesis model for synthesizing text feature vectors and voiceprint feature vectors, and a vocoder), personalized audio signals can be generated according to the needs of different users, reducing the cost of deploying the model; and because the difficulty of deploying the model is low, real-time voice processing can be achieved; it can be directly applied to AI follow-up or other scenarios, which can greatly reduce the time cost of users' recording and video production. All in all, adopting the solution of this application is beneficial to reducing the difficulty of voice synthesis, and because there is no need to repeatedly record audio, there is also no risk of audio recording errors.
[0081] Refer to Figure 3 , Figure 3 The flowchart of another voice generation method provided by an embodiment of this application. This method is applied to the above voice generation device, and in this embodiment, it is the same asFigure 2 The same content as that in the illustrated embodiment will not be described again here. The method of this embodiment includes the following steps:
[0082] 301. The voice generation device acquires a voiceprint feature model, a speech synthesis model, and a vocoder.
[0083] Among them, the voiceprint feature model is implemented based on a neural network, and the neural network can be a recurrent neural network, a fully connected neural network, or other types of neural networks.
[0084] Specifically, it can be trained by the voice generation device itself, or after the voiceprint feature model, the speech synthesis model, and the vocoder are trained on other devices, the voice generation device acquires the voiceprint feature model, the speech synthesis model, and the vocoder from other devices.
[0085] For the voiceprint feature model, it can be trained in the following manner:
[0086] Acquire a training set, which includes audio signals of multiple sample users and the voiceprint feature vectors corresponding to the audio signals. The audio signal of each sample user is an audio signal of a certain duration; that is, the audio signal of each sample user is the audio signal collected when the sample user speaks at least one sentence; the audio signal of each sample user includes audio signals corresponding to at least one sentence respectively. For example, the training set includes audio signals of 1000 sample users, and for the audio signal of each sample user, it is an audio signal with a duration of more than 1000 hours; input the audio signal of the sample user into the neural network for processing to obtain a predicted voiceprint feature vector; input the predicted voiceprint feature vector and the voiceprint feature vector corresponding to the audio signal of the sample user in the training set into a loss function for calculation to obtain a loss value; adjust the parameters in the neural network based on the loss value to obtain an adjusted neural network; then input the audio signal of another sample user into the neural network for processing to obtain a predicted voiceprint feature vector; input the predicted voiceprint feature vector and the voiceprint feature vector corresponding to the audio signal of the sample user in the training set into a loss function for calculation to obtain a loss value; adjust the parameters in the neural network based on the loss value to obtain an adjusted neural network; repeat the above steps until the loss value converges or the number of training times reaches a preset number; when the loss value converges or the number of training times reaches a preset number, use the adjusted neural network as the voiceprint feature model.
[0087] For the speech synthesis model and the vocoder, they can be trained in the above manner and will not be described here.
[0088] 302. The voice generation device acquires the text data, the reference speech signal, the first parameter, and the second parameter input by the user.
[0089] Among them, the first parameter is used to represent the user's similarity requirement for the finally generated audio signal, and the second parameter is used to represent the user's naturalness requirement for the finally generated audio signal.
[0090] 303. The voice generation device extracts features from the reference audio signal by using the voiceprint feature model to obtain the voiceprint feature vector of the reference audio signal.
[0091] 304. When the first parameter indicates that the similarity required by the user is higher than the preset similarity, and the second parameter represents that the naturalness required by the user is lower than the preset naturalness, the voice generation device uses the voiceprint feature vector of the reference voice signal as the target dimensionality-increased feature vector.
[0092] 305. When the first parameter indicates that the similarity required by the user is lower than the preset similarity, and the second parameter represents that the naturalness required by the user is higher than the preset naturalness, the voice generation device obtains the target voiceprint feature vector according to the audio data of the user samples in the training set.
[0093] It should be noted here that for the specific implementation process of step 305, reference can be made to the relevant description of step 203, which will not be elaborated here.
[0094] 306. The voice generation device extracts features from the text data to obtain the text feature vector.
[0095] 307. The voice generation device obtains the voice spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector.
[0096] 308. The voice generation device processes the voice spectrum corresponding to the text data by using the vocoder to obtain the target audio signal.
[0097] It should be noted here that for the specific implementation process of steps 306-308, reference can be made to the relevant description of steps 204-206, which will not be elaborated here.
[0098] It can be seen that in the embodiments of the present application, after obtaining the user's similarity requirement and naturalness requirement for the finally obtained target audio signal, a voiceprint feature model is used to extract features from the reference audio signal to obtain the voiceprint feature vector of the reference audio signal; the voiceprint feature vector of the reference audio signal is processed according to the user's similarity requirement and naturalness requirement for the finally obtained target audio signal to obtain the target voiceprint feature vector; the voice spectrum corresponding to the text data is obtained according to the text feature vector of the input text data and the target voiceprint feature vector; and a vocoder is used to process the voice spectrum to obtain the target audio signal. It can be seen that the solution of the present application can generate voices with similar styles based on the user's similarity requirement and naturalness requirement, meet the user's personalized needs, and achieve a good compromise between naturalness and similarity; moreover, with the solution of the present application, it is not necessary to train a separate model for each user. Based on a model of the present application (this model includes a voiceprint feature model, a speech synthesis model for synthesizing text feature vectors and voiceprint feature vectors, and a vocoder), personalized audio signals can be generated according to the needs of different users, reducing the cost of deploying the model; and since the difficulty of deploying the model is low, real-time speech processing can be realized; it can be directly applied to AI follow-up shooting or other scenarios, which can greatly reduce the time cost of user recording and video production. All in all, the solution of the present application is beneficial to reducing the difficulty of speech synthesis, and since there is no need to repeatedly record audio, there is also no risk of audio recording errors.
[0099] Refer to Figure 4 , Figure 4 FIG. 4 is a block diagram of the functional units of a speech generation device provided by an embodiment of the present application. The speech generation device 400 includes: an acquisition unit 401, a feature extraction unit 402, a determination unit 403, and a processing unit 404;
[0100] The acquisition unit 401 acquires the text data, the reference audio signal, the first parameter, and the second parameter input by the user, where the first parameter is used to characterize the similarity of the user's requirement, and the second parameter is used to characterize the naturalness of the user's requirement;
[0101] The feature extraction unit 402 is configured to extract features from the reference audio signal by using a voiceprint feature model to obtain the voiceprint feature vector of the reference audio signal;
[0102] The determination unit 403 is configured to obtain the target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal;
[0103] The feature extraction unit 402 is further configured to extract features from the text data to obtain the text feature vector;
[0104] The determination unit 403 is further configured to obtain the speech spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector;
[0105] The processing unit 404 is configured to process the speech spectrum by using a vocoder to obtain a target audio signal, and the text data represented by the target audio signal is the text data input by the user.
[0106] In some embodiments of the present application, in terms of obtaining the target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal, the determination unit 403 is specifically configured to perform the following operations:
[0107] When the first parameter indicates that the similarity of the user's requirement is higher than the preset similarity, and the second parameter characterizes that the naturalness of the user's requirement is lower than the preset naturalness, the voiceprint feature vector of the reference audio signal is determined as the target voiceprint feature vector.
[0108] In some embodiments of the present application, in terms of obtaining the target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal, the determination unit 403 is further specifically configured to perform the following operations:
[0109] When the first parameter indicates that the similarity of the user's requirement is lower than the preset similarity, and the second parameter characterizes that the naturalness of the user's requirement is higher than the preset naturalness, the average voiceprint feature vector of each of the M sample users in the M sample users is obtained according to the audio data of the M sample users in the training set for training the voiceprint feature model; M is an integer greater than 1; calculate the similarity between the voiceprint feature vector of the reference audio signal and the average voiceprint feature vector of each of the M sample users; the voiceprint feature vector of the target sample user is determined as the target voiceprint feature vector, and the target sample user is the sample user with the highest similarity between the voiceprint feature vector and the voiceprint feature vector of the reference audio signal among the M sample users.
[0110] In some embodiments of the present application, the audio signals of the M sample users include the audio signals of at least one sentence of each of the M sample users. In terms of obtaining the average voiceprint feature vector of each of the M sample users according to the audio data of the M sample users in the training set for training the voiceprint feature model, the determination unit 403 is specifically configured to perform the following operations:
[0111] Feature extraction is respectively performed on the audio signals of each sentence in at least one sentence of each sample user to obtain at least one voiceprint feature vector respectively corresponding to at least one sentence of each sample user; the at least one voiceprint feature vector respectively corresponding to at least one sentence of each sample user is averaged to obtain the average voiceprint feature vector of each sample user.
[0112] In some embodiments of the present application, the voiceprint feature vector of the target sample user is the average voiceprint feature vector of the target sample user, or,
[0113] The determining unit 403 is further configured to perform the following operations:
[0114] Calculate the similarity between the voiceprint feature vector of the reference audio signal and the voiceprint feature vector corresponding to each sentence in at least one sentence of the target sample user, and determine the voiceprint feature vector of the sentence with the highest similarity to the voiceprint feature vector of the reference audio signal as the voiceprint feature vector of the target sample user.
[0115] In some embodiments of the present application, in terms of obtaining the speech spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector, the determining unit 403 is specifically configured to perform the following operations:
[0116] Concatenate the text feature vector and the target voiceprint feature vector to obtain a concatenated feature vector; obtain the speech spectrum according to the concatenated feature vector.
[0117] In some embodiments of the present application, the duration of the reference audio signal is a preset duration.
[0118] Refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown, the electronic device 500 includes a transceiver 501, a processor 502, and a memory 503. They are connected through a bus 504. The memory 503 is used to store computer programs and data, and can transmit the data stored in the memory 503 to the processor 502.
[0119] The processor 502 is configured to read the computer program in the memory 503 and perform the following operations:
[0120] Obtain the text data, reference audio signal, first parameter, and second parameter input by the user, where the first parameter is used to characterize the similarity of the user's requirements, and the second parameter is used to characterize the naturalness of the user's requirements; use the voiceprint feature model to extract features from the reference audio signal to obtain the voiceprint feature vector of the reference audio signal; obtain the target voiceprint feature vector according to the first parameter, second parameter, and the voiceprint feature vector of the reference audio signal; extract features from the text data to obtain the text feature vector; obtain the speech spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector; use the vocoder to process the speech spectrum to obtain the target audio signal, and the text data represented by the target audio signal is the text data input by the user.
[0121] In some embodiments of the present application, in terms of obtaining a target voiceprint feature vector based on a first parameter, a second parameter, and the voiceprint feature vector of a reference audio signal, the processor 502 is specifically configured to perform the following operations:
[0122] When the first parameter indicates that the similarity of the user's requirement is higher than a preset similarity, and the second parameter indicates that the naturalness of the user's requirement is lower than a preset naturalness, determine the voiceprint feature vector of the reference audio signal as the target voiceprint feature vector.
[0123] In some embodiments of the present application, in terms of obtaining a target voiceprint feature vector based on a first parameter, a second parameter, and the voiceprint feature vector of a reference audio signal, the processor 502 is further specifically configured to perform the following operations:
[0124] When the first parameter indicates that the similarity of the user's requirement is lower than a preset similarity, and the second parameter indicates that the naturalness of the user's requirement is higher than a preset naturalness, obtain the average voiceprint feature vector of each of the M sample users in the training set of M sample users used to train the voiceprint feature model; M is an integer greater than 1; calculate the similarity between the voiceprint feature vector of the reference audio signal and the average voiceprint feature vector of each of the M sample users; determine the voiceprint feature vector of the target sample user as the target voiceprint feature vector, where the target sample user is the sample user with the highest similarity between the voiceprint feature vector and the voiceprint feature vector of the reference audio signal among the M sample users.
[0125] In some embodiments of the present application, the audio signals of the M sample users include the audio signals of at least one sentence of each of the M sample users. In terms of obtaining the average voiceprint feature vector of each of the M sample users based on the audio data of the M sample users in the training set used to train the voiceprint feature model, the processor 502 is specifically configured to perform the following operations:
[0126] Extract features from the audio signals of each sentence in at least one sentence of each sample user to obtain at least one voiceprint feature vector corresponding to each sentence of at least one sentence of each sample user; average the at least one voiceprint feature vector corresponding to each sentence of at least one sentence of each sample user to obtain the average voiceprint feature vector of each sample user.
[0127] In some embodiments of the present application, the voiceprint feature vector of the target sample user is the average voiceprint feature vector of the target sample user, or,
[0128] The processor 502 is further configured to perform the following operations:
[0129] Calculate the similarity between the voiceprint feature vector of the reference audio signal and the voiceprint feature vectors corresponding to each sentence in at least one sentence of the target sample user, and determine the voiceprint feature vector of the sentence with the highest similarity to the voiceprint feature vector of the reference audio signal as the voiceprint feature vector of the target sample user.
[0130] In some embodiments of the present application, in terms of obtaining the speech spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector, the processor 502 is specifically configured to perform the following operations:
[0131] Concatenate the text feature vector and the target voiceprint feature vector to obtain a concatenated feature vector; obtain the speech spectrum according to the concatenated feature vector.
[0132] In some embodiments of the present application, the duration of the reference audio signal is a preset duration.
[0133] Specifically, the above-mentioned processor 502 may be Figure 4 the feature extraction unit 402, the determination unit 403, and the processing unit 404 of the voice generation device 400 in the above-described embodiment.
[0134] It should be understood that the electronic device in the present application may include a smart phone (such as an Android phone, an iOS phone, a Windows Phone phone, etc.), a tablet computer, a handheld computer, a notebook computer, a mobile Internet device MID (Mobile Internet Devices, abbreviated as: MID), or a wearable device, etc. The above-mentioned electronic devices are only examples and not exhaustive, including but not limited to the above-mentioned electronic devices. In practical applications, the above-mentioned electronic devices may also include: intelligent vehicle terminals, computer devices, and so on.
[0135] The embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement part or all of the steps of any one of the voice generation methods described in the above method embodiments.
[0136] The embodiment of the present application further provides a computer program product, and the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps of any one of the voice generation methods described in the above method embodiments.
[0137] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0138] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0139] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0140] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software program modules.
[0142] When the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0143] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory, and the memory can include: flash drives, read-only memories (abbreviation: ROM, English: Read-Only Memory), random access memories (abbreviation: RAM, English: Random Access Memory), magnetic disks, or optical discs, etc.
[0144] The above has introduced the embodiments of the present application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A voice generation method, characterized in that, Including: Obtain text data, a reference audio signal, a first parameter, and a second parameter input by a user. The first parameter is used to characterize the similarity of the user's demand, and the second parameter is used to characterize the naturalness of the user's demand. The similarity of the user's demand is used to indicate the similarity between the voice signal that the user hopes to generate and the reference audio signal, and the naturalness of the user's demand is used to indicate the naturalness degree of the voice signal that the user hopes to generate. Use a voiceprint feature model to extract features from the reference audio signal to obtain a voiceprint feature vector of the reference audio signal. Obtain a target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal. The obtaining method of the target voiceprint feature vector includes: when the first parameter indicates that the similarity of the user's demand is higher than a preset similarity, and the second parameter characterizes that the naturalness of the user's demand is lower than a preset naturalness, determine the voiceprint feature vector of the reference audio signal as the target voiceprint feature vector; when the first parameter indicates that the similarity of the user's demand is lower than the preset similarity, and the second parameter characterizes that the naturalness of the user's demand is higher than the preset naturalness, obtain the average voiceprint feature vector of each of the M sample users in the M sample users according to the audio data of the M sample users in the training set used to train the voiceprint feature model, where M is an integer greater than 1, calculate the similarity between the voiceprint feature vector of the reference audio signal and the average voiceprint feature vector of each of the M sample users, and determine the voiceprint feature vector of the target sample user as the target voiceprint feature vector, where the target sample user is the sample user with the highest similarity between the voiceprint feature vector and the voiceprint feature vector of the reference audio signal among the M sample users. Extract features from the text data to obtain a text feature vector. Obtain a voice spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector. Use a vocoder to process the voice spectrum to obtain a target audio signal, and the text data represented by the target audio signal is the text data input by the user.
2. The method according to claim 1, wherein The audio signals of the M sample users include the audio signals of at least one sentence of each of the M sample users. The obtaining of the average voiceprint feature vector of each of the M sample users according to the audio data of the M sample users in the training set used to train the voiceprint feature model includes: Extract features from the audio signals of each sentence in at least one sentence of each sample user respectively to obtain at least one voiceprint feature vector corresponding to each sentence in at least one sentence of each sample user. Average the at least one voiceprint feature vector corresponding to each sentence in at least one sentence of each sample user respectively to obtain the average voiceprint feature vector of each sample user.
3. The method according to claim 2, wherein The voiceprint feature vector of the target sample user is the average voiceprint feature vector of the target sample user, or The method further includes: Calculate the similarity between the voiceprint feature vector of the reference audio signal and the voiceprint feature vectors corresponding to each sentence in at least one sentence of the target sample user. Determine the voiceprint feature vector of the sentence with the highest similarity to the voiceprint feature vector of the reference audio signal as the voiceprint feature vector of the target sample user.
4. The method according to any one of claims 1 to 3, characterized in that, The obtaining the speech spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector includes: Concatenate the text feature vector and the target voiceprint feature vector to obtain a concatenated feature vector. Obtain the speech spectrum according to the concatenated feature vector.
5. The method according to any one of claims 1 to 3, characterized in that, The duration of the reference audio signal is a preset duration.
6. A voice generation device, characterized in that, Including: An acquisition unit, a feature extraction unit, a determination unit, and a processing unit. The acquisition unit acquires the text data, the reference audio signal, a first parameter, and a second parameter input by the user. The first parameter is used to characterize the similarity of the user's requirement, and the second parameter is used to characterize the naturalness of the user's requirement. The similarity of the user's requirement is used to indicate the similarity between the voice signal that the user hopes to generate and the reference audio signal, and the naturalness of the user's requirement is used to indicate the naturalness degree of the voice signal that the user hopes to generate. The feature extraction unit is used to extract features from the reference audio signal by using a voiceprint feature model to obtain the voiceprint feature vector of the reference audio signal. The determination unit is used to obtain a target voiceprint feature vector according to the first parameter, the second parameter, and the voiceprint feature vector of the reference audio signal. The obtaining method of the target voiceprint feature vector includes: when the first parameter indicates that the similarity of the user's requirement is higher than a preset similarity, and the second parameter characterizes that the naturalness of the user's requirement is lower than a preset naturalness, determine the voiceprint feature vector of the reference audio signal as the target voiceprint feature vector; when the first parameter indicates that the similarity of the user's requirement is lower than a preset similarity, and the second parameter characterizes that the naturalness of the user's requirement is higher than a preset naturalness, obtain the average voiceprint feature vectors of each of the M sample users in the training set used to train the voiceprint feature model according to the audio data of the M sample users, where M is an integer greater than 1, calculate the similarity between the voiceprint feature vector of the reference audio signal and the average voiceprint feature vectors of each of the M sample users, and determine the voiceprint feature vector of the target sample user as the target voiceprint feature vector, where the target sample user is the sample user with the highest similarity between the voiceprint feature vector and the voiceprint feature vector of the reference audio signal among the M sample users. The feature extraction unit is further used to extract features from the text data to obtain a text feature vector. The determination unit is further used to obtain the speech spectrum corresponding to the text data according to the text feature vector and the target voiceprint feature vector. The processing unit is used to process the speech spectrum by using a vocoder to obtain a target audio signal, and the text data represented by the target audio signal is the text data input by the user.
7. An electronic device, characterized in that, Including: A processor and a memory, the processor being connected to the memory, the memory being configured to store a computer program, and the processor being configured to execute the computer program stored in the memory so that the electronic device performs the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1-5.
Citation Information
Patent Citations
Audio generation method and device based on artificial intelligence, equipment and storage medium
CN113822017A