Timbre card generation method, sound cloning method and 3D digital human video generation method
By generating tone cards to store audio features and description information, the problem of time-consuming and labor-intensive audio recording for 3D digital humans is solved, and efficient driving of 3D digital humans is achieved, improving user experience.
Patent Information
- Application Number
- CN202510537531.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-09-09
AI Technical Summary
In the existing technology, users need to record audio every time they drive a 3D digital human. As a result, the recording quality is affected by the external environment and personal status, which is time-consuming and labor-intensive, and the audio quality cannot be guaranteed.
By extracting the sound features, spectral features and discrete speech coding sequences of the target audio, timbre cards are generated and stored in association with the timbre description information. Users can directly select the timbre card to drive the 3D digital human, avoiding repeated recording.
It saves users’ time and costs, reduces workload, improves recording quality, enhances user experience, and records audio in a better environment to meet the needs of driving 3D digital humans.
Smart Images

Figure CN120612957A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sound processing, and in particular to a method for generating a timbre card, a method for generating a sound cloning method, and a method for generating a 3D digital human video. Background Art
[0002] A 3D digital human is a virtual, digital humanoid created using digital technology that closely resembles a human. It possesses specific characteristics such as appearance, gender, and personality, and the ability to express itself through language, facial expressions, and body movements. 3D digital humans can be categorized as interactive or non-interactive. Interactive 3D digital humans have the ability to recognize their environment and interact with humans.
[0003] To generate 3D digital human videos, related technologies allow users to upload a piece of audio as driving data, driving the 3D digital human to make sounds and perform corresponding expressions and body movements, making the 3D digital human the user's avatar in the digital world.
[0004] However, every time a user needs to control a 3D digital human to express something on their behalf, they must first record the corresponding audio to drive the 3D digital human. Firstly, audio recording still requires the user to take the time to complete, and no one else can do it for them. Secondly, due to the influence of external objective conditions and the user's own condition, the quality of the audio recording cannot be guaranteed. If the recording effect is not good, it needs to be repeated until it meets the requirements, which is time-consuming and laborious. Summary of the Invention
[0005] The present invention provides a timbre card generation method, a sound cloning method and a 3D digital human video generation method, which are used to solve the defects existing in the related technologies.
[0006] The present invention provides a method for generating a tone card, comprising: Receive the user's target audio, extract the sound features and spectral features of the target audio, and perform speech decomposition and encoding on the target audio to obtain a speech discrete code sequence; The timbre description information of the target audio is received, and a timbre card corresponding to the target audio is generated based on the sound feature, the spectrum feature, the speech discrete code sequence, and the timbre description information.
[0007] The present invention also provides a sound cloning method, comprising: Receive a user's target text and a target timbre card selected by the user; the target timbre card is obtained based on the above-mentioned timbre card generation method; Determine the phoneme sequence of the target text, and generate synthetic audio corresponding to the target text based on the phoneme sequence and the sound features, spectral features, and speech discrete coding sequence corresponding to the target timbre card.
[0008] The present invention also provides a 3D digital human video generation method, comprising: Obtaining a synthesized audio corresponding to a target text and a phoneme sequence of the target text; the synthesized audio is obtained based on the above-mentioned sound cloning method; Based on the phoneme sequence, marking the synthesized audio with a timestamp to obtain marked audio; Based on the marked audio, a 3D digital human video is generated.
[0009] The present invention also provides a device for generating a tone card, comprising: A feature extraction and encoding module is used to receive the user's target audio, extract the sound features and spectral features of the target audio, and perform voice decomposition and encoding on the target audio to obtain a voice discrete code sequence; The card generation module is used to receive the timbre description information of the target audio and generate a timbre card corresponding to the target audio based on the sound characteristics, the spectrum characteristics, the speech discrete code sequence and the timbre description information.
[0010] The present invention also provides a sound cloning device, comprising: A receiving module, configured to receive a user's target text and a target timbre card selected by the user; the target timbre card is obtained based on the above-mentioned timbre card generation method; The audio generation module is used to determine the phoneme sequence of the target text and generate the synthesized audio corresponding to the target text based on the phoneme sequence and the sound features, spectral features, and speech discrete coding sequence corresponding to the target timbre card.
[0011] The present invention also provides a 3D digital human video generation system, comprising: A phoneme sequence acquisition module, configured to acquire a synthesized audio corresponding to a target text and a phoneme sequence of the target text; the synthesized audio is obtained based on the above-mentioned sound cloning method; an audio tagging module, configured to tag the synthesized audio with a timestamp based on the phoneme sequence to obtain a tagged audio; The video generation module is used to generate a 3D digital human video based on the marked audio.
[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for generating a timbre card, the method for generating a sound cloning method, or the method for generating a 3D digital human video as described above is implemented.
[0013] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for generating a timbre card, a sound cloning method, or a 3D digital human video generation method as described above is implemented.
[0014] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned methods for generating a timbre card, or a sound cloning method, or a 3D digital human video generating method.
[0015] The timbre card generation method, sound cloning method, and 3D digital human video generation method provided by the present invention obtain sound characteristics, spectral characteristics, and discrete speech code sequences from target audio provided by the user. Combined with the timbre description information provided by the user, a timbre card corresponding to the target audio is generated. This allows for the associated storage of the audio characteristics and timbre description information of the target audio, and allows for the identification and differentiation of the audio characteristics of different target audios using the timbre description information. Furthermore, when the user needs to drive the 3D digital human, they can simply select the desired timbre card to determine the audio characteristics and timbre description information of the corresponding target audio, eliminating the need to record audio every time they drive the human. This saves the user time and reduces their workload. Furthermore, the target audio can be recorded when the external environment and the user's state are less affected, ensuring that the target audio meets processing requirements, saving time spent driving the 3D digital human and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is one of the flow charts of the method for generating a timbre card provided by the present invention.
[0018] Figure 2 This is the second flow chart of the method for generating a tone card provided by the present invention.
[0019] Figure 3It is a schematic diagram of the tone card provided by the present invention.
[0020] Figure 4 This is one of the flow charts of the sound cloning method provided by the present invention.
[0021] Figure 5 This is the second flow chart of the sound cloning method provided by the present invention.
[0022] Figure 6 It is a flow chart of the 3D digital human video generation method provided by the present invention.
[0023] Figure 7 It is a structural schematic diagram of the timbre card generating device provided by the present invention.
[0024] Figure 8 It is a structural schematic diagram of the sound cloning device provided by the present invention.
[0025] Figure 9 It is a structural diagram of the 3D digital human video generation system provided by the present invention.
[0026] Figure 10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0027] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0028] 3D digital humans require at least a specific appearance as their external image, as well as expressive abilities such as language, facial expressions, and body movements. The closer these appearance and expressive abilities are to real humans, the better. Regarding language expression, 3D digital humans primarily communicate and interact with people through voice.
[0029] With the continuous development of 3D digital human technology, users hope that 3D digital humans can become their own avatars in the digital world, capable of completing tasks such as information sharing, business presentations, and product promotions on their behalf. To achieve these technical effects, the 3D digital humans must be highly similar to the users in terms of appearance and voice.
[0030] In existing solutions, users must record the corresponding audio every time they want to drive a 3D digital human. This not only requires the user to take time to complete, but also affects the external objective environment and the user's condition, making the audio recording quality uncertain. This can lead to poor recording results and the need for repeated recordings, which is time-consuming and labor-intensive. Therefore, an embodiment of the present invention provides a method for generating tone cards, allowing users to directly select the desired tone card when driving a 3D digital human, eliminating the need to record audio every time.
[0031] Figure 1 FIG. 1 is a flow chart of a method for generating a tone card according to an embodiment of the present invention, Figure 1 As shown, the method includes: S11, receiving the user's target audio, extracting the sound features and spectral features of the target audio, and performing speech decomposition and encoding on the target audio to obtain a speech discrete code sequence; S12, receiving timbre description information of the target audio, and generating a timbre card corresponding to the target audio based on the sound features, spectrum features, speech discrete code sequence and timbre description information.
[0032] Specifically, the timbre card generation method provided in the embodiments of the present invention is implemented by a timbre card generation device. This device can be deployed in either an electronic device or a sound cloning device to provide the generated timbre cards to the sound cloning device. The electronic device can be a smartphone, tablet, or computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, tablet, or other computer, and the specific limitations are not provided herein.
[0033] First, step S11 is executed to receive the user's target audio. This target audio is the user's recorded audio that serves as a reference for subsequent voice cloning (also known as the "speaker's voice"). This target audio can be recorded online or pre-recorded and stored using a recording device. When a voice card is needed, it is uploaded to the voice card generation device, and the target audio is not specifically limited here.
[0034] Users can read reference text or text of their own choice while recording target audio.
[0035] After that, the received target audio can also be tested to determine whether the target audio meets the processing requirements in terms of duration, format, sensitive information, etc. If it meets the processing requirements, subsequent operations on the target audio will continue. If it does not meet the processing requirements, the user can be prompted to re-record or upload the target audio.
[0036] After receiving the user's target audio or if the target audio meets the processing requirements, the acoustic and spectral features of the target audio can be extracted. These acoustic features (also known as "speaker voice features") are high-level semantic parameters that describe the perceptual attributes or physical properties of the sound and can include at least one of timbre and style features. If these features are present, they can be derived by fusing the timbre and style features.
[0037] Timbre features represent the unique texture of a sound. They are determined by the physical properties of the sound source (such as vocal tract shape and vocal cord vibration pattern). They are used to distinguish the sound properties of different sound sources and can include spectral patterns, spectral structure, formant positions, and dynamic changes. Timbre features are extracted using a trained timbre feature extraction model. The structure of this timbre feature extraction model can be selected as needed and is not specifically limited here.
[0038] Stylistic features represent the paralinguistic attributes or expressions of speech, and are related to semantics, emotion, or cultural context. Stylistic features can include at least one of speech rate, volume, intonation, and language. Speech rate can be determined by segmenting speech segments using voice activity detection (VAD) and then counting the number of syllables or words per unit time. Volume can be determined by calculating short-term energy (RMS) or perceived loudness (e.g., ITU-R BS.1770) and extracting statistics such as energy variance and peak interval. Intonation can be determined by extracting fundamental frequency (F0) traces, normalizing them, and calculating the pitch range. Language type can be identified using pre-trained models (e.g., Whisper, VoxLingua, etc.). Furthermore, stylistic features can be extracted using a trained style feature extraction model. The structure of this style feature extraction model can be selected as needed and is not specifically limited here.
[0039] Spectral features are mathematical representations of the target audio in the frequency domain. They are extracted using time-frequency analysis techniques (such as Fourier transforms) and are used to quantify the physical properties of the sound and characterize the time-frequency characteristics of the target audio. Spectral features can include Mel-spectrogram, linear spectrum, constant-Q transform (CQT) spectrum, bark scale spectrum, gammatone filterbank spectrum, power spectral density (PSD), and wavelet scalogram.
[0040] Mel spectrum can be obtained by preprocessing the target audio, performing short-time Fourier transform (STFT), Mel filtering, logarithmic transformation, normalization and other operations.
[0041] A linear spectrum can be obtained using a short-time Fourier transform. For example, the target audio can be segmented into short-time frames. A window function is then applied to each frame to reduce spectral leakage. Each frame is then subjected to a fast Fourier transform to obtain a complex spectrum. The amplitude spectrum of each frame is then calculated to obtain a power spectrum. Finally, the power spectra of each frame are arranged in chronological order to form a linear spectrum.
[0042] The constant Q transform spectrum can be generated based on a filter bank with a logarithmic frequency scale. For example, the center frequencies can be set at logarithmic intervals, and the bandwidth can be designed for each center frequency. The signal can then be passed through the filter bank, and the energy of each frequency band can be calculated. The energy of each frequency band can be aligned in time to form the CQT spectrum.
[0043] Bark spectrum can be generated based on Bark scale filter bank, Gammatone spectrum can be generated by Gammatone filter bank, power spectral density can be generated by periodogram method or autocorrelation method, and wavelet spectrum can be obtained based on time-frequency analysis of wavelet transform.
[0044] Afterwards, the target audio can be speech-decomposed and encoded to obtain a discrete speech code sequence. Specifically, the target audio can be input into the integrated speech tokenizer (SpeechTokenizer). The integrated speech tokenizer uses an encoder-decoder architecture and incorporates residual vector quantization (RVQ) technology to achieve hierarchical decomposition and efficient encoding of the target audio, thereby obtaining and outputting a discrete speech code sequence. This discrete speech code sequence is the result of the hierarchical decomposition and encoding of the target audio.
[0045] At this time, sound features, spectrum features and speech discrete code sequences can be obtained through the target audio, which can be collectively referred to as audio features of the target audio.
[0046] Then, step S12 is executed. The user can input the timbre description information of the target audio into the timbre card generation device. The timbre description information is user-defined information that describes the timbre of the target audio. For example, it may include the gender of the timbre, the timbre avatar, the timbre name, and the timbre style, which are not specifically limited here. The timbre style can be steady, sunny, intellectual, deep, etc.
[0047] After receiving the timbre description information of the target audio, the timbre card generation device can use the sound features, spectral features, discrete speech code sequence, and timbre description information to generate a timbre card corresponding to the target audio. For example, the timbre description information can be used as the display content of the timbre card, and the audio features of the target audio (i.e., sound features, spectral features, and discrete speech code sequence) can be associated with the timbre card and persistently stored. The timbre card can be used to identify and distinguish the audio features of different target audio. Furthermore, the target audio, along with the sound features, spectral features, and discrete speech code sequence, can be associated with the timbre card and persistently stored, so that the required information can be subsequently extracted directly from the target audio.
[0048] The timbre card generation method provided in an embodiment of the present invention obtains sound characteristics, spectral characteristics, and discrete speech code sequences from a user-provided target audio, and combines this with the user-provided timbre description information to generate a timbre card corresponding to the target audio. This method enables the associated storage of the target audio's audio characteristics and timbre description information, and allows the identification and differentiation of the audio characteristics of different target audios using the timbre description information. Furthermore, when the user needs to drive a 3D digital human, they can simply select the desired timbre card to determine the audio characteristics and timbre description information of the corresponding target audio, eliminating the need to record audio every time they drive the human. This saves the user time and reduces workload. Furthermore, the target audio can be recorded when the external environment and the user's state are less affected, ensuring that the target audio meets processing requirements, saving time spent driving the 3D digital human and improving the user experience.
[0049] Based on the above embodiment, extracting the sound features and spectrum features of the target audio previously includes: Receive the user's online recorded authorization audio for the authorization text; Performing voiceprint similarity detection on the target audio and the online recorded authorized audio to obtain a first detection result, and performing text content detection on the online recorded authorized audio to obtain a second detection result; Based on the first detection result and the second detection result, authorization authentication is performed on the target audio.
[0050] Specifically, considering that the audio features obtained through the user's target audio are the user's private information, the user's authorization is required before persistent storage and application. Therefore, after receiving the user's target audio in step S11 and before extracting the sound features and spectral features of the target audio, the target audio can also be authorized and authenticated.
[0051] During the authorization and authentication process, the timbre card generation device can display a voice authorization interface to the user, through which the user can record the authorized audio for online recording and confirm the authorization text. The authorization text can be a text-based authorization agreement. After viewing the authorization text, the user can read the authorization text aloud to obtain the authorized audio for online recording. The timbre card generation device can obtain the authorized audio for online recording via the audio acquisition module of the electronic device in which it is located.
[0052] Afterwards, the timbre card generation device can perform voiceprint similarity detection on the target audio and the online recorded authorized audio, that is, determine whether the target audio and the online recorded authorized audio are from the same speaker, and obtain a first detection result. For example, the target audio and the online recorded authorized audio can first be subjected to preprocessing operations such as noise reduction, frame segmentation and windowing. Then, based on traditional models such as i-vector or d-vector, or deep learning models such as ECAPA-TDNN, ResNet / Transformer, the target audio and the online recorded authorized audio are respectively voiceprint embedded, obtaining a first embedding vector of the target audio and a second embedding vector of the online recorded authorized audio. The similarity between the first embedding vector and the second embedding vector is then calculated. Finally, the similarity is compared with a first preset threshold. If the similarity is greater than or equal to the first preset threshold, the first detection result can be determined to be passed, that is, the target audio and the online recorded authorized audio are considered to belong to the same speaker. Otherwise, the first detection result is determined to be failed, that is, the target audio and the online recorded authorized audio are not from the same speaker. Here, the first preset threshold can be set as needed, for example, it can be set to any value greater than or equal to 70%, which is not specifically limited here.
[0053] The timbre card generation device also needs to perform text content detection on the online-recorded authorized audio, specifically determining whether the text content read by the user is in the authorized text, and obtaining a second detection result. For example, the online-recorded authorized audio can first be subjected to automatic speech recognition (ASR) to obtain the text content of the online-recorded authorized audio. The text content of the online-recorded authorized audio is then compared with the authorized text, and the similarity between the two is calculated. Finally, the similarity is compared with a second preset threshold. If the similarity is greater than or equal to the second preset threshold, the second detection result can be determined to be a pass, meaning that the text content read by the user is in the authorized text. Otherwise, the second detection result can be determined to be a fail, meaning that the text content read by the user is not in the authorized text. The second preset threshold can be set as needed, for example, to any value greater than or equal to 90%, and is not specifically limited here.
[0054] Finally, the target audio can be authorized and authenticated by combining the first and second detection results. For example, if both the first and second detection results pass, the target audio is deemed to have been authorized and authenticated; if at least one of the first and second detection results fails, the target audio is deemed to have failed.
[0055] Furthermore, if the first test result is a failure, a first reminder message is sent to the user, which is used to remind the user to confirm whether the online recorded authorized audio and the target audio are spoken by the same person. If the second test result is a failure, a second reminder message is sent to the user, which is used to remind the user to confirm whether the text content read aloud is the content in the authorized text. If both the first and second test results fail, the first and second reminder messages are sent to the user at the same time.
[0056] In the embodiment of the present invention, dual authorization of sound and text content can ensure the user's own operation, avoid the leakage of private information such as audio characteristics, and improve the security of user information.
[0057] Based on the above embodiment, the sound features and spectrum features of the target audio are extracted, which also includes: Detect the duration, signal-to-noise ratio, format, and sensitive information of the target audio; Determine whether the target audio is available based on duration information, signal-to-noise ratio information, format information, and sensitive information.
[0058] Specifically, the target audio used in the timbre card generation process needs to meet the processing requirements. To ensure the smooth extraction of audio features, the processing requirements may include the total duration requirement of the target audio, for example, the total duration does not exceed the first preset duration; since the target audio may contain blanks and excessive noise, which are invalid parts and cannot identify the content, the processing requirements may also include the effective duration of the target audio being greater than or equal to the second preset duration; since the target audio may contain noise, noise, reverberation, etc., it is necessary to determine the audio quality of the effective part by calculating the signal-to-noise ratio, so the processing requirements may also include the signal-to-noise ratio of the target audio being greater than or equal to the signal-to-noise ratio threshold. In addition, the processing requirements may also include audio format requirements and the absence of sensitive information. The audio format requirement may be that the audio format of the target audio needs to be a target format that the timbre card generation device can support and the sampling rate is greater than or equal to the sampling rate threshold.
[0059] Here, the first preset duration, the second preset duration, the signal-to-noise ratio threshold, and the sampling rate threshold can all be set as needed. For example, the first preset duration can be set to 2 minutes, the second preset duration can be set to 10 seconds, the signal-to-noise ratio threshold can be set to 5dB, and the sampling rate threshold can be set to 16kHz. Target formats include WAV, FLAC, MP3, AAC, Ogg, Opus, M4A, etc. Sensitive information can include information such as the voice of public figures such as politicians and celebrities.
[0060] Therefore, before extracting the sound and spectrum features of the target audio, the target audio can be preprocessed to detect the duration information, signal-to-noise ratio information, format information, and sensitive information of the target audio. The duration information can include information such as the total duration and effective duration of the target audio. The format information can include information such as the audio format and sampling rate.
[0061] Thereafter, the duration information, signal-to-noise ratio information, format information, and sensitive information can be used to determine whether the target audio is available. Specifically, the duration information, signal-to-noise ratio information, format information, and sensitive information are determined to meet processing requirements. If all of the information meet the processing requirements, the target audio is determined to be available, and subsequent operations can proceed. If one or more of the duration information, signal-to-noise ratio information, format information, and sensitive information do not meet the processing requirements, the target audio is determined to be unavailable, and a third reminder message is sent to the user, prompting the user to redefine the target audio.
[0062] In an embodiment of the present invention, by detecting the duration information, signal-to-noise ratio information, format information and sensitive information of the target audio to determine whether the target audio is usable, it is possible to avoid the inability to extract audio features and ensure the accuracy of the extracted audio features.
[0063] Based on the above embodiment, the target audio is local audio uploaded by the user or online audio recorded by the user.
[0064] Specifically, the target audio of the user received in step S11 can be either local audio uploaded by the user or online audio recorded by the user.
[0065] The timbre card generating device can be configured with a human-computer interaction interface, and the human-computer interaction interface can be configured with an online audio recording option and a local audio upload option, so that users can record audio online and upload local audio.
[0066] If the user selects the online audio recording option, the human-computer interaction interface jumps to the audio upload operation interface. In this operation interface, the user can upload the target audio by dragging the audio file or by selecting the audio file in the dialog box.
[0067] If the user chooses to upload local audio, the human-computer interaction interface will jump to the online audio recording operation interface, which can display the reference text for audio recording to the user. The user can directly read the reference text for audio recording to prevent the user from not knowing what to say when recording audio. If the user is not satisfied with the reference text, he or she can also choose to generate a new reference text. It should be noted that the role of the reference text is only to provide assistance for users to record audio, and it does not force users to read the reference text aloud. In other words, users can also read other selected texts at will, and there is no restriction here.
[0068] In the embodiment of the present invention, two methods for obtaining target audio can be provided, so that the user can choose an appropriate method to provide the target audio according to needs.
[0069] Based on the above embodiment, when the sound feature includes at least one of a timbre feature and a style feature, extracting the sound feature of the target audio includes: The target audio is input into the sound feature extraction model to obtain the timbre features output by the timbre feature extraction model and the style features output by the style feature extraction model in the sound feature extraction model.
[0070] Specifically, the extraction of timbre features and / or style features can be achieved through a sound feature extraction model. The sound feature extraction model may include a timbre feature extraction model and a style feature extraction model. When the target audio is input into the sound feature extraction model, the timbre features can be extracted by the timbre feature extraction model, and the style features can be extracted by the style feature extraction model.
[0071] Here, the timbre feature extraction model can be obtained by training an initial timbre feature extraction model using timbre training samples with speaker identity information labels. The initial timbre feature extraction model can be a CAM++ voiceprint model or other models.
[0072] Voice training samples can be audio from open-source datasets such as VoxCeleb1 / 2, CN-Celeb, and LibriSpeech, or from self-built databases. The audio in self-built databases can be collected using different recording devices (mobile phones, microphones) and environments (quiet or noisy). The audio can also include special timbres such as hoarseness, children's voices, and dialects to enhance generalization.
[0073] When training the initial timbre feature extraction model, the timbre training samples can be first divided into short speech segments to facilitate processing by the initial timbre feature extraction model.
[0074] Each short speech segment is then fed into an initial timbre feature extraction model, which is trained using the speaker identity labels. The model optimization objective during training is to minimize a loss function for speaker classification. This loss function can be an AAM-softmax loss or other loss function, which is not specifically defined here.
[0075] During the training process, the initial timbre feature extraction model learns the timbre features of each speaker and stores them in the embedding layer of the initial timbre feature extraction model.
[0076] Therefore, when applying the timbre feature extraction model to extract timbre features, the audio segments obtained by segmenting the target audio are input into the timbre feature extraction model, and the timbre features can be obtained from the embedding layer of the timbre feature extraction model.
[0077] The style feature extraction model can be obtained by training an initial style feature extraction model using style training samples with style labels. The initial style feature extraction model can be a Global Style Tokens (GST) model or other models.
[0078] Style training samples can be speech data with obvious style characteristics, and style labels can be speech speed level, volume range, etc.
[0079] When training the initial style feature extraction model, the style training samples can be first divided into short speech segments to facilitate processing by the initial style feature extraction model.
[0080] Then, each short speech segment is input into the initial style feature extraction model to obtain the predicted style features output by the initial style feature extraction model. The loss function is calculated using the style label and the predicted style features, and the initial style feature extraction model is iteratively trained using the loss function. When the loss function converges or the number of iterations reaches a preset number, the style feature extraction model is obtained.
[0081] Therefore, when applying the style feature extraction model to extract style features, each audio segment obtained by segmenting the target audio is input into the style feature extraction model to obtain the style features output by the style feature extraction model.
[0082] The sound feature extraction model further includes a fusion unit. After extracting the timbre feature and the style feature, the fusion unit can fuse the timbre feature and the style feature to obtain the sound feature. Here, the sound feature is in vector form.
[0083] In the embodiment of the present invention, the timbre feature extraction model and the style feature extraction model in the sound feature extraction model are used to extract timbre features and style features respectively, thereby improving the extraction efficiency and accuracy of the sound features of the target audio.
[0084] Figure 2 Schematic diagram of a complete flow chart of a method for generating a timbre card according to an embodiment of the present invention. The method includes: The user enters the human-computer interaction interface; The user selects the online audio recording option or the local audio upload option in the human-computer interaction interface; If the user selects the option to upload local audio, the user can upload the target audio by dragging the audio file or by selecting the audio file in the dialog box; If the user selects the online audio recording option, the user can first select a reference text and then record the target audio online by reading the reference script aloud; After receiving the target audio, the timbre card generating device can pre-process the target audio, that is, detect the duration information, signal-to-noise ratio information, format information and sensitive information of the target audio to determine whether the target audio is usable; When the target audio is available, audio features are extracted from the target audio to obtain timbre features, style features, speech discrete code sequence and spectrum features. The timbre features and style features constitute the sound features. The sound features, speech discrete code sequence and spectrum features are associated with the timbre card and stored, and the timbre description information of the target audio is used as the display content of the timbre card to obtain the final timbre card.
[0085] like Figure 3 As shown, users can use the timbre description information to distinguish different timbre cards, and then use the timbre cards to determine the associated sound characteristics, speech discrete code sequences, and spectral characteristics. The timbre description information of each timbre card may include language, timbre gender, application scenario, style characteristics (i.e., timbre style), timbre name, and timbre avatar. Timbre cards can be filtered by language, timbre gender, application scenario, and style characteristics. Figure 3 Only some voice cards with Chinese as the language and female as the voice gender are shown, with no restrictions on application scenarios or styles. For example, voice names may include Professional Female Voice, Female Explanator, Sharp Female Commentator, Spirited Female Voice, Enthusiastic Senior, Mature Sister, Girl Next Door, Sweet Girl, Pure and Sweet, Intellectual Host, Smart Commentator, and Gentle Lecturer. Personality traits can be displayed below each voice name, such as "Youthful Explanator" and "Professional and Capable" for a Professional Female Voice.
[0086] Based on the above embodiments, Figure 4As shown, an embodiment of the present invention further provides a sound cloning method, comprising: S21, receiving a user's target text and a target timbre card selected by the user; The target timbre card is obtained based on the timbre card generation method provided in the above embodiments; S22, determining the phoneme sequence of the target text, and generating synthetic audio corresponding to the target text based on the phoneme sequence and the sound features, spectral features, and speech discrete coding sequence corresponding to the target timbre card.
[0087] Specifically, the sound cloning method provided in the embodiments of the present invention is performed by a sound cloning device, which can be configured within an electronic device. The electronic device can be a smartphone, tablet, or computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, tablet, or the like, without specific limitations herein. The sound cloning device can be configured with a timbre card generation device to apply timbre cards generated and stored by the timbre card generation device. The timbre cards can also be directly stored within the sound cloning device, without specific limitations herein.
[0088] First, step S21 is executed to receive a user's target text and a target timbre card selected by the user. The target text can be the driving text required for voice cloning, containing the content required for voice cloning. In other words, the content of the target text is the content of the cloned synthetic voice reading. The target timbre card can be generated using the timbre card generation method provided in the above embodiments. Please refer to the above embodiments for details, and will not be repeated here.
[0089] The user can select a tone card as a target tone card in the tone cloning operation interface of the tone cloning device.
[0090] Then, step S22 is performed. Text-to-Speech (TTS) technology can be used to determine the phoneme sequence of the target text. This phoneme sequence is an ordered set of phonemes within the phonetic representation of each language unit in the target text. A language unit refers to the smallest unit of language representation in the target text, such as a single Chinese character or an English word. A phonetic representation refers to the pronunciation basis of a language unit, such as the pinyin of a Chinese character or the phonetic symbol of an English word. A phoneme is the smallest unit of speech defined by the natural properties of speech. For example, Mandarin Chinese has 32 phonemes and English has 48 phonemes. Each phonetic representation contains one or more phonemes, so the target text can correspond to a phoneme sequence consisting of multiple phonemes.
[0091] The synthesized audio corresponding to the target text can then be generated using this phoneme sequence and the acoustic, spectral, and discrete speech code sequences corresponding to the target timbre card. For example, attention or dynamic time warping (DTW) can be used to align the phoneme sequence with the temporal features of the target audio uploaded or recorded online when the user generated the timbre card. This temporal feature could be the duration of each phoneme in the target audio.
[0092] Then, the phoneme sequence, sound features, spectral features and speech discrete coding sequence are input into the speech synthesis model, which uses mechanisms such as cross-attention or conditional layer normalization to perform multimodal fusion on the phoneme sequence, sound features, spectral features and speech discrete coding sequence to obtain the synthesized speech output by the speech synthesis model.
[0093] The speech synthesis model used here can be an end-to-end model, such as VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech) and FastSpeech, or a two-stage model, such as Tacotron 2 + HiFi-GAN. The first stage generates a mel-spectrogram based on the input phoneme sequence, acoustic features, spectral features, and discrete speech coding sequence. The second stage uses a vocoder to convert the mel-spectrogram into a waveform, thus generating the synthesized speech. Alternatively, the speech synthesis model can be based on discrete coding, such as VQ-TTS. The discrete codes of the VQ-VAE are used as an intermediate representation and combined with the phoneme sequence to generate synthesized speech.
[0094] In the voice cloning method provided in the embodiments of the present invention, first, the target text of the user and the target voice color card selected by the user are received. Then, the phoneme sequence of the target text is determined, and the synthetic audio corresponding to the target text is generated by using the phoneme sequence, the voice features, spectral features, and voice discrete coding sequence corresponding to the target voice color card. This method can extract the audio features of the user from a target audio containing the user's voice for persistent storage and associate them with the voice color card. When the user needs to let the 3D digital human replace himself for language expression, there is no need for the user to record audio. Just upload the target text corresponding to the audio to be cloned and combine it with the pre-persistently stored audio features corresponding to the target voice color card selected by the user to generate the synthetic audio corresponding to the target text, realizing the cloning of the user's voice. This can avoid the user providing the drive by recording audio when cloning the voice, reduce the user's time cost, reduce the user's cumbersome operations, and improve the user experience.
[0095] Based on the above embodiments, generating the synthetic audio corresponding to the target text based on the phoneme sequence, the voice features, spectral features, and voice discrete coding sequence corresponding to the target voice color card includes: Perform sentence segmentation and regularization processing on the target text respectively to obtain multiple sentence texts; Input the multiple sentence texts, the phoneme sequence, and the voice features in the target voice color card into the voice cloning large model to obtain the target voice coding sequence output by the voice cloning large model; Input the target voice coding sequence, the voice features, spectral features, and voice discrete coding sequence in the target voice color card into the spectral reconstruction network to obtain the target spectrum corresponding to the target voice coding sequence output by the spectral reconstruction network; Generate the synthetic audio based on the target spectrum.
[0096] Specifically, in the process of generating the synthetic audio, in the embodiments of the present invention, the language of the content of the target text can be recognized first. For example, it can be Chinese, English, Chinese-English mixed, other languages, etc. Synchronously, the target text can also be segmented into sentences according to punctuation marks to obtain multiple sentences, and each sentence is subjected to regularization processing to convert the characters in each sentence that are not in the standard form of a certain language into the standard form of a certain language. For example, the Arabic numerals 123, etc. can be converted into the corresponding Chinese characters one, two, three. Furthermore, the sentence text corresponding to each sentence is obtained.
[0097] The speech synthesis model adopted above may include a sound cloning model, a spectrum reconstruction network, a synthesis module, a vocoder and other structures. After obtaining each sentence text, the sound features of each sentence text, the phoneme sequence and the target timbre card can be input into the sound cloning model to obtain the target speech coding sequence output by the sound cloning model. The sound cloning model can be a Qwen2.5-0.5B model, or obtained by training the Qwen2.5-0.5B model. For example, the sample text can be first subjected to sentence segmentation and regularization processing to obtain each sample text sentence, and the phoneme sequence of the sample text can be determined. The Qwen2.5-0.5B model can be trained through each sample text sentence, the phoneme sequence of the sample text, the sound features of the sample audio, and the target speech coding sequence of the sample audio.
[0098] The target speech code sequence output by the large voice cloning model, along with the acoustic features, spectral features, and discrete speech code sequence from the target timbre card, is then fed into a spectral reconstruction network. This network can be based on the Flow Matching module of CosyVoice2, and the spectral features and target spectrum are of the same type, either linear or mel-spectrographic.
[0099] Finally, the target spectrum can be used to generate synthesized audio. For example, if the target spectrum is a linear spectrum, and the phase of the linear spectrum is known, the synthesized audio can be directly obtained by performing an inverse short-time Fourier transform (ISTFT) or inverse fast Fourier transform (IFFT) through the synthesis module. If the phase of the linear spectrum is unknown, the synthesis module can iteratively optimize the phase using the Griffin-Lim algorithm, and then obtain the synthesized audio through an inverse short-time Fourier transform (ISTFT) or inverse fast Fourier transform (IFFT).
[0100] If the target spectrum is a Mel spectrum, a vocoder can be used to obtain synthesized audio.
[0101] In the embodiment of the present invention, the combination of the sound cloning model and the spectrum reconstruction network can simplify the complexity of generating synthetic audio and reduce the difficulty of generation.
[0102] Based on the above embodiment, determining the phoneme sequence of the target text includes: Perform language recognition on the target text to obtain the target language of the target text, and perform sentence segmentation and regularization on the target text to obtain multiple sentence texts; Based on the target language, prosody prediction is performed on multiple sentence texts to determine the pause duration between adjacent language units in the multiple sentence texts and the pause markers corresponding to the pause duration; A language unit sequence of the target text is determined based on the adjacent language units and pause marks, and a phoneme sequence is determined based on the language unit sequence and the phonetic representation of the adjacent language units in the target language.
[0103] Specifically, when determining the phoneme sequence of the target text, the target text can first be language identified to obtain the target language of the target text. The target language can include Chinese, English, Korean, Japanese, and a combination of at least two languages. Thereafter, the target text can be sentence segmented and regularized to obtain multiple sentence texts.
[0104] Using the target language, we perform prosody prediction on multiple sentences, determining the pause durations between adjacent language units and the corresponding pause markers. Prosody prediction can be achieved using a prosody prediction model trained on Roformer. By feeding multiple sentences into the prosody prediction model, we can determine the pause durations between adjacent language units and use different pause markers to identify different pause durations.
[0105] Furthermore, using adjacent language units and pause markers, a language unit sequence consisting of language units and pause markers can be obtained. Using the language unit sequence and the phonetic representations of each adjacent language unit in the target language, a phoneme sequence can be determined. For example, if the target language includes Chinese and English, all Chinese characters can be converted to pinyin, resulting in a phonetic representation sequence consisting of pinyin / English words. Combining the language unit sequence and the phonetic representation sequence yields a phoneme sequence that includes pause markers.
[0106] On this basis, if the target language is a language with polyphones, the polyphone prediction model generated after Roformer training can be used to predict the phonetic representations corresponding to the polyphones in the target text to ensure the pronunciation accuracy of the polyphones in the target text.
[0107] In the embodiment of the present invention, pause marks are introduced into the phoneme sequence through prosody prediction, so that the synthesized audio conforms to the pause habits of the target audio.
[0108] Figure 5 FIG. 1 is a complete flow chart of the sound cloning method provided in an embodiment of the present invention. Figure 5 As shown, the method includes: Obtain the target text, perform language recognition, sentence segmentation, and regularization on the target text in sequence to obtain the target language of the target text and multiple sentences of the target text; Predict polyphones for each sentence and determine the speech representation sequence based on the speech representation in the target language. Perform prosody prediction on each sentence text to obtain the pause duration between adjacent language units, and then determine the language unit sequence; The phoneme sequence can be obtained by combining the speech representation sequence and the language unit sequence; By using the sound features in the timbre cards obtained by the timbre card generation method provided in the above embodiments, combining multiple sentence texts and phoneme sequences, and applying the sound cloning model, the target speech coding sequence can be obtained; The target speech coding sequence is input into the spectrum reconstruction network together with the speech discrete coding sequence, spectrum features and sound features in the above-mentioned timbre card to obtain the target spectrum, and then the synthesized speech is obtained through the target spectrum.
[0109] like Figure 6 As shown, based on the above embodiment, an embodiment of the present invention further provides a method for generating a 3D digital human video, the method comprising: S31, obtaining a synthesized audio corresponding to the target text and a phoneme sequence of the target text; the synthesized audio is obtained based on the sound cloning method provided in the above embodiments; S32, based on the phoneme sequence, marking the synthesized audio with a timestamp to obtain a marked audio; S33, generating 3D digital human video based on the labeled audio.
[0110] Specifically, the 3D digital human video generation method provided in the embodiment of the present invention is executed by a 3D digital human video generation system, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not specifically limited here.
[0111] First, execute step S31 to obtain the synthesized audio corresponding to the target text and the phoneme sequence of the target text. The synthesized audio can be obtained by the sound cloning method provided in the above embodiments. The phoneme sequence of the target text can be determined by TTS technology or in combination with prosody prediction. Please refer to the above embodiments for details and will not be repeated here.
[0112] Then, step S32 is performed to timestamp the synthesized audio using the phoneme sequence to generate labeled audio. Using the phoneme sequence to timestamp the synthesized audio involves aligning the phoneme sequence with the timeline of the synthesized audio, annotating the start and end times for each phoneme in the phoneme sequence, and generating timestamped labeled audio. The labeled audio can be a TextGrid file or a JSON annotation file.
[0113] Timestamps can be implemented using Dynamic Time Warping (DTW) or Hidden Markov Model (HMM), or using a timestamp tagging model.
[0114] This timestamp labeling model can be built based on the native encoder-decoder of the Transformer network. When training the timestamp labeling model, the sample phoneme sequence corresponding to the sample text and the sample audio can be input into the initial timestamp labeling model, which then outputs the timestamp labeling result. The labeling result and the timestamp label carried by the sample audio are then used to calculate a loss function. Using this loss function, the timestamp labeling model is iteratively trained until the loss function converges or the number of iterations reaches a preset number, resulting in the trained timestamp labeling model.
[0115] Here, the encoder consists of N stacked encoder layers, each of which consists of a two-sublayer structure: the first sublayer is a multi-head self-attention sublayer, and the second sublayer is a feed-forward fully connected sublayer. Each sublayer is followed by a normalization layer and a residual connection. The decoder also consists of N stacked decoder layers, each of which consists of a three-sublayer structure: the first sublayer is a masked multi-head self-attention sublayer, the second sublayer is a multi-head attention sublayer (encoder-to-decoder), and the third sublayer is a feed-forward fully connected sublayer. Each sublayer is followed by a normalization layer and a residual connection.
[0116] Finally, step S33 is executed to generate a 3D digital human video using the tagged audio. For example, the tagged audio can be processed first, using Whisper to extract the text content and phoneme timestamps of the tagged audio. Rhubarb can then be used to generate a JSON file for the lip-sync animation. 3D modeling can then be performed, creating a character in Blender and rigging the facial skeleton. Animation can then be imported, converting the Rhubarb JSON file into Blender shape keyframes. Emotional actions can then be added, manually or through scripts to add expressions (e.g., smiles, frowns) based on the audio emotion tags. Finally, the character and animation are exported to FBX format and imported into Unreal Engine. Lighting and cameras are configured in Unreal Engine, and the scene is built and rendered. Sequencer is used to synchronize the audio and animation, and the 3D digital human video is rendered and output.
[0117] The 3D digital human video generation method provided in the embodiments of the present invention, by marking the timestamp of the synthesized audio obtained in the above embodiments, and then using the marked audio to generate the 3D digital human video, can make the process of generating the 3D digital human video only require the user to input the target text and select the target tone card, without the user having to perform additional operations, which can greatly simplify the user's operation steps, reduce the user's operation difficulty, and improve the user experience.
[0118] On the basis of the above embodiment, generating a 3D digital human video based on the marked audio includes: Based on the timbre description information in the target timbre card, determining an action style that matches the timbre description information; Generate 3D digital human videos based on action styles and labeled audio.
[0119] Specifically, during the process of generating a 3D digital human video, the timbre description information in the target timbre card can be used to determine a matching action style. This action style can be selected from an action style library based on the timbre description information. For example, based on the gender of the timbre in the timbre description information, an action style matching the gender of the timbre can be selected from the action style library. The matched action style and the tagged audio are then combined to generate a 3D digital human video. This ensures that the 3D digital human video better meets the user's expectations.
[0120] like Figure 7 As shown, based on the above embodiment, an embodiment of the present invention provides a timbre card generating device, comprising: The feature extraction and encoding module 61 is used to receive the user's target audio, extract the sound features and spectral features of the target audio, and perform voice decomposition and encoding on the target audio to obtain a voice discrete code sequence; The card generation module 62 is configured to receive the timbre description information of the target audio and generate a timbre card corresponding to the target audio based on the sound features, spectrum features, speech discrete code sequence and the timbre description information.
[0121] On the basis of the above embodiment, the timbre card generating device provided in the embodiment of the present invention further includes an authorization and authentication module for: Receive the user's online recorded authorization audio for the authorization text; Performing voiceprint similarity detection on the target audio and the online recorded authorized audio to obtain a first detection result, and performing text content detection on the online recorded authorized audio to obtain a second detection result; Based on the first detection result and the second detection result, authorization authentication is performed on the target audio.
[0122] On the basis of the above embodiment, the timbre card generating device provided in the embodiment of the present invention further includes a pre-processing module for: Detect the duration, signal-to-noise ratio, format, and sensitive information of the target audio; Determine whether the target audio is available based on duration information, signal-to-noise ratio information, format information, and sensitive information.
[0123] On the basis of the above-mentioned embodiment, in the apparatus for generating a timbre card provided in the embodiment of the present invention, the target audio is local audio uploaded by the user or online audio recorded by the user.
[0124] On the basis of the above embodiment, in the timbre card generating device provided in the embodiment of the present invention, the sound feature includes at least one of a timbre feature and a style feature.
[0125] Based on the above embodiments, the timbre card generation device and feature extraction and encoding module provided in the embodiments of the present invention are specifically used to: The target audio is input into the sound feature extraction model to obtain the timbre features output by the timbre feature extraction model and the style features output by the style feature extraction model in the sound feature extraction model, and the fusion unit in the sound feature extraction model fuses the timbre features and style features to obtain the sound features.
[0126] Specifically, the functions of each module in the timbre card generating device provided in the embodiment of the present invention correspond one-to-one to the operation procedures of each step in the above method embodiment, and the effects achieved are also consistent. Please refer to the above embodiment for details, which will not be described in detail in the embodiment of the present invention.
[0127] like Figure 8 As shown, based on the above embodiment, an embodiment of the present invention provides a sound cloning device, including: A receiving module 71 is configured to receive a user's target text and a target timbre card selected by the user; the target timbre card is obtained based on the timbre card generation method provided in the above embodiments; The audio generation module 72 is used to determine the phoneme sequence of the target text and generate the synthesized audio corresponding to the target text based on the phoneme sequence and the sound features, spectral features, and speech discrete code sequence corresponding to the target timbre card.
[0128] Based on the above embodiment, in the sound cloning device provided in the embodiment of the present invention, the audio generation module is specifically configured to: Perform sentence segmentation and regularization on the target text to obtain multiple sentence texts; Input multiple sentence texts, phoneme sequences, and sound features in target timbre cards into the sound cloning model to obtain the target speech coding sequence output by the sound cloning model; Input the target speech coding sequence and the sound features, spectrum features and speech discrete coding sequence in the target timbre card into the spectrum reconstruction network, and obtain the target spectrum corresponding to the target speech coding sequence output by the spectrum reconstruction network; Generates synthetic audio based on the target spectrum.
[0129] Based on the above embodiment, in the sound cloning device provided in the embodiment of the present invention, the audio generation module is specifically configured to: Perform language recognition on the target text to obtain the target language of the target text, and perform sentence segmentation and regularization on the target text to obtain multiple sentence texts; Based on the target language, prosody prediction is performed on multiple sentence texts to determine the pause duration between adjacent language units in the multiple sentence texts and the pause markers corresponding to the pause duration; A language unit sequence of the target text is determined based on the adjacent language units and pause marks, and a phoneme sequence is determined based on the language unit sequence and the phonetic representation of the adjacent language units in the target language.
[0130] Specifically, the functions of the modules in the sound cloning device provided in the embodiment of the present invention correspond one-to-one to the operation procedures of the steps in the above method embodiment, and the effects achieved are also consistent. Please refer to the above embodiment for details, which will not be described in detail in the embodiment of the present invention.
[0131] like Figure 9 As shown, based on the above embodiment, an embodiment of the present invention provides a 3D digital human video generation system, including: The phoneme sequence acquisition module 81 is used to obtain the synthesized audio corresponding to the target text and the phoneme sequence of the target text; the synthesized audio is obtained based on the sound cloning method provided in the above embodiments; An audio tagging module 82 is configured to tag the synthesized audio with a timestamp based on the phoneme sequence to obtain a tagged audio; The video generation module 83 is used to generate a 3D digital human video based on the marked audio.
[0132] On the basis of the above embodiments, in the 3D digital human video generation system provided in the embodiments of the present invention, the video generation module is specifically used for: Based on the timbre description information in the target timbre card, determining an action style that matches the timbre description information; Generate 3D digital human videos based on action styles and labeled audio.
[0133] Specifically, the functions of each module in the 3D digital human video generation system provided in the embodiment of the present invention correspond one-to-one to the operation process of each step in the above method embodiment, and the effects achieved are also consistent. Please refer to the above embodiment for details, and no further details will be given in the embodiment of the present invention.
[0134] Figure 10 An example of a physical structure diagram of an electronic device is shown below. Figure 10 As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the timbre card generation method provided in the above embodiments, the sound cloning method provided in the above embodiments, or the 3D digital human video generation method provided in the above embodiments.
[0135] Furthermore, the logic instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the relevant art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0136] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the tone card generation method provided in the above embodiments, or the sound cloning method provided in the above embodiments, or the 3D digital human video generation method provided in the above embodiments.
[0137] In yet another aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program is used to perform the timbre card generation method, the voice cloning method, or the 3D digital human video generation method provided in the aforementioned embodiments. The computer-readable storage medium may be either a non-transitory or transient computer-readable storage medium, and is not specifically limited herein.
[0138] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0139] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for generating a tone card, characterized in that: include: Receive the user's target audio, extract the sound features and spectral features of the target audio, and perform speech decomposition and encoding on the target audio to obtain a speech discrete code sequence; The timbre description information of the target audio is received, and a timbre card corresponding to the target audio is generated based on the sound feature, the spectrum feature, the speech discrete code sequence, and the timbre description information.
2. The method for generating a timbre card according to claim 1, wherein: The extracting of the sound features and spectral features of the target audio includes: Receiving an online recorded authorization audio of the user for the authorization text; Performing voiceprint similarity detection on the target audio and the online recorded authorized audio to obtain a first detection result, and performing text content detection on the online recorded authorized audio to obtain a second detection result; Based on the first detection result and the second detection result, authorization authentication is performed on the target audio.
3. The method for generating a timbre card according to claim 1, wherein: The extracting of the sound features and spectral features of the target audio further includes: Detecting duration information, signal-to-noise ratio information, format information, and sensitive information of the target audio; Whether the target audio is available is determined based on the duration information, the signal-to-noise ratio information, the format information, and the sensitive information.
4. The method for generating a tone card according to any one of claims 1 to 3, characterized in that: The target audio is local audio uploaded by the user or online audio recorded by the user.
5. The method for generating a tone card according to any one of claims 1 to 3, characterized in that: The sound feature includes at least one of a timbre feature and a style feature.
6. The method for generating a timbre card according to claim 5, wherein: The extracting the sound features of the target audio includes: The target audio is input into a sound feature extraction model to obtain the timbre features output by the timbre feature extraction model in the sound feature extraction model and the style features output by the style feature extraction model, and the fusion unit in the sound feature extraction model fuses the timbre features and the style features to obtain the sound features.
7. A sound cloning method, characterized in that: include: receiving a user's target text and a target timbre card selected by the user; The target timbre card is obtained based on the timbre card generation method according to any one of claims 1 to 6; Determine the phoneme sequence of the target text, and generate synthetic audio corresponding to the target text based on the phoneme sequence and the sound features, spectral features, and speech discrete coding sequence corresponding to the target timbre card.
8. The sound cloning method according to claim 7, characterized in that: The step of generating a synthesized audio corresponding to the target text based on the phoneme sequence and the sound features, spectral features, and speech discrete code sequence corresponding to the target timbre card includes: Performing sentence segmentation and regularization processing on the target text to obtain multiple sentence texts; Inputting the plurality of sentence texts, the phoneme sequence, and the sound features in the target timbre card into a sound cloning model to obtain a target speech coding sequence output by the sound cloning model; Inputting the target speech coding sequence and the sound features, spectrum features and speech discrete coding sequence in the target timbre card into a spectrum reconstruction network, and obtaining a target spectrum corresponding to the target speech coding sequence output by the spectrum reconstruction network; The synthesized audio is generated based on the target frequency spectrum.
9. The sound cloning method according to claim 7, characterized in that: Determining the phoneme sequence of the target text includes: Performing language recognition on the target text to obtain the target language of the target text, and performing sentence segmentation and regularization processing on the target text to obtain multiple sentence texts; Based on the target language, prosody prediction is performed on the plurality of sentence texts to determine pause durations between adjacent language units in the plurality of sentence texts and pause markers corresponding to the pause durations; A language unit sequence of the target text is determined based on the adjacent language units and the pause marks, and the phoneme sequence is determined based on the language unit sequence and the phonetic representations of the adjacent language units in the target language.
10. A 3D digital human video generation method, characterized in that: include: Obtaining a synthesized audio corresponding to a target text and a phoneme sequence of the target text; the synthesized audio is obtained based on the sound cloning method according to any one of claims 7 to 9; Based on the phoneme sequence, marking the synthesized audio with a timestamp to obtain marked audio; Based on the marked audio, a 3D digital human video is generated.
11. The 3D digital human video generation method according to claim 10, characterized in that: Generating a 3D digital human video based on the marked audio includes: Based on the timbre description information in the target timbre card, determining an action style that matches the timbre description information; The 3D digital human video is generated based on the action style and the marked audio.
12. A device for generating a tone card, characterized in that: include: A feature extraction and encoding module is used to receive the user's target audio, extract the sound features and spectral features of the target audio, and perform voice decomposition and encoding on the target audio to obtain a voice discrete code sequence; The card generation module is used to receive the timbre description information of the target audio and generate a timbre card corresponding to the target audio based on the sound characteristics, the spectrum characteristics, the speech discrete code sequence and the timbre description information.
13. A sound cloning device, characterized in that: include: A receiving module, configured to receive a user's target text and a target timbre card selected by the user; The target timbre card is obtained based on the timbre card generation method according to any one of claims 1 to 6; The audio generation module is used to determine the phoneme sequence of the target text and generate the synthesized audio corresponding to the target text based on the phoneme sequence and the sound features, spectral features, and speech discrete coding sequence corresponding to the target timbre card.
14. A 3D digital human video generation system, characterized in that: include: a phoneme sequence acquisition module, configured to acquire a synthesized audio corresponding to a target text and a phoneme sequence of the target text; the synthesized audio is obtained based on the sound cloning method according to any one of claims 7 to 9; an audio tagging module, configured to tag the synthesized audio with a timestamp based on the phoneme sequence to obtain a tagged audio; The video generation module is used to generate a 3D digital human video based on the marked audio.
15. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for generating a timbre card according to any one of claims 1 to 6, the method for generating a sound cloning according to any one of claims 7 to 9, or the method for generating a 3D digital human video according to any one of claims 10 to 11 is implemented.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a timbre card according to any one of claims 1 to 6, the method for generating a sound cloning according to any one of claims 7 to 9, or the method for generating a 3D digital human video according to any one of claims 10 to 11 is implemented.
17. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a timbre card according to any one of claims 1 to 6, the method for generating a sound cloning according to any one of claims 7 to 9, or the method for generating a 3D digital human video according to any one of claims 10 to 11 is implemented.