Method and device for driving virtual digital image to sing
By obtaining virtual digital images, melody and text data, obtaining rhythm and tone data, correcting the songs, determining the lip coefficient sequence, and driving virtual digital image singing, it solves the problem that virtual digital people cannot perform accurately and naturally, and expands its application scenarios.
Patent Information
- Application Number
- CN202210637106.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-07
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-06-07
AI Technical Summary
The existing technology cannot model song melody and lyric text, resulting in virtual digital people being unable to accurately and naturally lip-driven, limiting their application scenarios.
By obtaining virtual digital images, target melody and text data, obtaining rhythm data and tone data, correcting the initial song based on these data, determining the target lip coefficient sequence, and driving the virtual digital images to sing.
The melody and lyrics of the song are modeled, the target songs with a specific rhythm are generated, and the accurate and natural singing of virtual digital images is realized, which expands its use scenarios.
Smart Images

Figure CN115083371B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as virtual digital images and intelligent media, and specifically to a method and device for driving a virtual digital image to sing. Background Art
[0002] Virtual digital images, such as virtual humans, have a wide range of industrial applications. The most common areas of application include virtual anchors, virtual customer service representatives, virtual assistants, virtual teachers, virtual idols, and other interactive games and entertainment. For example, in the related art of virtual humans, single-tone 3D facial lip movement driving methods can only drive machine sounds and real-person audio. It is impossible to model song melodies and lyrics to generate machine sounds with specific rhythms to accurately drive virtual digital human figures. This lack of capability has led to significant limitations in the actual application scenarios of virtual humans. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device and storage medium for driving a virtual digital image to sing.
[0004] According to one aspect of the present disclosure, a method for driving a virtual digital image to sing is provided, comprising: acquiring a virtual digital image, a target melody, and text data; acquiring rhythm data of the target melody, and processing the text data based on the rhythm data to acquire an initial song; acquiring pitch data of the target melody and frequency data of the target melody, and modifying the initial song based on the pitch data and the frequency data to acquire a target song; determining a target lip-shape coefficient sequence corresponding to the virtual digital image based on the text data, and driving the virtual digital image to sing the target song based on the target lip-shape coefficient sequence.
[0005] The method of driving a virtual digital image to sing provided by the present invention realizes modeling of song melody and lyrics text to generate a target song with a specific rhythm, and thereby drives the virtual digital image to accurately and naturally follow the lip movements, thereby realizing the virtual digital image singing and increasing the usage scenarios of the virtual digital image.
[0006] According to another aspect of the present disclosure, a device for driving a virtual digital image to sing is provided, including: an acquisition module for acquiring a virtual digital image, a target melody and text data; a processing module for acquiring rhythm data of the target melody, and processing the text data based on the rhythm data to obtain an initial song; a correction module for acquiring pitch data and frequency data of the target melody, and correcting the initial song based on the pitch data and the frequency data to obtain a target song; a driving module for determining a target lip-shape coefficient sequence corresponding to the virtual digital image based on the text data, and driving the virtual digital image to sing the target song based on the target lip-shape coefficient sequence.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to execute the above-mentioned method of driving a virtual digital image to sing.
[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned method of driving a virtual digital image to sing.
[0009] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the above-mentioned method of driving a virtual digital image to sing when executed by a processor.
[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0012] Figure 1 This is an exemplary implementation of a method for driving a virtual digital image to sing according to an exemplary embodiment of the present disclosure.
[0013] Figure 2 3D face schematic diagrams corresponding to different blenshape coefficients according to an exemplary embodiment of the present disclosure.
[0014] Figure 3 FIG. 1 is a schematic diagram of key points on the face of a target object according to an exemplary embodiment of the present disclosure.
[0015] Figure 4FIG. 3 is a schematic diagram of a 3D face corresponding to some blenshape coefficients that are not related to lip shape changes according to an exemplary embodiment of the present disclosure.
[0016] Figure 5 FIG. 2 is a schematic diagram of a process of determining an initial song according to an exemplary embodiment of the present disclosure.
[0017] Figure 6 FIG. 4 is a schematic diagram of rhythmically marking a target melody according to an exemplary embodiment of the present disclosure.
[0018] Figure 7 It is a schematic diagram of stretching or compressing the pronunciation duration of each target entity word in audio data according to an exemplary embodiment of the present disclosure.
[0019] Figure 8 FIG. 4 is a schematic diagram of modifying an initial song based on pitch data and frequency data according to an exemplary embodiment of the present disclosure.
[0020] Figure 9 FIG. 4 is a schematic diagram of a process for determining a target lip coefficient sequence according to an exemplary embodiment of the present disclosure.
[0021] Figure 10 Schematic diagram of the opening and closing amplitude according to an exemplary embodiment of the present disclosure.
[0022] Figure 11 3 is a schematic diagram of driving a virtual digital image to play a target song based on a target lip-sync coefficient sequence according to an exemplary embodiment of the present disclosure.
[0023] Figure 12 It is an overall flow chart of a method for driving a virtual digital image to sing according to an exemplary embodiment of the present disclosure.
[0024] Figure 13 4 is a schematic diagram of a device for driving a virtual digital image to sing according to an exemplary embodiment of the present disclosure.
[0025] Figure 14 is a schematic diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0027] Artificial Intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This discipline encompasses both hardware and software technologies. AI hardware technologies generally include computer vision, speech recognition, natural language processing, as well as deep learning / learning, big data processing, and knowledge graphs.
[0028] Intelligent media is an online social information dissemination system that integrates artificial intelligence and human intelligence. It is the sum of information clients and servers that perceive users and provide them with an enhanced experience. The core of intelligent media is to intelligently deliver products to users in real time based on their needs. The goal is to better serve users and thus inject strong competitiveness into their development.
[0029] Speech recognition technology, also known as automatic speech recognition (ASR), aims to convert the lexical content of human speech into computer-readable input, such as keystrokes, binary codes, or character sequences. This differs from speaker identification and speaker verification, which attempt to identify or verify the speaker of the speech rather than the lexical content.
[0030] Figure 1 This is an exemplary embodiment of a method for driving a virtual digital image to sing shown in the present disclosure, such as Figure 1 As shown, the method for driving a virtual digital image to sing comprises the following steps:
[0031] S101, obtaining a virtual digital image, a target melody and text data.
[0032] A virtual digital avatar is a hybrid of human-like characteristics, such as appearance, performance, and interaction, that exists outside the physical world. It is created and used using a variety of techniques, including motion capture, computer graphics, graphics rendering, deep learning, and speech synthesis. It can also be referred to as an avatar, virtual human, or digital human. Representative applications include virtual assistants, virtual customer service representatives, and virtual idols / anchors.
[0033] In the present disclosure, the acquired virtual digital image may be a virtual digital image selected from a plurality of virtual digital images to be selected, or may be a virtual digital image corresponding to the user who uses the virtual digital image.
[0034] Alternatively, if the user using the virtual digital avatar does not already have a corresponding virtual digital avatar, a video of the user can be recorded, and based on the recorded video, a 3D facial parameterized model can be used to reconstruct the user's face, thereby obtaining a corresponding three-dimensional facial model of the user. The 3D facial parameterized model is a vertex-based additive model learned from a large amount of facial data. It consists of a certain number of vertices and corresponding triangular facets, and includes a blendshape coefficient with different facial expressions. By weighting the different blendshape coefficients, the facial model can be driven to produce various facial expressions. Figure 2 This is a schematic diagram of a 3D face corresponding to different blenshape coefficients, such as Figure 2 As shown in the figure, the three 3D faces on the left, middle and right correspond to different blenshape coefficients, and different blenshape coefficients correspond to different 3D face shapes.
[0035] The key points of the target object’s face are detected by the facial key point model to obtain the two-dimensional facial key point data of the target object. Each key point detected by the facial key point model is accompanied by the confidence data of the key point. Figure 3 As shown, Figure 3 It is a schematic diagram of the facial key points of the target object detected by the facial key point model.
[0036] In order to effectively eliminate the position constraints of erroneous points on the three-dimensional face model and improve the robustness and stability of the virtual digital image fitting results, the three-dimensional face model and two-dimensional face key point data obtained above are fused. That is, the error between the 2D projection of the fitted three-dimensional face model and the detected two-dimensional face key point data is calculated, and the blenshape coefficients of the 3D face model are gradually generated to obtain the virtual digital image corresponding to the target object.
[0037] Preferably, since the movement of eyebrows and eyes is unrelated to the motion posture and lip shape changes, the 3D facial parametric model used in the present disclosure constrains the blenshape coefficients of eyebrows and eyes to zero, and removes the motion posture of the face to accurately drive the lip shape changes of the virtual digital image. Figure 4 This is a schematic diagram of a 3D face corresponding to some blenshape coefficients that are not related to the change of mouth shape, such as Figure 4 As shown in the figure, the mouth shape of the three 3D faces on the left, center, and right remains unchanged; only the eye state changes. In practical applications, the virtual digital image's mouth shape can be accurately driven, while the changes in the virtual digital image's eyebrows and eyes can be driven in a fixed or random manner, for example, controlling the virtual digital image to blink every 5 seconds.
[0038] Obtain target melody and text data. The target melody refers to the melody used to drive the virtual digital avatar to sing, and the text data refers to the lyrics used for singing. The text data can be built-in Chinese lyrics or foreign lyrics, or randomly specified text data.
[0039] S102: Acquire rhythm data of the target melody, and process the text data based on the rhythm data to obtain an initial song.
[0040] The target melody is rhythmically tapped to obtain the target melody's rhythm data. The rhythm data includes the rhythm point position and rhythm duration. The rhythm point position refers to the position of each rhythm point on the target melody, and the rhythm duration refers to the duration of the target melody between two adjacent rhythm points.
[0041] After obtaining the rhythm point position, the text data is converted into audio data using the Text To Speech (TTS) model. Since the pronunciation duration of each target entity word may not be exactly the same as the rhythm duration after the text data is converted into audio data, in order to achieve the accurate rhythm of the target song, the pronunciation duration of each target entity word in the audio data can be stretched or compressed to match the rhythm duration, and the audio matching the rhythm duration is obtained as the initial audio.
[0042] S103 , acquiring pitch data and frequency data of a target melody, and modifying the initial song based on the pitch data and the frequency data to acquire the target song.
[0043] The pitch data and frequency data of a given target melody in different rhythm point ranges are extracted, and the initial song is modified based on the pitch data and frequency data to obtain the target song.
[0044] For example, the voice style conversion model can effectively extract the pitch data and frequency data of the target melody in different rhythm point ranges, and the pitch data and frequency data in the target melody can be converted to the original song, thereby generating a machine sound with specific pitch data and frequency data as the target song.
[0045] S104: determining a target lip-sync coefficient sequence corresponding to the virtual digital image based on the text data, and driving the virtual digital image to sing the target song based on the target lip-sync coefficient sequence.
[0046] The coefficient corresponding to the lip shape posture of the virtual digital image is called the lip shape coefficient. According to the target entity words in the text data obtained above, the initial lip shape coefficient corresponding to each animation frame of the virtual digital image corresponding to each target entity word can be determined.
[0047] Optionally, after converting the text data into machine audio data, a preset mapping relationship between candidate speakers, candidate entity words, and candidate lip shape coefficients can be queried to obtain the initial lip shape coefficients corresponding to each animation frame of the virtual digital avatar corresponding to each target entity word. Each candidate speaker corresponds to a machine voice. For example, the candidate speakers may include a little boy, a little girl, an older man, etc.
[0048] Optionally, after obtaining the initial lip shape coefficient corresponding to each frame of animation, in order to improve the richness of the lip shape of the virtual digital image in the animation frame, the initial lip shape coefficient corresponding to each frame of animation can be optimized to obtain the target lip shape coefficient corresponding to each frame of animation after optimization, and all target lip shape coefficients are arranged based on the order of the target entity words in the text data to obtain a target lip shape coefficient sequence.
[0049] Based on the target lip shape coefficient sequence obtained above, combined with the selected virtual digital image, a virtual digital image animation frame corresponding to each target lip shape coefficient in the target lip shape coefficient sequence is generated, and all animation frames are spliced according to the order of the target entity words in the text data and played in the splicing order to generate a virtual digital image animation.
[0050] While playing all the animation frames in the splicing order, the target song obtained above is played synchronously, so that each target entity word currently playing in the target song corresponds one-to-one to the target lip coefficient of the current animation frame in the virtual digital image animation, that is, the virtual digital image is shown to be singing the target song.
[0051] The method for driving a virtual digital image to sing proposed in an embodiment of the present disclosure comprises the following steps: obtaining a virtual digital image, a target melody, and text data; obtaining rhythm data of the target melody, and processing the text data based on the rhythm data to obtain an initial song; obtaining pitch data and frequency data of the target melody, and modifying the initial song based on the pitch data and frequency data to obtain a target song; determining a target lip-sync coefficient sequence corresponding to the virtual digital image based on the text data, and driving the virtual digital image to sing the target song based on the target lip-sync coefficient sequence. The embodiment of the present disclosure realizes modeling of song melody and lyrics text to generate a target song with a specific rhythm, and thereby accurately and naturally drives the virtual digital image to sing with its lip shape, thereby realizing the virtual digital image singing and increasing the use scenarios of the virtual digital image.
[0052] Figure 5 This is an exemplary embodiment of a method for driving a virtual digital image to sing shown in the present disclosure, such as Figure 5 As shown, based on the above embodiment, the text data is processed based on the rhythm data to obtain the initial song, including the following steps:
[0053] S501: Generate initial audio based on text data and rhythm data.
[0054] Find all the rhythm points in the target melody and tap them. Figure 6 This is a schematic diagram of dotting the target melody with rhythmic points, such as Figure 6 As shown, all the rhythm point positions are found from the target melody and marked. After the rhythm point positions are given, the text data is converted into audio data using the Text To Speech (TTS) model. Since the pronunciation duration of each target entity word is not necessarily exactly the same as the rhythm duration after the text data is converted into audio data, in order to achieve the accurate rhythm of the target song, the pronunciation duration of each target entity word in the audio data can be stretched or compressed to match the rhythm duration, and the audio that matches the rhythm duration is obtained as the initial audio. Among them, the Chinese lyrics in the built-in text data need to be segmented in advance. When the built-in text data is in English, the space between two words can be used as the separation mark of the entity words.
[0055] Figure 7 It is a schematic diagram of stretching or compressing the pronunciation duration of each target entity word in the audio data, such as Figure 7 As shown, the audio data converted from the text data is not exactly the same as the rhythm point duration. In order to match the rhythm point duration, the pronunciation duration of each target entity word in the audio data is stretched or compressed to match the rhythm point duration.
[0056] S502: Determine a target speaker and pronunciation feature information of the target speaker.
[0057] After obtaining the initial audio, it is necessary to determine the target speaker and the pronunciation feature information of the target speaker.
[0058] As an achievable approach, when determining a target speaker, a selection instruction may be obtained, the selection instruction being used to instruct selection of a target speaker from a plurality of candidate speakers, the target speaker being determined according to the selection instruction, and pronunciation feature information of the target speaker being obtained. For example, if there are five candidate speakers, namely, candidate speaker 1, candidate speaker 2, candidate speaker 3, candidate speaker 4, and candidate speaker 5, candidate speaker 4 may be selected as the target speaker.
[0059] As another possible implementation, when determining the target speaker, text feature information of the text data is determined. Based on the scene or content corresponding to the text feature information, a suitable target speaker is determined from multiple candidate speakers, and the pronunciation feature information of the target speaker is obtained. For example, if the text data is lyrics of a children's song, a child can be selected as the target speaker.
[0060] S503: Adjust the initial audio according to the pronunciation feature information to generate an initial song that matches the pronunciation feature of the target speaker.
[0061] The target speaker and their pronunciation characteristics are obtained, and the initial audio obtained above is adjusted according to the pronunciation characteristics to generate an initial song that matches the target speaker's pronunciation characteristics. It should be noted that since the pitch and frequency data of the target melody have not yet been introduced into the initial song, the initial song obtained at this point can be understood as a machine sound corresponding to the target speaker that matches the rhythm points.
[0062] The embodiment of the present application determines the target speaker and the pronunciation feature information of the target speaker, adjusts the initial audio according to the pronunciation feature information, and generates an initial song that matches the pronunciation features of the target speaker, so that the initial song can be of different styles and more generalizable, increasing the applicable scenarios of virtual digital images singing.
[0063] Furthermore, after adjusting the initial audio according to the pronunciation feature information and generating an initial song that matches the pronunciation features of the target speaker, the pitch data and frequency data of the given target melody can be extracted, and the initial song can be corrected based on the pitch data and frequency data to obtain the target song. Figure 8 It is a schematic diagram of modifying the initial song based on pitch data and frequency data, such as Figure 8 As shown, the initial song is modified based on the pitch data and the frequency data to obtain the target song.
[0064] Figure 9 This is an exemplary embodiment of a method for driving a virtual digital image to sing shown in the present disclosure, such as Figure 9 As shown, based on the above embodiment, the process of determining the target lip coefficient sequence includes the following steps:
[0065] S901, obtaining multiple target entity words and the pronunciation order of the target entity words in text data.
[0066] The word segmentation model is used to perform a word segmentation operation on the text data to effectively segment the entity word range of the stylized machine sound, thereby obtaining multiple target entity words after the word segmentation operation, and obtaining the pronunciation order of the target entity words based on the arrangement order of the multiple target entity words in the text data. For example, if the text data is "Welcome to visit my hometown", then after performing the word segmentation operation on the text data using the word segmentation model, the multiple target entity words that can be obtained are "welcome", "you", "visit", "me", "of", and "hometown".
[0067] S902: Obtain the target mouth shape coefficient corresponding to each target entity word.
[0068] Each selectable speaker is used as a candidate speaker. For example, the candidate speakers may include a little boy, a little girl, an old man, etc. For example, each entity word that appears in the historical audio and video data of the target object or in the sampled audio and video data may be used as a candidate entity word, and the lip shape coefficient of each frame of the virtual digital image animation corresponding to each candidate entity word in the historical audio and video data of the target object or in the sampled audio and video data may be used as the candidate lip shape coefficient.
[0069] In order to obtain the target lip shape coefficient based on the target entity word and target speaker corresponding to the target object to drive the movement of the virtual digital image, a mapping relationship between each candidate speaker, candidate entity word and candidate lip shape coefficient is pre-established in the embodiment of the present disclosure.
[0070] For any candidate speaker, the process of determining the mapping relationship includes:
[0071] Obtain candidate video data for the candidate speaker. Optionally, the candidate video data can be sampled video data recorded specifically for the candidate speaker for establishing a mapping relationship, or can be historical video data for the candidate speaker. During the process of obtaining the candidate video data, the candidate speaker's voice information is synchronously recorded and used as the candidate audio data for the candidate speaker.
[0072] After obtaining candidate video data for the candidate speaker, the candidate speaker's facial key points and the candidate video frame sequence are obtained in each frame of the candidate video data. Based on the candidate speaker's facial key points in each frame of the candidate video data, the candidate speaker's face is reconstructed, and candidate animation frames corresponding to each candidate video frame of the candidate speaker are generated. The lip shape coefficients corresponding to the candidate virtual digital image in each candidate animation frame are used as candidate lip shape coefficients. Optionally, a 3D parametric facial model may be used when reconstructing the candidate speaker's face.
[0073] After obtaining the candidate animation frames corresponding to each candidate video frame of the candidate speaker, all candidate animation frames are spliced in the order of the candidate video frames, and the obtained animation sequence of the candidate speaker is used as the candidate animation sequence corresponding to the candidate speaker.
[0074] For the candidate audio data recorded synchronously during the process of obtaining the candidate video data of the candidate speaker, text recognition is performed on the candidate audio data to obtain candidate text data corresponding to the candidate audio data. Preferably, the candidate audio data may include multiple pronunciation syllables. Optionally, when performing text recognition on the candidate audio data, an automatic speech recognition (ASR) model may be used to perform text recognition on the candidate audio data.
[0075] After obtaining candidate text data corresponding to the candidate audio data, a word segmentation operation is performed on the candidate text data using a word segmentation model to effectively segment the entity word range of the candidate audio data, thereby obtaining multiple candidate entity words obtained after the word segmentation operation. Based on the candidate speaker, the candidate entity words corresponding to the candidate speaker, and the candidate lip shape coefficients corresponding to the candidate animation frame corresponding to the candidate speaker, a mapping relationship between the candidate speaker, the candidate entity words, and the candidate lip shape coefficients is established for subsequent use.
[0076] According to the target entity word in the text data corresponding to the target object determined above and the target speaker corresponding to the target object, the mapping relationship is queried, and the candidate lip shape coefficients corresponding to the target speaker and the target entity word obtained from the query are used as the initial lip shape coefficients corresponding to the target entity word.
[0077] As a feasible approach, the initial lip shape coefficients may be directly used as target lip shape coefficients corresponding to the virtual digital image.
[0078] As another feasible optimization method, in order to make the lip shape changes of the generated virtual digital image richer and fuller, the initial lip shape coefficient can be optimized to obtain the optimized target lip shape coefficient.
[0079] Optionally, when optimizing the initial mouth shape coefficients, vector information corresponding to the pronunciation mouth shape of each pronunciation unit in the candidate audio data corresponding to the target speaker can be obtained based on the candidate audio data of the target speaker. Based on the vector information corresponding to the pronunciation mouth shapes of all pronunciation units corresponding to the target speaker, the pronunciation mouth shape with the largest opening and closing amplitude is determined as the target pronunciation mouth shape, and the vector corresponding to the target pronunciation mouth shape is used as the target vector corresponding to the target speaker. Based on the target vector, the initial mouth shape coefficients are optimized to obtain the optimized target mouth shape coefficients.
[0080] For example, if the candidate audio data of the target speaker includes 1000 pronunciation units, the vector information corresponding to the pronunciation mouth shape of the target speaker when expressing these 1000 pronunciation units is obtained, each pronunciation unit corresponds to a vector, and from the 1000 vectors corresponding to these 1000 pronunciation units, a vector with the largest opening and closing amplitude is selected as the target vector corresponding to the target speaker, and based on the target vector, the initial mouth shape coefficient is optimized to obtain the optimized target mouth shape coefficient. Figure 10 is a schematic diagram of the opening and closing amplitude, such as Figure 10 As shown, the line connecting the two points of the upper and lower lips represents the opening and closing amplitude of the mouth. The opening amplitude should be as large as possible to increase the richness of the mouth shape changes of the virtual digital image, making the mouth shape changes of the generated virtual digital image richer and fuller.
[0081] Exemplarily, the initial lip shape coefficient is optimized based on the target vector. When obtaining the target lip shape coefficient, the first weight corresponding to the target vector and the second weight corresponding to the initial lip shape coefficient can be obtained, and based on the first weight and the second weight, the target vector and the initial lip shape coefficient are weighted to obtain the target lip shape coefficient, so that the lip shape changes of the generated virtual digital image are richer and fuller.
[0082] Optionally, when establishing the mapping relationship between the candidate speakers, candidate entity words and candidate lip shape coefficients, a convolutional neural network lip movement model can be trained based on the candidate video data and candidate audio data of multiple candidate speakers, and the candidate video data and the corresponding candidate audio data are input into the lip movement model. For each speech window of 385ms, the speech is divided into 64 speech segments, and the autocorrelation coefficient of each speech segment with a length of 32 components is extracted to form a 64x32-dimensional feature as the speech feature input of the model. In addition to using the 64x32 autocorrelation speech feature as input, in order to distinguish different candidate speakers, different candidate speakers can be given different ID codes. Optionally, different candidate speakers can be ID-coded based on random Gaussian sampling. For example, each candidate speaker is represented by an ID code of length 32. For each candidate speaker, the vector information corresponding to the pronunciation lip shape of each word in the candidate audio data corresponding to the candidate speaker is established. A vector information library corresponding to the candidate speaker is established. During actual training, the training data uses the ID code corresponding to each candidate speaker, performs a dot multiplication operation with the corresponding vector information library vector, and is used together with the candidate audio data as the model input to train the convolutional neural network lip movement model. The output of the convolutional neural network lip movement model is the candidate lip shape coefficient corresponding to the candidate entity word in the candidate audio data corresponding to each candidate speaker.
[0083] S903 , generating a target lip-shape coefficient sequence corresponding to the virtual digital image based on the target lip-shape coefficients according to the pronunciation order.
[0084] Based on the order of the target entity words in the text data, all target lip coefficients are arranged to obtain a target lip coefficient sequence.
[0085] The disclosed embodiment obtains candidate video data of multiple candidate speakers, establishes a mapping relationship between candidate speakers, candidate entity words and candidate lip shape coefficients, is compatible with the timbre of multiple speakers, determines the target vector with the best opening and closing effect corresponding to the target speaker for optimizing the initial lip shape coefficient, optimizes the initial lip shape coefficient, increases the rhythm and richness of the virtual digital image's lip shape, and realizes the determination of the target lip shape coefficient sequence corresponding to the virtual digital image based on text data, so as to accurately drive the virtual digital image.
[0086] Figure 11 This is an exemplary embodiment of a method for driving a virtual digital image to sing shown in the present disclosure, such as Figure 11 As shown, based on the above embodiment, the method of driving a virtual digital image to sing a target song based on a target lip coefficient sequence includes the following steps:
[0087] S1101 , driving a virtual digital image based on a target lip coefficient sequence to generate multiple animation frames.
[0088] Based on the target lip coefficient sequence obtained above and in combination with the selected virtual digital image, a virtual digital image animation frame corresponding to each target lip coefficient in the target lip coefficient sequence is generated.
[0089] S1102, splicing the animation frames based on the target song, and driving the virtual digital image to sing the target song.
[0090] All animation frames are spliced together according to the order of the target entity words in the text data and played in this order to generate a virtual digital avatar animation. While playing all animation frames in this order, the target song obtained above is played synchronously, so that each target entity word currently playing in the target song corresponds to a target lip-sync coefficient in the current animation frame of the virtual digital avatar animation, thereby representing the virtual digital avatar singing the target song.
[0091] The disclosed embodiment realizes modeling of song melody and lyrics text to generate machine sounds with specific rhythms, and thereby drives the virtual digital image with precise and natural lip movements, enabling the virtual digital image to sing and increasing the usage scenarios of the virtual digital image.
[0092] Figure 12 This is a general flow chart of a method for driving a virtual digital image to sing, as shown in the present disclosure. Figure 12As shown, the method for driving a virtual digital image to sing comprises the following steps:
[0093] S1201, obtaining a virtual digital image, a target melody and text data.
[0094] S1202: Generate initial audio based on the text data and rhythm data.
[0095] S1203: Determine the target speaker and the pronunciation feature information of the target speaker.
[0096] S1204: Adjust the initial audio according to the pronunciation feature information to generate an initial song that matches the pronunciation features of the target speaker.
[0097] Regarding the implementation of steps S1201 to S1204, reference may be made to the introduction of the relevant parts in the above embodiment, which will not be repeated here.
[0098] S1205 , obtaining pitch data and frequency data of a target melody, and modifying the initial song based on the pitch data and the frequency data to obtain the target song.
[0099] S1206 , obtaining multiple target entity words of the text data and the pronunciation order of each target entity word.
[0100] S1207: Obtain a mapping relationship between candidate speakers, candidate entity words, and candidate lip shape coefficients.
[0101] S1208: Based on the target entity word and the target speaker, a mapping relationship is searched to determine an initial lip shape coefficient corresponding to the target entity word.
[0102] S1209: Obtain target lip shape coefficients based on the initial lip shape coefficients.
[0103] S1210 , generating a target lip shape coefficient sequence corresponding to the virtual digital image based on the target lip shape coefficients in accordance with the pronunciation order.
[0104] S1211, driving the virtual digital image to sing the target song based on the target lip coefficient sequence.
[0105] Regarding the implementation of steps S1205 to S1211, reference may be made to the introduction of the relevant parts in the above embodiment, which will not be repeated here.
[0106] The method for driving a virtual digital image to sing, proposed in an embodiment of the present disclosure, comprises obtaining a virtual digital image, a target melody, and text data; obtaining rhythm data of the target melody, and processing the text data based on the rhythm data to obtain an initial song; obtaining pitch data and frequency data of the target melody, and modifying the initial song based on the pitch data and frequency data to obtain a target song; determining a target lip-sync coefficient sequence corresponding to the virtual digital image based on the text data, and driving the virtual digital image to sing the target song based on the target lip-sync coefficient sequence. The embodiment of the present disclosure realizes modeling of song melody and lyrics text to generate a target song with a specific rhythm, and thereby accurately and naturally drives the virtual digital image to sing with its lip shape, thereby realizing the virtual digital image's singing and increasing the use scenarios of the virtual digital image.
[0107] Figure 13 Schematic diagram of a device for driving a virtual digital image to sing, as shown in the present disclosure. Figure 13 As shown, the device 1300 for driving a virtual digital image to sing includes an acquisition module 1301, a processing module 1302, a correction module 1303 and a driving module 1304, wherein:
[0108] An acquisition module 1301 is used to acquire a virtual digital image, a target melody, and text data;
[0109] The processing module 1302 is used to obtain rhythm data of the target melody and process the text data based on the rhythm data to obtain an initial song;
[0110] The correction module 1303 is used to obtain pitch data and frequency data of the target melody, and correct the initial song based on the pitch data and frequency data to obtain the target song;
[0111] The driving module 1304 is used to determine a target lip-sync coefficient sequence corresponding to the virtual digital image based on the text data, and drive the virtual digital image to sing the target song based on the target lip-sync coefficient sequence.
[0112] The device for driving a virtual digital image to sing, proposed in an embodiment of the present disclosure, obtains a virtual digital image, a target melody, and text data; obtains rhythm data of the target melody, and processes the text data based on the rhythm data to obtain an initial song; obtains pitch data and frequency data of the target melody, and modifies the initial song based on the pitch data and frequency data to obtain a target song; determines a target lip-sync coefficient sequence corresponding to the virtual digital image based on the text data, and drives the virtual digital image to sing the target song based on the target lip-sync coefficient sequence. The embodiment of the present disclosure realizes modeling of song melody and lyrics text to generate a target song with a specific rhythm, and thereby drives the virtual digital image to sing with a precise and natural lip-sync, thereby realizing the virtual digital image singing and increasing the use scenarios of the virtual digital image.
[0113] Furthermore, the processing module 1302 is also used to: generate initial audio based on text data and rhythm data; determine the target speaker and the pronunciation feature information of the target speaker; adjust the initial audio according to the pronunciation feature information to generate an initial song that matches the pronunciation features of the target speaker.
[0114] Furthermore, the processing module 1302 is also used to: obtain a selection instruction, the selection instruction is used to instruct the selection of a target speaker from multiple candidate speakers; determine the target speaker according to the selection instruction, and obtain the pronunciation feature information of the target speaker; or, determine the text feature information of the text data, determine the target speaker from multiple candidate speakers based on the text feature information, and obtain the pronunciation feature information of the target speaker.
[0115] Furthermore, the driving module 1304 is also used to: obtain multiple target entity words of the text data and the pronunciation order of each target entity word; obtain the target lip shape coefficient corresponding to each target entity word; and generate a target lip shape coefficient sequence corresponding to the virtual digital image based on the target lip shape coefficient according to the pronunciation order.
[0116] Furthermore, the driving module 1304 is also used to: obtain the mapping relationship between the candidate speakers, candidate entity words and candidate lip shape coefficients; based on the target entity word and the target speaker, query the mapping relationship to determine the initial lip shape coefficient corresponding to the target entity word; and obtain the target lip shape coefficient based on the initial lip shape coefficient.
[0117] Furthermore, the driving module 1304 is also used to optimize the initial lip shape coefficient to obtain the target lip shape coefficient.
[0118] Furthermore, the driving module 1304 is also used to: determine a target vector for optimizing the initial lip shape coefficient based on the candidate audio data of the target speaker; and optimize the initial lip shape coefficient based on the target vector to obtain the target lip shape coefficient.
[0119] Furthermore, the driving module 1304 is also used to: obtain candidate audio data of the target speaker, and obtain vector information corresponding to the pronunciation mouth shape of each pronunciation unit in the candidate audio data based on the candidate audio data; compare the pronunciation mouth shapes of each pronunciation unit in the candidate audio data, select the pronunciation mouth shape with the largest opening and closing amplitude as the target pronunciation mouth shape, and use the vector information corresponding to the target pronunciation mouth shape as the target vector.
[0120] Furthermore, the driving module 1304 is also used to: obtain a first weight corresponding to the target vector and a second weight corresponding to the initial lip shape coefficient; based on the first weight and the second weight, perform weighted processing on the target vector and the initial lip shape coefficient to obtain the target lip shape coefficient.
[0121] Furthermore, the driving module 1304 is also used to: perform the following steps for any candidate speaker: obtain the candidate video data of the candidate speaker, and generate a candidate animation sequence corresponding to the candidate speaker based on the candidate video data; obtain the candidate lip shape coefficient corresponding to each candidate animation frame in the candidate animation sequence; obtain the candidate audio data of the candidate speaker, and based on the candidate audio data, obtain multiple candidate entity words contained in the candidate audio data; based on the candidate speaker, the candidate entity words and the candidate lip shape coefficients, establish a mapping relationship between the candidate speaker, the candidate entity words and the candidate lip shape coefficients.
[0122] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0123] Figure 14 A schematic block diagram of an example electronic device 1400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0124] like Figure 14As shown, device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1402 or a computer program loaded from a storage unit 1408 into a random access memory (RAM) 1403. Various programs and data required for the operation of device 1400 can also be stored in RAM 1403. Computing unit 1401, ROM 1402, and RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to bus 1404.
[0125] Various components in device 1400 are connected to I / O interface 1405, including: an input unit 1406, such as a keyboard, mouse, etc.; an output unit 1407, such as various types of displays, speakers, etc.; a storage unit 1408, such as a magnetic disk, optical disk, etc.; and a communication unit 1409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1409 allows device 1400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0126] Computing unit 1401 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1401 performs the various methods and processes described above, such as the method for driving a virtual digital avatar to sing. For example, in some embodiments, the method for driving a virtual digital avatar to sing can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1408. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1400 via ROM 1402 and / or communication unit 1409. When the computer program is loaded into RAM 1403 and executed by computing unit 1401, one or more steps of the method for driving a virtual digital avatar to sing described above can be performed. Alternatively, in other embodiments, the computing unit 1401 may be configured in any other appropriate manner (eg, by means of firmware) to execute the method of driving the virtual digital image to sing.
[0127] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0128] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0129] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0131] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0132] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0133] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0134] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for driving a virtual digital image to sing, comprising: obtaining a virtual digital image, a target melody, and text data; Acquire rhythm data of the target melody, and process the text data based on the rhythm data to obtain an initial song, wherein the rhythm data includes rhythm point positions and rhythm durations; Acquiring pitch data and frequency data of the target melody, and modifying the initial song based on the pitch data and the frequency data to obtain the target song; A target lip-sync coefficient sequence corresponding to the virtual digital image is determined based on the text data, and the virtual digital image is driven to sing the target song based on the target lip-sync coefficient sequence.
2. The method according to claim 1, wherein The processing of the text data based on the rhythm data to obtain an initial song includes: generating an initial audio based on the text data and the rhythm data; Determining a target speaker and pronunciation feature information of the target speaker; The initial audio is adjusted according to the pronunciation feature information to generate the initial song that matches the pronunciation feature of the target speaker.
3. The method according to claim 2, wherein: The determining of the target speaker and the pronunciation feature information of the target speaker includes: Obtaining a selection instruction, the selection instruction is used to instruct to select the target speaker from multiple candidate speakers; determining the target speaker according to the selection instruction, and obtaining pronunciation feature information of the target speaker; or, The text feature information of the text data is determined, the target speaker is determined from a plurality of candidate speakers based on the text feature information, and pronunciation feature information of the target speaker is acquired.
4. The method according to claim 1, wherein The step of determining a target lip-sync coefficient sequence corresponding to the virtual digital image based on the text data includes: Acquire multiple target entity words and a pronunciation order of each target entity word from the text data; Obtaining a target lip shape coefficient corresponding to each target entity word; According to the pronunciation order and based on the target lip shape coefficients, a target lip shape coefficient sequence corresponding to the virtual digital image is generated.
5. The method according to claim 4, wherein The step of obtaining a target lip shape coefficient corresponding to each target entity word includes: Obtaining the mapping relationship between candidate speakers, candidate entity words, and candidate lip shape coefficients; Based on the target entity word and the target speaker, query the mapping relationship to determine the initial lip shape coefficient corresponding to the target entity word; Based on the initial lip shape coefficient, the target lip shape coefficient is obtained.
6. The method according to claim 5, wherein: The acquiring the target lip shape coefficient based on the initial lip shape coefficient comprises: The initial lip shape coefficient is optimized to obtain the target lip shape coefficient.
7. The method according to claim 6, wherein: The step of optimizing the initial lip shape coefficient to obtain the target lip shape coefficient includes: Determining a target vector for optimizing the initial lip shape coefficients based on the candidate audio data of the target speaker; Based on the target vector, the initial lip coefficient is optimized to obtain the target lip coefficient.
8. The method according to claim 7, wherein: The determining of a target vector for optimizing the initial lip coefficients includes: Acquire candidate audio data of the target speaker, and acquire vector information corresponding to the pronunciation mouth shape of each pronunciation unit in the candidate audio data based on the candidate audio data; The pronunciation mouth shapes of the pronunciation units in the candidate audio data are compared, and the pronunciation mouth shape with the largest opening and closing amplitude is selected as the target pronunciation mouth shape, and the vector information corresponding to the target pronunciation mouth shape is used as the target vector.
9. The method according to claim 7 or 8, wherein The step of optimizing the initial lip shape coefficient based on the target vector to obtain the target lip shape coefficient includes: Obtaining a first weight corresponding to the target vector and a second weight corresponding to the initial lip shape coefficient; Based on the first weight and the second weight, weighted processing is performed on the target vector and the initial lip shape coefficient to obtain the target lip shape coefficient.
10. The method according to claim 5, wherein The obtaining of the mapping relationship between the candidate speakers, the candidate entity words and the candidate lip shape coefficients includes: For any candidate speaker, perform the following steps: Obtaining candidate video data of the candidate speaker, and generating a candidate animation sequence corresponding to the candidate speaker based on the candidate video data; Obtaining a candidate lip-sync coefficient corresponding to each candidate animation frame in the candidate animation sequence; Acquire candidate audio data of the candidate speaker, and based on the candidate audio data, acquire multiple candidate entity words contained in the candidate audio data; Based on the candidate speaker, the candidate entity word and the candidate lip shape coefficient, the mapping relationship between the candidate speaker, the candidate entity word and the candidate lip shape coefficient is established.
11. A device for driving a virtual digital image to sing, comprising: An acquisition module, used to acquire a virtual digital image, a target melody and text data; a processing module, configured to obtain rhythm data of the target melody and process the text data based on the rhythm data to obtain an initial song, wherein the rhythm data includes rhythm point positions and rhythm durations; a correction module, configured to obtain pitch data and frequency data of the target melody, and correct the initial song based on the pitch data and the frequency data to obtain the target song; A driving module is used to determine a target lip-sync coefficient sequence corresponding to the virtual digital image based on the text data, and drive the virtual digital image to sing the target song based on the target lip-sync coefficient sequence.
12. The device according to claim 11, wherein The processing module is further configured to: generating an initial audio based on the text data and the rhythm data; Determining a target speaker and pronunciation feature information of the target speaker; The initial audio is adjusted according to the pronunciation feature information to generate the initial song that matches the pronunciation feature of the target speaker.
13. The device according to claim 12, wherein The processing module is further configured to: Obtaining a selection instruction, the selection instruction being used to instruct selection of the target speaker from a plurality of candidate speakers; determining the target speaker according to the selection instruction, and obtaining pronunciation feature information of the target speaker; or, The text feature information of the text data is determined, the target speaker is determined from a plurality of candidate speakers based on the text feature information, and pronunciation feature information of the target speaker is acquired.
14. The device according to claim 11, wherein The driving module is further used for: Acquire multiple target entity words and a pronunciation order of each target entity word from the text data; Obtaining a target lip shape coefficient corresponding to each target entity word; According to the pronunciation order and based on the target lip shape coefficients, a target lip shape coefficient sequence corresponding to the virtual digital image is generated.
15. The device according to claim 14, wherein The driving module is further used for: Obtaining the mapping relationship between candidate speakers, candidate entity words, and candidate lip shape coefficients; Based on the target entity word and the target speaker, query the mapping relationship to determine the initial lip shape coefficient corresponding to the target entity word; Based on the initial lip shape coefficient, the target lip shape coefficient is obtained.
16. The device according to claim 15, wherein The driving module is further used for: The initial lip shape coefficient is optimized to obtain the target lip shape coefficient.
17. The device according to claim 16, wherein The driving module is further used for: Determining a target vector for optimizing the initial lip shape coefficients based on the candidate audio data of the target speaker; Based on the target vector, the initial lip coefficient is optimized to obtain the target lip coefficient.
18. The device according to claim 17, wherein The driving module is further used for: Acquire candidate audio data of the target speaker, and acquire vector information corresponding to the pronunciation mouth shape of each pronunciation unit in the candidate audio data based on the candidate audio data; The pronunciation mouth shapes of the pronunciation units in the candidate audio data are compared, and the pronunciation mouth shape with the largest opening and closing amplitude is selected as the target pronunciation mouth shape, and the vector information corresponding to the target pronunciation mouth shape is used as the target vector.
19. The device according to claim 17 or 18, wherein The driving module is further used for: Obtaining a first weight corresponding to the target vector and a second weight corresponding to the initial lip shape coefficient; Based on the first weight and the second weight, weighted processing is performed on the target vector and the initial lip shape coefficient to obtain the target lip shape coefficient.
20. The apparatus according to claim 15, wherein The driving module is further used for: For any candidate speaker, perform the following steps: Obtaining candidate video data of the candidate speaker, and generating a candidate animation sequence corresponding to the candidate speaker based on the candidate video data; Obtaining a candidate lip-sync coefficient corresponding to each candidate animation frame in the candidate animation sequence; Acquire candidate audio data of the candidate speaker, and based on the candidate audio data, acquire multiple candidate entity words contained in the candidate audio data; Based on the candidate speaker, the candidate entity word and the candidate lip shape coefficient, the mapping relationship between the candidate speaker, the candidate entity word and the candidate lip shape coefficient is established.
21. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
23. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Audio synthesis method and device, storage medium and computer device
CN110189741A
Text and audio-based real-time face reenactment
WO2020150688A1