Video generation method, electronic equipment, readable storage medium and program product
By generating audio features and face key points images and rendering faces, the problem of inaccurate synchronization between mouth shape and audio in digital human technology is solved, and the accurate synchronization between mouth shape and audio in video is achieved and the smooth transition of face images is improved, which is augmented with the reality and quality of the video.
Patent Information
- Application Number
- CN202510154065.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, when digital human technology generates virtual characters, the mouth shape and audio synchronization is not accurate enough, resulting in jumping or incoherent face images in the generated video.
By acquiring audio data, M-frame audio features are generated, and corresponding N-frame face key point images are generated based on these features. For audio features adjacent to each m frame, face rendering is performed to obtain face images, and finally encode multi-frame face images and audio data to generate synchronized videos.
Accurate synchronization of mouth shape and audio in video is achieved, avoiding jumping or incoherence of face images, making the generated video more vivid, realistic and natural.
Smart Images

Figure CN120034705A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to technical fields such as artificial intelligence, and in particular to a video generation method, an electronic device, a readable storage medium, and a program product. Background Art
[0002] Digital human technology can create highly realistic virtual characters that not only look like real people, but can also simulate various human behaviors, expressions, and voices. Digital human technology has been widely used in many fields such as the metaverse, live broadcasting, education, culture, tourism, and finance, and has effectively promoted the development of the digital economy.
[0003] In the related art, computer graphics and motion capture are used to generate virtual characters, but there is a problem that the lip shape and audio synchronization are not accurate enough. Summary of the invention
[0004] The present disclosure provides a video generation method, an electronic device, a readable storage medium, and a program product.
[0005] According to one aspect of the present disclosure, there is provided a video generation method, comprising: Get audio data; Generate M frames of audio features according to the audio data; Generate N frames of facial key point images corresponding to the M frames of audio features according to the M frames of audio features; For each m frames of audio features that are adjacent to each other in the M frames of audio features, perform face rendering on n frames of facial key point images corresponding to the m frames of audio features that are adjacent to each other according to the m frames of audio features that are adjacent to each other, to obtain a facial image, wherein M, N, m and n are integers greater than 1 respectively; and Encode multiple frames of the face image and the audio data to obtain a video.
[0006] According to at least one embodiment of the present disclosure, a video generation method generates M frames of audio features according to the audio data, including: Extracting the Mel spectrum corresponding to the audio data; and The Mel spectrum is encoded to obtain the M-frame audio features.
[0007] According to at least one embodiment of the present disclosure, a video generation method generates N frames of facial key point images corresponding to the M frames of audio features according to the M frames of audio features, including: Inputting the M frames of audio features into a trained first model, predicting the facial key point information corresponding to the audio features through the first model, and obtaining N frames of facial key point information corresponding to the M frames of audio features; and The N frames of facial key point information are respectively drawn on drawings to obtain the N frames of facial key point images.
[0008] According to at least one embodiment of the present disclosure, a video generation method performs face rendering on n frames of face key point images corresponding to the m frames of audio features that are adjacent to each other, to obtain a face image, including: splicing the audio features of each m frames that are adjacent to each other to obtain the target audio features; splicing n frames of facial key point images corresponding to the m frames of audio features that are adjacent to each other to obtain a target facial key point image; and Performing face rendering on the target face key point image according to the target audio feature to obtain the face image.
[0009] According to at least one embodiment of the video generation method of the present disclosure, performing face rendering on the target face key point image according to the target audio feature to obtain the face image includes: The target audio features and the target facial key point image are input into a trained second model, and the face rendering is performed through the second model to obtain the face image.
[0010] According to at least one embodiment of the present disclosure, a video generation method encodes multiple frames of facial images and the audio data to obtain a video, including: Get the preset image; splicing the multiple frames of face images into the preset image respectively to obtain multiple frames of target images; and A plurality of frames of the target image and the audio data are encoded to obtain the video.
[0011] According to at least one embodiment of the present disclosure, a video generation method for acquiring audio data includes: Get text messages; and Perform speech synthesis on the text information to obtain the audio data.
[0012] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, so that the processor executes the video generation method of any one embodiment of the present disclosure.
[0013] According to another aspect of the present disclosure, a readable storage medium is provided, wherein the readable storage medium stores execution instructions, and when the execution instructions are executed by a processor, the video generating method according to any embodiment of the present disclosure is implemented.
[0014] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the video generating method according to any one of the embodiments of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings illustrate exemplary embodiments of the present disclosure and together with the description serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0016] Figure 1 It is a flowchart of a video generation method according to an embodiment of the present invention.
[0017] Figure 2 It is a schematic diagram of a process of acquiring audio data according to an embodiment of the present disclosure.
[0018] Figure 3 It is a schematic diagram of a process of generating audio features according to an embodiment of the present disclosure.
[0019] Figure 4 It is a schematic diagram of a process of generating a facial key point image according to an embodiment of the present disclosure.
[0020] Figure 5 It is a schematic diagram of a process of generating a facial image according to an embodiment of the present disclosure.
[0021] Figure 6 It is a schematic diagram of a process of generating a video according to an embodiment of the present disclosure.
[0022] Figure 7 It is a flowchart of a video generation method according to another embodiment of the present invention.
[0023] Figure 8 It is a schematic block diagram of the structure of a video generating device according to an embodiment of the present disclosure.
[0024] Fig. 9 It is a schematic block diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] The present disclosure is further described in detail below in conjunction with the accompanying drawings and examples. It is understood that the specific examples described herein are only used to explain the relevant content, rather than to limit the present disclosure. It should also be noted that, for ease of description, only the parts related to the present disclosure are shown in the accompanying drawings.
[0026] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure can be combined with each other. The technical solution of the present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0027] Digital human technology can not only enhance the user's sense of presence and reality in the communication process by creating realistic virtual characters, but also promote the digital transformation of related industries and inject new impetus into the development of the digital economy. In related technologies, virtual characters are generated based on computer graphics and motion capture. Due to the large amount of data demand, difficulty in obtaining sufficient data, high data acquisition costs, and difficulty in capturing facial movements, it is difficult for the final generated virtual character to accurately synchronize the audio and mouth shape during speech.
[0028] To this end, the present disclosure proposes a video generation method.
[0029] The video generation method disclosed in the present invention can be used for electronic devices to automatically generate speech videos with high synchronization rates between audio and virtual character lip shapes. In the present invention, electronic devices include but are not limited to servers, mobile phones, tablet computers, laptops, personal computers, wearable devices, ATMs, etc.
[0030] For the convenience of description and to make the technical solutions of the specific implementation methods of the present disclosure easier to understand, before describing the video generation method implemented in the present disclosure, the technical terms involved in the specific implementation methods of the present disclosure are explained as follows: A facial key point image is an image that marks the locations of specific parts of the face (such as eyes, nose, mouth, etc.). Facial key points are the locations of specific parts of the face.
[0031] A face image is an image that represents the detailed appearance information of a face (such as skin color, brightness, texture, contour, and shape of facial features).
[0032] Figure 1 FIG. 1 is a schematic diagram showing the overall process of a video generation method M100 according to an embodiment of the present disclosure. Figure 1 The method shown includes steps S110 to S150. The method can be executed by electronic devices such as a server, a mobile phone, and a computer.
[0033] Specifically, Figure 1 The methods shown include: S110, obtaining audio data; The audio data includes the audio in the video that is expected to be generated.
[0034] The audio data may be directly uploaded by the user to the electronic device, thereby being acquired by the electronic device; the audio data may also be acquired by the electronic device after processing the text information input by the user, which is not limited here.
[0035] S120, generating M frames of audio features according to the audio data; Audio data usually has a certain duration. Converting it into multi-frame audio features helps electronic devices better understand and process audio data.
[0036] Exemplarily, feature extraction can be performed directly on audio data to obtain M frames of audio features. For example, the audio data can be directly input into a pre-trained audio feature extraction model, and then M frames of audio features can be directly generated through the audio feature extraction model. Alternatively, the audio data can be processed to a certain extent, and then M frames of audio features can be generated based on the processed data.
[0037] S130, generating N frames of facial key point images corresponding to the M frames of audio features according to the M frames of audio features; The facial key point image can represent the facial key point information that matches the audio features. By generating the facial key point image corresponding to the audio features, it is helpful to more accurately determine the facial key point information at each stage of the audio, thereby providing data support for the lip shape and audio of the virtual character in the synchronized video.
[0038] M and N can be the same, that is, the number of frames of the audio features is the same as the number of frames of the facial key point image, and then the audio features correspond one-to-one to the facial key point image, which is convenient for subsequent data processing and is suitable for situations where the number of frames of audio features generated by the audio per second is greater than or equal to the preset frame number threshold.
[0039] M and N can also be different. For example, M is less than N. Then one frame of audio features corresponds to multiple frames of facial key point images. The facial key point information corresponding to one frame of audio features is represented by multiple frames of facial key point images, making the facial key point information richer. This is suitable for situations where the number of frames of audio features generated per second is less than a preset frame number threshold.
[0040] The preset frame number threshold can be set based on the speech rate. In one example, considering that the speech rate of Chinese is usually 3 to 5 words per second, and when 25 frames of audio features are generated per second, there is almost no situation where one frame of audio features involves two words, that is, one frame of facial key points corresponding to one frame of audio features will basically not change, so the preset frame number threshold is set to 25.
[0041] It should be noted that the specific numerical values mentioned in the present disclosure are only used as examples to illustrate the implementation of the present disclosure in detail, and should not be understood as limiting the present disclosure. In other examples or implementations or embodiments, other numerical values can be selected according to the present disclosure, and are not specifically limited here.
[0042] S140, for every m frames of audio features that are adjacent to each other in the M frames of audio features, perform face rendering on n frames of facial key point images corresponding to every m frames of audio features that are adjacent to each other according to the m frames of audio features that are adjacent to each other, to obtain a facial image, wherein M, N, m and n are integers greater than 1, M is greater than m, and N is greater than n; Since there is a corresponding relationship between the facial key point image and the audio feature, after determining the adjacent audio features of m frames, the corresponding n frames of adjacent facial key point images can also be determined. Furthermore, since the face rendering is based on multiple frames of audio features and multiple frames of facial key point images, and the multiple frames of audio features are adjacent to each other, and the multiple frames of facial key point images are also adjacent to each other, the coherence between frames is strong, and the audio features correspond to the facial key point images, the rendered facial image can incorporate more contextual information, and has a smooth transition effect on the timeline, so that the mouth shape in the facial image is more synchronized with the audio, avoiding the phenomenon of jumping or discontinuity in the facial image in the final generated video.
[0043] The ratio of m to n is the same as that of M to N. If M and N are 1:1, then m and n are also 1:1. In one example, M and N are in a 1:1 relationship, and then the M frames of audio features can be sorted based on the time sequence, and a sliding window with a length of m and a step size of 1 is slid on the sorted M frames of audio features in turn, so as to determine the audio features of every m frames of the M frames of audio features that are adjacent to each other; similarly, the N frames of facial key point images can be sorted based on the time sequence, and a sliding window with a length of n and a step size of 1 is slid on the sorted N frames of facial key point images in turn, so as to determine the facial key point images of every n frames of the N frames of facial key point images that are adjacent to each other. Among them, m can be 7, so that context information can be effectively integrated into the facial image and computing costs can be saved.
[0044] S150: Encode multiple frames of face images and audio data to obtain a video.
[0045] Since the audio features of each m frames of M-frame audio features are adjacent to each other, they will participate in the face rendering, that is, multiple face renderings will be performed in total, so multiple frames of face images will be obtained. Then, video processing tools (such as FFmpeg) can be used to encode the multiple frames of face images and audio data, so as to obtain a video including face images and audio, and the mouth shape and audio are accurately synchronized.
[0046] In one example, multiple frames of facial images are encoded into a video stream based on a specified frame rate, and audio data is encoded into an audio stream that can be merged with the video stream, and then the video stream and the audio stream are merged to obtain a video.
[0047] The video generation method of the disclosed embodiment automatically generates N frames of facial key point images based on M frames of audio features of audio data, and performs facial rendering on n frames of facial key point images corresponding to m frames of audio features that are adjacent to each other according to the audio features that are adjacent to each other in the M frames of audio features, thereby obtaining facial images, and encoding multiple frames of facial images and audio data to obtain a video. Since the facial key point images are generated through audio features, the facial key point images can accurately represent the facial key point information that matches the audio. Because face rendering is based on multi-frame audio features and multi-frame facial key point images, and multi-frame audio features are adjacent to each other, and multi-frame facial key point images are also adjacent to each other, there is a strong coherence between frames, and the audio features correspond to the facial key point images. Therefore, the rendered facial image takes into account the timing information and can incorporate more contextual information, so that the final encoded video has a smooth transition effect on the timeline, the mouth shape and audio in the video are more synchronized, and the face image in the video is avoided from jumping or incoherent, making the video more vivid, realistic and natural.
[0048] Regarding step S110, in some embodiments of the present disclosure, it may include: Figure 2 Steps S111 and S112 are shown.
[0049] S111. Obtain text information.
[0050] Exemplarily, a user may input text information through a human-computer interaction interface of an electronic device.
[0051] S112: Perform speech synthesis on the text information to obtain audio data.
[0052] TTS (Text To Speech) tools can be used to perform speech synthesis on text information to obtain audio data.
[0053] Exemplarily, the text information is first preprocessed so that the text information can be cleaned and standardized so that subsequent modules can correctly understand the text content, wherein the preprocessing includes but is not limited to text cleaning, word segmentation, sentence segmentation and normalization, text cleaning is used to remove unnecessary characters, symbols or formatting tags (such as HTML tags), word segmentation and sentence segmentation are used to divide the text into words, phrases or sentences for subsequent processing, and normalization is used to convert non-standard text (such as numbers, abbreviations, dates, times, etc.) into a pronounceable form; then, the preprocessed text information is subjected to phonological analysis to generate a phoneme sequence suitable for speech synthesis, and predict the rhythm, pauses, stress and intonation of the speech and other prosodic information; then, the phoneme sequence and prosodic information are input into a deep learning model (such as Tacotron, FastSpeech, etc.) to predict acoustic features; then, the acoustic features are converted into actual audio waveforms through a vocoder (such as WaveRNN, WaveGlow, HiFi-GAN, etc.), and the audio waveforms are saved in a specific audio format (such as WAV, MP3, etc.) to obtain audio data.
[0054] The video generation method of the above embodiment converts text information into audio data through speech synthesis technology, expands the input form of the video generation method, realizes end-to-end generation from text to video, and improves the practicality and scope of application of the method. At the same time, the automatic generation of audio data through speech synthesis reduces the need for manual audio recording, improves the efficiency of video generation, and has a higher degree of automation.
[0055] Regarding step S120, in some embodiments of the present disclosure, it may include the following steps: Figure 3 Step S121 and step S122 are shown.
[0056] S121, extracting the Mel spectrum corresponding to the audio data.
[0057] Mel spectrum can be understood as the spectrum obtained by converting the audio data to the Mel frequency scale. The horizontal axis is time and the vertical axis is Mel frequency. Since the Mel frequency scale is a nonlinear scale that simulates the human ear's perception of sounds of different frequencies, the M-frame audio features obtained by extracting the Mel spectrum and then encoding it are more in line with human auditory perception, thus providing support for generating accurate facial key point images.
[0058] S122: Encode the Mel spectrum to obtain M frames of audio features.
[0059] Exemplarily, a whisper encoder may be used to encode the Mel spectrum, thereby obtaining M frames of audio features.
[0060] The video generation method of the above embodiment can effectively capture the key information in the audio by encoding the Mel spectrum corresponding to the audio data to obtain M frames of audio features, and provide high-quality data support for generating facial key point images. At the same time, the Mel spectrum has strong robustness to noise and environmental changes, which is conducive to improving the stability of the obtained audio features.
[0061] Regarding step S130, in some embodiments of the present disclosure, it may include the following steps: Figure 4 Step S131 and step S132 are shown.
[0062] S131. Input M frames of audio features into a trained first model, and predict facial key point information corresponding to the audio features through the first model to obtain N frames of facial key point information corresponding to the M frames of audio features.
[0063] The first model may be a model designed based on transfermer, the input of the first model includes audio features, and the output of the first model includes facial key point information. It is understandable that compared with training a model that directly generates facial key point images, training a model that directly generates facial key point information requires less cost and takes less time.
[0064] In some embodiments, the trained first model can be quantized through a quantization tool, thereby reducing the size of the first model and speeding up the inference speed, so that the quantized first model can respond quickly and run smoothly on terminal devices such as mobile phones and tablets, which is conducive to solving the problem that the video generation method is difficult to implement on the terminal device.
[0065] S132, drawing N frames of facial key point information on drawings respectively to obtain N frames of facial key point images.
[0066] Exemplarily, a frame of facial key point information may include the coordinates of multiple facial key points. Based on the coordinates, the multiple key points in the frame of facial key point information are drawn on a drawing of the same coordinate system, and the key points in the drawing are connected with curves through interpolation or fitting methods to obtain a frame of facial key point image.
[0067] The video generation method of the above-mentioned embodiment can automatically predict the facial key point information corresponding to the audio features, and draw the predicted facial key points in a drawing to obtain a facial key point image, thereby reducing the need for manual intervention, having a high degree of automation, and the generated facial key point image can be highly matched with the audio content.
[0068] Regarding step S130, in other implementations, the M frames of audio features may be input into a trained third model, and the facial key point images corresponding to the audio features may be predicted by the third model, thereby directly obtaining N frames of facial key point images corresponding to the M frames of audio features. The third model may also be a model designed based on transfermer, the input of the third model includes the audio features, and the output of the third model includes the facial key point images.
[0069] In some embodiments, the trained third model can be quantized through a quantization tool, thereby reducing the size of the third model and speeding up the inference speed, so that the quantized third model can respond quickly and run smoothly on terminal devices such as mobile phones and tablets, which is conducive to solving the problem that the video generation method is difficult to implement on the terminal device.
[0070] Regarding step S140, in some embodiments of the present disclosure, it may include: Figure 5 Steps S141 to S143 are shown.
[0071] S141. Concatenate two adjacent audio features of every m frames to obtain target audio features.
[0072] Exemplarily, based on the time sequence, the audio features of m frames that are adjacent to each other are spliced together to obtain a target audio feature, and then the audio features of m frames that are adjacent to each other are represented by one target audio feature to facilitate face rendering.
[0073] S142, splicing n frames of facial key point images corresponding to m frames of audio features that are adjacent to each other to obtain a target facial key point image.
[0074] Exemplarily, based on the chronological order, n frames of facial key point images corresponding to m frames of adjacent audio features are spliced together to obtain a target facial key point image, and then the n frames of facial key point images are represented by one target facial key point image to facilitate face rendering.
[0075] It is worth noting that step S141 may be executed before step S142, or after step S142, or simultaneously with step S142, and this is not limited in the present disclosure.
[0076] S143. Perform face rendering on the target face key point image according to the target audio features to obtain a face image.
[0077] Since the target facial key point image lacks detailed facial appearance information (such as skin color, brightness, and texture), and the detailed facial appearances corresponding to different speaking actions are also different, rendering the target facial key point image according to the target audio features can ensure that the obtained facial image has detailed facial appearance information and is highly matched with the audio.
[0078] The video generation method of the above embodiment enhances the correlation between different frames by splicing the audio features of each m frames and the facial key point images of n frames respectively, so that the generated facial image is more coherent and smooth. At the same time, by rendering the facial key point images through multi-frame audio features, it is possible to better handle complex voice rhythms and expression changes. In addition, by splicing the audio features of each m frames and the facial key point images of n frames respectively, the face rendering is performed, which avoids the redundant calculation caused by independent frame processing, which is conducive to saving the amount of calculation.
[0079] Regarding step S143, in some embodiments of the present disclosure, it may specifically be: inputting the target audio features and the target facial key point image into a trained second model, performing face rendering through the second model, and obtaining a facial image.
[0080] The second model can be obtained by designing a multimodal fusion network, a multimodal Transformer, and a conditional generative adversarial network. The input of the second model can include target audio features and target facial key point images, and the output of the second model can include facial images.
[0081] In some embodiments, the trained second model can be quantized through a quantization tool, thereby reducing the size of the second model and speeding up the inference speed, so that the quantized second model can respond quickly and run smoothly on terminal devices such as mobile phones and tablets, which is conducive to solving the problem that the video generation method is difficult to implement on the terminal device.
[0082] The video generation method of the above embodiment, by performing face rendering through the trained second model, can capture more detailed expression and action features, improve the realism and quality of the generated face image, and help improve the overall viewing experience of the video. In addition, the second model can be customized according to specific needs, which is conducive to achieving different styles or personalized face rendering.
[0083] Regarding step S150, in some embodiments of the present disclosure, it may include: Figure 6 Steps S151 to S153 are shown.
[0084] S151, obtaining a preset image.
[0085] S152, stitching the multiple frames of face images into the preset images respectively to obtain multiple frames of target images.
[0086] The preset image may include a background and / or a human body. By splicing a face image into the preset image, the target image obtained is more realistic and natural. A frame of a face image is spliced into the preset image to obtain a frame of a target image, and the number of frames of the target image is the same as the number of frames of the face image.
[0087] S153: Encode multiple frames of target images and audio data to obtain a video.
[0088] A video processing tool (such as FFmpeg) may be used to encode multiple frames of target image and audio data, thereby obtaining a video including a facial image and audio, with the mouth shape and audio accurately synchronized.
[0089] In one example, multiple frames of target images are encoded into a video stream based on a specified frame rate, and audio data is encoded into an audio stream that can be merged with the video stream, and then the video stream and the audio stream are merged to obtain a video.
[0090] The video generation method of the above embodiment first splices the face image into the preset image to obtain the target image, and then encodes the target image and audio data to obtain the video, making the video content richer and more vivid. At the same time, the preset image can be flexibly adjusted to support different scene settings (such as indoor, outdoor, virtual scenes, etc.), which can increase the diversity of the generated video.
[0091] Please combine Figure 7 In one example, the video generation method may include the following steps S201 to S215. The contents related to steps S201 to S215 can refer to the description of the above implementation method. For the sake of brevity, they will not be repeated here.
[0092] In step S201, text information is obtained.
[0093] In step S202, speech synthesis is performed on the text information to obtain audio data.
[0094] In step S203, the Mel spectrum corresponding to the audio data is extracted.
[0095] In step S204, the Mel spectrum is encoded to obtain M frames of audio features.
[0096] In step S205, M frames of audio features are input into the trained first model, and the facial key point information corresponding to the audio features is predicted by the first model to obtain N frames of facial key point information corresponding to the M frames of audio features.
[0097] In step S206, N frames of facial key point information are respectively drawn on drawings to obtain N frames of facial key point images.
[0098] In step S207, audio features of m frames adjacent to each other are determined from the audio features of the M frames based on the sliding window.
[0099] In step S208, n frames of adjacent facial key point images are determined from the N frames of facial key point images based on the sliding window. The determined m frames of audio features have a corresponding relationship with the n frames of facial key point images. It is worth noting that Figure 7 In the example, step S207 is executed before step S208. In other implementations, step S207 may also be executed after step S208, or executed simultaneously with step S208, which is not limited here.
[0100] In step S209, the audio features of each m frames that are adjacent to each other are concatenated to obtain the target audio features.
[0101] In step S210, n frames of facial key point images corresponding to m frames of adjacent audio features are spliced to obtain the target facial key point image. Figure 7 In the example, step S209 is executed before step S210. In other implementations, step S209 may also be executed after step S210, or executed simultaneously with step S210, which is not limited here.
[0102] In step S211, the target audio features and the target facial key point image are input into the trained second model, and the face rendering is performed through the second model to obtain a face image.
[0103] In step S212, it is determined whether the traversal of the audio features and the facial key point images is complete, if so, the process proceeds to step S213, otherwise, the process proceeds to step S207.
[0104] In step S213, a preset image is acquired.
[0105] In step S214, the obtained multiple frames of face images are respectively spliced into the preset image to obtain multiple frames of target images.
[0106] In step S215, multiple frames of target images and audio data are encoded to obtain a video.
[0107] Based on any of the above implementations, the present disclosure also provides a video generating device.
[0108] Figure 8 It is a schematic block diagram of the structure of a video generating device according to an embodiment of the present disclosure.
[0109] like Figure 8 As shown, the video generating device comprises: An acquisition module 110, used to acquire audio data; A first generating module 120, configured to generate M frames of audio features according to the audio data; The second generating module 130 is used to generate N frames of facial key point images corresponding to the M frames of audio features according to the M frames of audio features; A rendering module 140 is used to perform face rendering on n frames of facial key point images corresponding to m frames of audio features that are adjacent to each other in every m frames of audio features, according to the audio features that are adjacent to each other in every m frames, to obtain a facial image, wherein M, N, m and n are integers greater than 1 respectively; The encoding module 150 is used to encode multiple frames of face images and audio data to obtain a video.
[0110] The video generating device may be in the form of computer software, and each module of the video generating device may be implemented by a computer software module.
[0111] In some implementations of the present disclosure, the acquisition module 110 is used to: acquire text information; and perform speech synthesis on the text information to obtain audio data.
[0112] In some embodiments of the present disclosure, the first generating module 120 is used to: extract the Mel spectrum corresponding to the audio data; and encode the Mel spectrum to obtain M frames of audio features.
[0113] In some embodiments of the present disclosure, the second generation module 130 is used to: input M frames of audio features into a trained first model, predict the facial key point information corresponding to the audio features through the first model, and obtain N frames of facial key point information corresponding to the M frames of audio features; and draw the N frames of facial key point information on a drawing respectively to obtain N frames of facial key point images.
[0114] In some embodiments of the present disclosure, the rendering module 140 is used to: splice audio features that are adjacent to each other in every m frames to obtain target audio features; splice n frames of facial key point images corresponding to audio features that are adjacent to each other in every m frames to obtain target facial key point images; and perform face rendering on the target facial key point images according to the target audio features to obtain a facial image.
[0115] In some embodiments of the present disclosure, the rendering module 140 is used to: input the target audio features and the target facial key point image into a trained second model, perform facial rendering through the second model, and obtain a facial image.
[0116] In some embodiments of the present disclosure, the encoding module 150 is used to: obtain a preset image; splice multiple frames of facial images into the preset image respectively to obtain multiple frames of target images; and encode the multiple frames of target images and audio data to obtain a video.
[0117] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0118] The execution subject of the video generation method in the specific implementation manner of the present disclosure may be a server, a mobile phone, a computer or other electronic device.
[0119] Therefore, based on any one of the above embodiments, the present disclosure further provides an electronic device, which can execute the video generating method of any one of the above embodiments of the present disclosure.
[0120] Fig. 9 1 is a schematic block diagram of the structure of an electronic device 1000 according to an embodiment of the present disclosure.
[0121] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application and overall design constraints of the hardware. The bus 1100 connects various circuits including one or more processors 1200, memory 1300 and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripherals, voltage regulators, power management circuits, external antennas, etc.
[0122] The bus 1100 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the figure only uses one connecting line, but does not mean that there is only one bus or one type of bus.
[0123] The present disclosure also provides a readable storage medium, in which a computer program is stored, and the computer program is used to implement the above method when executed by a processor. "Readable storage medium" can be any device that can contain, store, communicate, propagate or transmit a program for use in an instruction execution system, device or equipment or in combination with these instruction execution systems, devices or equipment. More specific examples of readable storage media include the following: an electrical connection portion (electronic device) with one or more wirings, a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0124] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed, the process or function of the present disclosure is executed in whole or in part.
[0125] The computer program or instructions may be stored in a readable storage medium or transmitted from one readable storage medium to another readable storage medium, for example, the computer program or instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The readable storage medium may be any available medium that can be accessed or a data storage device such as a server, data center, etc. that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disk; it may also be a semiconductor medium, such as a solid state drive. The computer readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0126] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0127] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0128] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0130] In the description of this specification, the description with reference to terms such as "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples" means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.
[0131] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0132] Those skilled in the art should understand that the above embodiments are only for the purpose of clearly illustrating the present disclosure, and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications may be made based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. A video generation method, characterized in that: include: Get audio data; Generate M frames of audio features according to the audio data; Generate N frames of facial key point images corresponding to the M frames of audio features according to the M frames of audio features; For the audio features of every m frames that are adjacent to each other in the M frames of audio features, according to the audio features of every m frames that are adjacent to each other in the M frames, face rendering is performed on n frames of facial key point images corresponding to the audio features of every m frames that are adjacent to each other in the M frames, to obtain a facial image, wherein M, N, m and n are integers greater than 1 respectively; as well as Encode multiple frames of the face image and the audio data to obtain a video.
2. The video generation method according to claim 1, characterized in that: Generating M frames of audio features according to the audio data includes: Extracting the Mel spectrum corresponding to the audio data; and The Mel spectrum is encoded to obtain the M-frame audio features.
3. The video generation method according to claim 1, characterized in that: Generating N frames of facial key point images corresponding to the M frames of audio features according to the M frames of audio features, including: Inputting the M frames of audio features into a trained first model, predicting the facial key point information corresponding to the audio features through the first model, and obtaining N frames of facial key point information corresponding to the M frames of audio features; and The N frames of facial key point information are respectively drawn on drawings to obtain the N frames of facial key point images.
4. The video generation method according to claim 1, characterized in that: According to the audio features of every m frames that are adjacent to each other, face rendering is performed on n frames of face key point images corresponding to the audio features of every m frames that are adjacent to each other, to obtain a face image, including: splicing the audio features of each m frames that are adjacent to each other to obtain the target audio features; splicing n frames of facial key point images corresponding to the m frames of audio features that are adjacent to each other to obtain a target facial key point image; and Performing face rendering on the target face key point image according to the target audio feature to obtain the face image.
5. The video generation method according to claim 4, characterized in that: Performing face rendering on the target face key point image according to the target audio feature to obtain the face image includes: The target audio features and the target facial key point image are input into a trained second model, and the face rendering is performed through the second model to obtain the face image.
6. The video generation method according to claim 1, characterized in that: Encoding multiple frames of the face image and the audio data to obtain a video includes: Get the preset image; splicing the multiple frames of face images into the preset image respectively to obtain multiple frames of target images; and A plurality of frames of the target image and the audio data are encoded to obtain the video.
7. The video generation method according to claim 1, characterized in that: Get audio data, including: Get text messages; and Perform speech synthesis on the text information to obtain the audio data.
8. An electronic device, characterized in that: include: A memory storing execution instructions; as well as A processor, wherein the processor executes the execution instruction stored in the memory, so that the processor executes the video generation method according to any one of claims 1 to 7.
9. A readable storage medium, characterized in that: The readable storage medium stores execution instructions, which are used to implement the video generation method according to any one of claims 1 to 7 when executed by a processor.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the video generating method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Digital human video generation method and system based on audio driving
CN120602740A