Digital human video synthesis method, server and storage medium

The method simplifies digital human video synthesis by allowing users to input descriptions for automated video generation, addressing complexity and resource limitations, and improving video naturalness and realism through detailed facial and head movements.

CN120321350APending Publication Date: 2025-07-15CHONGQING ZHONGKE YUNCONG TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510465561.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing digital human video synthesis platform is complex in operation, requires professional skills, and is difficult to meet the personalized and innovative needs of users.

Method used

By inputting image description text and video description information on the interactive interface, a digital human image is generated using a preset keyword library and image generation model, and rendering information is generated based on facial information and voice information, and finally a digital human video is synthesized.

Benefits of technology

It simplifies the operation process, meets users' personalized needs, improves the natural fluency and fidelity of the video, and reduces the dependence on professional skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321350A_ABST
    Figure CN120321350A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, particularly provides a digital human video synthesis method, a server and a storage medium, and aims to solve the problems of how to simplify the operation of digital human video synthesis and how to meet the personalized demand of a user for a video. The method provided by the invention comprises the following steps: extracting a first keyword in an image description text, obtaining a second keyword with similar semantics, and inputting the first keyword and the second keyword into an image generation model for image generation to obtain a first digital human image; converting the video description information into voice information, and generating rendering information according to the face information of the digital person in the first digital person image and the voice information; generating a second digital human image according to the rendering information and the first digital human image; and performing video synthesis on the voice information and the second digital human image to obtain a digital human video. Based on the method, the operation of synthesizing the digital human video can be simplified, the individual demand of a user on the video can be met, and meanwhile, the natural fluency and fidelity of the video are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a digital human video synthesis method, a server and a storage medium. Background Art

[0002] Digital Human video synthesis technology has become an important part of the field of digital content creation and is widely used in various media production, advertising promotion, education and training, etc. However, the existing conventional video synthesis platform has a series of problems, which seriously limit its application convenience and compatibility.

[0003] Specifically, first of all, existing video synthesis platforms usually require users to perform a lot of manual operations, such as designing images, adjusting special effects, editing, etc., which not only requires users to have certain professional skills, but also has high operational complexity. Therefore, for ordinary users without relevant background knowledge and skills, using these platforms for video synthesis will face huge barriers. Secondly, the content created by users on conventional video synthesis platforms is often limited by the type and quality of existing resources. Although some platforms provide rich material libraries and templates for users to choose from, these predefined resources often cannot meet users' needs for personalization and innovation.

[0004] Accordingly, the art needs a new technical solution to solve the above problems. Summary of the invention

[0005] In order to overcome the above-mentioned defects, the present application is proposed to solve or at least partially solve the following technical problems: how to simplify the operation of digital human video synthesis and meet the personalized needs of users for digital human videos.

[0006] In a first aspect, a digital human video synthesis method is provided, the method comprising:

[0007] In response to a user input operation on an image description text on an interactive interface, acquiring the image description text input by the user, wherein the image description text is used to describe feature information of a digital human image;

[0008] Extracting a first keyword from the image description text, and acquiring a second keyword having semantic similarity to the first keyword from a preset keyword library, inputting the first keyword and the second keyword into a preset image generation model for image generation, and obtaining a first digital human image;

[0009] Acquiring facial information of the digital human in the first digital human image;

[0010] In response to a user's input operation on video description information in an interactive interface, obtain the video description information input by the user, and convert the video description information into multi-frame voice information;

[0011] According to the facial information and the multi-frame voice information, respectively generate rendering information corresponding to each frame of voice information, where the rendering information includes the mouth shape, expression, and head movement of the digital human;

[0012] According to the rendering information corresponding to each frame of voice information and the first digital human image, respectively generate second digital human images corresponding to each frame of voice information;

[0013] Perform video synthesis on each frame of voice information and the second digital human images corresponding to each frame of voice information to obtain a digital human video.

[0014] In a technical solution of the above digital human video synthesis method, the obtaining of the second keyword semantically similar to the first keyword from a preset keyword library includes: obtaining the first word vector feature of the first keyword, and calculating the similarity between the first word vector feature and the second word vector feature of the second keyword in the preset keyword library to obtain the similarity between the first keyword and the second keyword; obtaining the second keyword with the similarity greater than a set threshold as the second keyword that is semantically similar.

[0015] In a technical solution of the above digital human video synthesis method, the obtaining of the facial information of the digital human in the first digital human image includes: identifying the face area of the digital human in the first digital human image; extracting the face key points and facial bone points of the face area; obtaining the facial information according to the face key points and the facial bone points.

[0016] In a technical solution of the above digital human video synthesis method, before the obtaining of the facial information of the digital human in the first digital human image, the method further includes: optimizing the image information of the first digital human image, where the image information includes color, background, and the facial information of the digital human.

[0017] In a technical solution of the above digital human video synthesis method, the generating of the rendering information corresponding to each frame of voice information according to the facial information and the multi-frame voice information includes: respectively generating the mouth shape of the digital human corresponding to each frame of voice information according to the pronunciation of each frame of voice information and based on the facial information; using a pre-trained deep learning model to respectively generate the expression and head movement of the digital human corresponding to each frame of voice information according to the facial information and each frame of voice information.

[0018] In one technical solution of the above digital human video synthesis method, the video synthesis of each frame of voice information and the second digital human image corresponding to each frame of voice information includes: generating subtitles according to the video description information, and synchronizing the subtitles with the mouth shapes of the digital humans corresponding to each frame of voice information; performing video synthesis on the synchronized subtitles, each frame of voice information, and the second digital human image corresponding to each frame of voice information to obtain a digital human video.

[0019] In one technical solution of the above digital human video synthesis method, the video synthesis of each frame of voice information and the second digital human image corresponding to each frame of voice information includes: in response to a user's configuration operation on an image element on an interactive interface, obtaining the image element to be added and its configuration information, where the configuration information includes the time and position at which the image element appears in the digital human video; performing video synthesis on the image element, each frame of voice information, and the second digital human image corresponding to each frame of voice information according to the configuration information to obtain a digital human video.

[0020] In one technical solution of the above digital human video synthesis method, the preset keyword library is a keyword library constructed based on a knowledge graph; and / or, the video description information is audio information or text information.

[0021] In a second aspect, a server is provided, which includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the method described in any one of the technical solutions provided in the above first aspect is implemented.

[0022] In a third aspect, a computer-readable storage medium is provided, in which multiple program codes are stored, and the program codes are adapted to be loaded and run by a processor to execute the method described in any one of the technical solutions provided in the above first aspect.

[0023] One or more of the above technical solutions of the present application have at least one or more of the following Beneficial effects:

[0024] In a technical solution of implementing the digital human video synthesis method provided by the present application, in response to a user's input operation of image description text on an interactive interface, the image description text input by the user is obtained. The image description text is used to describe the feature information of the digital human image; the first keyword in the image description text is extracted, and the second keyword semantically similar to the first keyword is obtained from a preset keyword library. The first keyword and the second keyword are input into a preset image generation model for image generation to obtain the first digital human image; the facial information of the digital human in the first digital human image is obtained; in response to the user's input operation of video description information on the interactive interface, the video description information input by the user is obtained, and the video description information is converted into multi-frame voice information; according to the facial information and the multi-frame voice information, rendering information corresponding to each frame of voice information is respectively generated. The rendering information includes the mouth shape, expression, and head movement of the digital human; according to the rendering information corresponding to each frame of voice information and the first digital human image, second digital human images corresponding to each frame of voice information are respectively generated; the video synthesis of each frame of voice information and the second digital human image corresponding to each frame of voice information is performed to obtain a digital human video.

[0025] In the above implementation, the user only needs to input the image description text and the video description information in sequence through the interactive interface, and the digital human video can be automatically generated, greatly simplifying the operation process. Moreover, the user does not need to have professional skills such as image design, special effect adjustment, and video editing. In addition, the image description text reflects the user's personalized and innovative requirements for the digital human image. The digital human image automatically generated by the image generation model according to the image description text actually meets the above requirements. There is no need to preset a material library and templates for the user to choose, and then generate the digital human image according to the materials or templates selected by the user. In this way, there is no problem that the preset material library and templates are difficult to meet the user's personalized and innovative requirements.

[0026] In addition, the above implementation can respectively generate the rendering information corresponding to each frame of voice information according to the facial information and the multi-frame voice information, that is, generate the rendering information frame by frame, improving the refinement degree of the rendering information. In this way, it is beneficial to improve the natural smoothness and vividness of the video when performing video synthesis according to the rendering information; further, the rendering information includes various rendering details of the digital human (including mouth shape, expression, and head movement). In this way, when performing video synthesis, various rendering information of the digital human can be taken into account at the same time, further improving the natural smoothness and vividness of the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Referring to the accompanying drawings, the disclosure of the present application will become easier to understand. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present application. Among them:

[0028] Figure 1 It is a schematic diagram of the main steps of a digital human video synthesis method according to an embodiment of the present application;

[0029] Figure 2 It is a schematic diagram of an interactive interface according to an embodiment of the present application;

[0030] Figure 3 It is a schematic diagram of the main steps of obtaining a second keyword semantically similar to the first keyword according to an embodiment of the present application;

[0031] Figure 4 It is a schematic diagram of the main steps of obtaining the facial information of the digital human in the first digital human image according to an embodiment of the present application;

[0032] Figure 5 It is a schematic diagram of the main steps of generating rendering information corresponding to each frame of voice information according to an embodiment of the present application;

[0033] Figure 6 It is a schematic diagram of an image element according to an embodiment of the present application;

[0034] Figure 7 It is a schematic diagram of the main steps of a digital human video synthesis method according to another embodiment of the present application;

[0035] Figure 8 It is a schematic diagram of the main structure of a server according to an embodiment of the present application;

[0036] Figure 9 It is a schematic diagram of the main structure of a digital human video synthesis system according to an embodiment of the present application;

[0037] Figure 10 It is a schematic diagram of the main structure of a digital human video synthesis system according to another embodiment of the present application.

[0038] Reference numerals:

[0039] 11: Memory; 12: Processor; 21: Keyword generation module; 22: Image rendering module; 23: Video synthesis module; 24: Image preprocessing module; 25: Pipeline module. Detailed implementation manners

[0040] The following describes some embodiments of the present application with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principle of the present application and are not intended to limit the protection scope of the present application.

[0041] In the description of this application, a "module" and a "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various appropriate sensors, communication ports, memory, and may also include a software part, such as program code, or may be a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other appropriate processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. A computer-readable storage medium includes any appropriate medium that can store program code, such as magnetic disks, hard disks, optical discs, flash memories, read-only memories, random access memories, and so on. The term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B.

[0042] The following describes an embodiment of the digital human video synthesis method provided by this application.

[0043] Refer to the attached Figure 1 , Figure 1 which is a schematic diagram of the main steps of the digital human video synthesis method according to an embodiment of this application. As Figure 1 shown, the digital human video synthesis method in the embodiment of this application mainly includes the following steps S101 to S108.

[0044] Step S101: In response to the user's input operation of the image description text on the interaction interface, obtain the image description text input by the user. The image description text is used to describe the feature information of the digital human image, and the feature information includes the human body features of the digital human and the features of the environment where the digital human is located.

[0045] The digital human video synthesis method provided by this application can be applied to a server. The above interaction interface can be displayed on the display interface of the server. The interaction interface can display an area for inputting the image description text, and the user can input the image description text in this area. The user can freely set the specific content of the image description text according to their actual needs. For example, the content of the image description text is "Draw a 20-year-old Chinese handsome guy with silver-gray hair, brown eyes, a street style of black and blue, on the moon's surface, near-future city, OC rendering, cyberpunk style". The meaning of OC rendering is an image style rendered using the Octane Render renderer.

[0046] Step S102: Extract the first keyword from the image description text, and obtain a second keyword with a similar semantics to the first keyword from a preset keyword library.

[0047] In this embodiment, a conventional keyword extraction method can be used to extract the first keywords from the image description text. In some embodiments, a pre-trained deep learning model (such as a large language model) can be used to perform semantic understanding and analysis on the image description text, obtain the entities in the image description text and the relationships between different entities, and determine the first keywords of the image description text according to the entities and the relationships between different entities. The first keywords can include entities and the relationships between different entities. For example, if the relationship between two entities is an action (such as holding, carrying, etc.), the first keyword will include both of these two entities and the action. An entity can refer to a specific object, such as a certain type of person, a certain type of object, a certain type of thing, etc. In this embodiment, a conventional model training method can be used to train the above-mentioned deep learning model. This embodiment does not specifically limit the training method of the model, as long as it has the above-mentioned speech understanding and analysis capabilities.

[0048] If the semantics of the first keyword is an action, then the second keyword similar in semantics to it represents an action similar to the action. For example, if the first keyword is waving, the second keyword can be waving one's hand.

[0049] Step S103: Input the first keyword and the second keyword into a preset image generation model for image generation to obtain the first digital human image.

[0050] The preset image generation model is a pre-trained deep learning model, which has the ability to generate images according to the input text. In this embodiment, a conventional model training method can be used to train the deep learning model. This embodiment does not specifically limit the training method of the model, as long as it has the above-mentioned image generation ability.

[0051] Step S104: Obtain the facial information of the digital human in the first digital human image.

[0052] Step S105: In response to the user's input operation on the video description information on the interaction interface, obtain the video description information input by the user, and convert the video description information into multi-frame voice information.

[0053] The video description information is used to describe the content spoken by the digital human in the digital human video. The video description information is audio information or text information, that is, the spoken content is expressed in the form of audio or text. For example, the video description information is a piece of text "Company A is an artificial intelligence technology enterprise headquartered in B, founded by C, and its business covers fields such as D".

[0054] If the video description information is text information, the TTS (Text-to-Speech) technology can be used to convert the video description information into multi-frame voice information.

[0055] The interactive interface can display an area for inputting video description information, where users can input video description information. For example, refer to Figure 2 the interactive interface shown. There is a text preview area in this interactive interface, where users can input video description information. Additionally, in some embodiments, users can also configure the voice of the digital human on the interactive interface. The method provided in this application can obtain the selected voice type in response to the user's selection of the voice type on the interactive interface, and in subsequent step S108, video synthesis can be performed based on this voice type and other information (including the voice information and the second digital human image mentioned in step S108), so that the digital human in the digital human video can output the specific content of the video description information according to the above voice type, that is, perform voice broadcast according to this voice type. For the voice type in this embodiment, no specific limitation is made. For example, Figure 2 exemplarily shows various voice types such as female voice and male voice.

[0056] Step S106: Generate rendering information corresponding to each frame of voice information respectively according to the facial information and multiple frames of voice information. The rendering information includes the mouth shape, expression, and head movement of the digital human.

[0057] The mouth shape corresponding to each frame of voice information is respectively matched with the pronunciation of each frame of voice information, that is, this mouth shape can simulate the mouth shape presented by a human when making the sound of this voice information. The expression can include smiling, serious, etc., and the head movement can include nodding, shaking the head, etc.

[0058] Step S107: Generate a second digital human image corresponding to each frame of voice information respectively according to the rendering information corresponding to each frame of voice information and the first digital human image.

[0059] For each frame of voice information, in the second digital human image corresponding to the voice information, the mouth shape, expression, and head movement of the digital human are respectively the same as the mouth shape, expression, and head movement in the rendering information corresponding to this voice information.

[0060] Step S108: Perform video synthesis on each frame of voice information and the second digital human image corresponding to each frame of voice information to obtain a digital human video.

[0061] The digital human video includes multiple video frames arranged in sequence according to time from the earliest to the latest. Each video frame is respectively synthesized by a frame of voice information and its corresponding second digital human image, where the time of the video frame is the time corresponding to the voice information in the video frame.

[0062] For the video format of the digital human video, this embodiment does not make specific limitations. In some embodiments, the format of the digital human video can be MOV (QuickTime Movie) format, and the digital human video in this format can be applied to various scenarios such as social media, short videos, and virtual customer service.

[0063] Based on the method described in the above steps S101 to S108, the user only needs to input the image description text and the video description information in sequence through the interaction interface, and then the digital human video can be automatically generated, which greatly simplifies the operation process and does not require the user to have professional skills such as designing images, adjusting special effects, and editing. In addition, the image description text reflects the user's personalized and innovative needs for the digital human image, and the digital human image automatically generated by the image generation model according to the image description text actually meets the above needs. Moreover, the above method can generate the rendering information corresponding to each frame of voice information respectively according to the facial information and multiple frames of voice information, that is, generate the rendering information frame by frame, improving the refinement degree of the rendering information, which is beneficial to improving the natural smoothness and realism of the video when synthesizing the video according to the rendering information; the rendering information includes various rendering details of the digital human (including mouth shape, expression, and head movement), so that various rendering information of the digital human can be taken into account simultaneously when synthesizing the video, further improving the natural smoothness and realism of the video.

[0064] Next, the embodiments of the digital human video synthesis method provided by the present application will be further described, specifically, steps S102, S104, S106, and S108 will be described.

[0065] 1. Explanation of step S102.

[0066] In some embodiments of the above step S102, it can be obtained through Figure 3 the following steps S1021 to S1023 shown, the second keyword semantically similar to the first keyword.

[0067] Step S1021: Obtain the first word vector feature of the first keyword.

[0068] Specifically, a word vector model can be used to extract the features of the first keyword to obtain the first word vector feature. The word vector model can be a pre-trained deep learning model, which has the ability to extract the word vector feature from the keyword. This embodiment does not make specific limitations on the training method of the above deep learning model, as long as it has the above feature extraction ability. In some embodiments, the word vector model can adopt the CLIP (Contrastive Language-Image Pre-Training) model.

[0069] Step S1022: Calculate the similarity between the first word vector feature and the second word vector feature of the second keyword in the preset keyword library to obtain the similarity between the first keyword and the second keyword. The second keywords in the preset keyword library are stored in the form of word vector features, and the second word vector features of the second keywords can also be obtained by extracting features from the second keywords using a word vector model.

[0070] In this embodiment, a conventional similarity calculation method can be used to calculate the similarity between the first and second word vector features, and the similarity between the two word vector features is used as the similarity between the corresponding two keywords. The greater the similarity, the more similar the two keywords are; the smaller the similarity, the less similar the two keywords are.

[0071] In some embodiments, the preset keyword library is a keyword library constructed based on a Knowledge Graph, that is, the preset keyword library stores the second keywords in the form of a Knowledge Graph, and the entities and the relationships between the entities in the Knowledge Graph are all second keywords. For example, the relationship between two entities is an action (such as leading, holding, etc.), and the second keyword will include both of these entities and the action at the same time.

[0072] Step S1023: Obtain the second keywords with a similarity greater than the set threshold as the second keywords with semantic similarity. If the similarity is greater than the set threshold, it indicates that the similarity degree between the first and second keywords is relatively high. Therefore, the second keyword is used as the keyword with semantic similarity.

[0073] The value of the set threshold represents the size of the similarity. When determining the value of the set threshold, those skilled in the art can count the similarities between a large number of semantically similar keywords, and then obtain the minimum similarity, and determine the value of the set threshold according to this minimum similarity.

[0074] Based on the method described in the above steps S1021 to S1023, the second keywords semantically similar to the first keyword can be conveniently and accurately obtained by using the similarity of the word vector features between the keywords.

[0075] II. Explain step S104.

[0076] In some embodiments of the above step S104, the following steps S1041 to S1043 can be used to obtain the facial information of the digital human in the first digital human image. Figure 4 As shown in the figure.

[0077] Step S1041: Identify the face area of the digital human in the first digital human image.

[0078] In this embodiment, a conventional face detection method (such as a deep learning model for face detection) can be used to perform face detection on the first digital human image to obtain the face region of the digital human. In some embodiments, the head region of the digital human can also be detected first on the first digital human image to obtain the head region of the digital human, and then the face region can be extracted from the head region. For example, information unrelated to the face, such as the background and hair in the head region, is removed to obtain the face region.

[0079] Step S1042: Extract the face key points and facial bone points of the face region.

[0080] The face key points include eyes, nose, mouth, chin, etc., and the facial bone points include glabella point, nasion point, orbital point, zygomatic point, nasal tip point, mental point, mandibular angle point, pterion point, etc.

[0081] In addition, the 3D coordinates of the face key points and facial bone points are extracted in this step. When obtaining the rendering information in the subsequent step S106, a rendering engine can be used to perform 3D rendering based on the 3D coordinates of the face key points and facial bone points to obtain the 3D model of the digital human, and then the mouth shape, expression, and head movement of the digital human can be generated using this 3D model.

[0082] Step S1043: Obtain the facial information according to the face key points and facial bone points.

[0083] Specifically, the face key points and facial bone points are combined as the facial information.

[0084] Based on the methods described in the above steps S1041 to S1043, the facial information of the digital human can be accurately obtained, which is beneficial to improving the accuracy of obtaining the rendering information according to the facial information.

[0085] In some embodiments of the above step S104, before obtaining the facial information of the digital human in the first digital human image, the image information of the first digital human image can also be optimized to obtain an optimized image, and then the facial information of the digital human in the optimized image is obtained. Among them, the image information of the first digital human image can include color, background, and the facial information of the digital human.

[0086] Optimizing the color and background can ensure the overall aesthetic appearance of the image. In this embodiment, conventional image processing methods can be used to optimize the color and background of the image to ensure the aesthetic appearance of the image. For example, in some embodiments, histogram equalization can be used to adjust the color of the first digital human image, and the gamma correction method can be used to correct the background color of the first digital human image.

[0087] Optimizing the facial information can make the facial information more natural and have a higher degree of realism. In this embodiment, super-resolution image reconstruction can be performed according to the facial information to obtain a reconstructed image, and the facial information can be repaired or adjusted according to the reconstructed image.

[0088] III. Description of step S106.

[0089] In some embodiments of the above step S106, it can be achieved through Figure 5 the following steps S1061 to S1062 as shown, to generate rendering information corresponding to each frame of voice information.

[0090] Step S1061: Generate the mouth shapes of the digital human corresponding to each frame of voice information respectively according to the pronunciation of each frame of voice information and based on the facial information.

[0091] The pronunciation of the voice information can be represented as a phoneme sequence, and this phoneme sequence includes at least one phoneme. When generating the mouth shape of the digital human, the mouth shapes matching each phoneme in the phoneme sequence can be obtained respectively, and the mouth shape corresponding to the voice information can be generated according to the mouth shapes matching each phoneme.

[0092] When obtaining the mouth shape matching the phoneme, the matching relationship between different phonemes and different mouth shapes can be preset in advance, and the phoneme can be matched according to this matching relationship to obtain the matching mouth shape.

[0093] In addition, synchronization processing needs to be performed on each frame of voice information and their corresponding mouth shapes to ensure that the mouth shapes corresponding to each frame of voice information are respectively synchronized with each frame of voice information.

[0094] Step S1062: Use a pre-trained deep learning model to generate the expressions and head movements of the digital human corresponding to each frame of voice information respectively according to the facial information and each frame of voice information. The deep learning model can perform semantic understanding and analysis on each frame of voice information to determine the types of expressions and head movements to be generated, and then generate the expressions and head movements of the digital human based on the facial information and the types of expressions and head movements to be generated. In this embodiment, the above deep learning model can be trained using conventional model training methods, and this embodiment does not specifically limit the training method of this model, as long as it has the above-mentioned ability to generate expressions and head movements. For example, in some embodiments, the above deep learning model can use a multimodal large language model, and the data processed by this multimodal large language model includes various types of data such as images, texts, and voices.

[0095] In some embodiments, after generating the expressions and head movements of the digital human, the expressions and head movements can be further optimized to make them more natural and realistic. In this embodiment, super-resolution image reconstruction can be performed based on the first digital human image according to the expressions and head movements of the digital human to obtain a reconstructed image, and the expressions and head movements of the digital human can be repaired or adjusted according to the reconstructed image.

[0096] Based on the method described in the above steps S1061 to S1062, natural and highly realistic expressions and head movements of the digital human can be generated, which is conducive to improving the natural smoothness and realism of the digital human video.

[0097] IV. Description of step S108.

[0098] In some embodiments of the above step S108, the video synthesis of each frame of voice information and the second digital human image corresponding to each frame of voice information can be performed through the following steps S1081 to S1082 to obtain a digital human video.

[0099] Step S1081: Generate subtitles according to the video description information and synchronize the subtitles with the mouth shapes of the digital human corresponding to each frame of voice information.

[0100] Step S1082: Perform video synthesis according to the synchronized subtitles, each frame of voice information and the second digital human image corresponding to each frame of voice information to obtain a digital human video.

[0101] The digital human video includes a plurality of video frames arranged in chronological order. Each video frame is synthesized by a frame of voice information, its corresponding second digital human image and subtitles. Among them, the time of the video frame is the time corresponding to the voice information in the video frame.

[0102] Based on the method described in the above steps S1081 to S1082, subtitles can be additionally added on the basis of images and audio to improve the information richness of the digital human video.

[0103] In some embodiments of the above step S108, the video synthesis of each frame of voice information and the second digital human image corresponding to each frame of voice information can be performed through the following steps S1083 to S1084 to obtain a digital human video.

[0104] Step S1083: In response to the user's configuration operation on the image elements on the interaction interface, obtain the image elements to be added and their configuration information. The configuration information includes the time and position where the image elements appear in the digital human video;

[0105] The interactive interface may provide an operation area for configuring image elements, and the user may set the configuration information of the image elements in the operation area. In this embodiment, an image element library may be provided, and the image element library stores a variety of different image elements, which may be taken by the user himself and stored in the image element library, or downloaded by the user from the Internet and stored in the image element library. The user may select an image element in the above operation area, and the image element selected by the user for configuration may be determined according to the selection operation.

[0106] See attached Figure 6 , Figure 6 An image element is shown as an example. The image in the dotted frame is the image element, and the rest of the image area is a frame of the digital human video. Figure 6 The position of the image element in this frame of image is shown by way of example, but the time when the image element appears in the digital human video is not shown.

[0107] Step S1084: Perform video synthesis on the image elements, each frame of voice information and the second digital human image corresponding to each frame of voice information according to the configuration information to obtain a digital human video.

[0108] Specifically, according to the time when the image element in the configuration information appears, the voice information corresponding to the time is obtained, and the voice information, the image element and the second digital human image corresponding to the voice information are subjected to video synthesis, wherein during the synthesis, the position of the image element in the second digital human image is determined according to the position of the image element in the configuration information.

[0109] In some implementations, the digital human video and image elements are freely combined through layered rendering to enhance the information richness of the digital human video. In addition, special effects (such as light and shadow changes, transition effects) can also be added to enhance the visual effect.

[0110] Based on the method described in steps S1083 to S1084 above, additional image elements can be added on the basis of images and audio to enhance the information richness of the digital human video.

[0111] In some implementations of the above step S108, the methods described in the above steps S1081 to S1084 may be simultaneously used to perform video synthesis to obtain a digital human video, thereby further improving the information richness of the digital human video.

[0112] The following is combined with Figure 7 , an embodiment of the digital human video synthesis method provided by the present application is described. Figure 7 , Figure 7 The overall process of the digital human video synthesis method according to some embodiments of the present application is exemplified. Figure 7As shown, video synthesis can be performed through the following steps S201 to S205.

[0113] Step S201: Generate an image based on keywords to obtain an AI image. Specifically, in response to the user's input operation of the image description text on the interaction interface, obtain the image description text input by the user. The image description text is used to describe the feature information of the image to be generated (i.e., the above-mentioned AI image); extract the first keyword from the image description text, and obtain the second keyword semantically similar to the first keyword from the preset keyword library; input the first keyword and the second keyword into the preset image generation model for image generation to obtain the AI image. Among them, the AI image can be a digital human image or an image that does not contain a digital human. Step S202: Determine whether the AI image is a portrait image; if it is a portrait image (i.e., a digital human image), then go to step S203; otherwise, go to step S205. Step S203: Image preprocessing. Specifically, use the method described in the aforementioned step S104 to obtain the facial information of the digital human in the AI image. Step S204: Image rendering. Specifically, use the methods described in the aforementioned steps S105 to S107 to obtain multiple frames of voice information and the corresponding second digital human images for each frame of voice information. Arrange the second digital human images corresponding to each frame of voice information according to the order of the time of the voice information from first to last to obtain the first video. Step S205: Video synthesis to obtain the second video. Specifically, use the method described in the aforementioned step S108 to perform video synthesis on the AI image, image elements, subtitles, audio, and the first video to obtain the second video, that is, the final digital human video. Among them, the AI image in step S205 is a non-portrait image. When performing video synthesis, steps such as obtaining facial information and obtaining rendering information do not need to be executed, and the multi-frame voice information converted from the video description information can be directly synthesized with the AI image for video synthesis.

[0114] Similar to the foregoing method embodiments, based on the method described in the above steps S201 to S205, it is also possible to simplify the operation steps of video synthesis, meet the user's personalized needs for video synthesis, and improve the natural smoothness and realism of the synthesized video.

[0115] It should be noted that although the above embodiments describe the various steps in a specific order, those skilled in the art can understand that in order to achieve the effects of the present application, it is not necessary to execute the different steps in such an order. They can be executed simultaneously (in parallel) or in other orders. These adjusted solutions are equivalent technical solutions to the technical solutions described in the present application, and therefore will also fall within the protection scope of the present application.

[0116] Those skilled in the art can understand that all or part of the processes in the methods of the above-mentioned embodiments of the present application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate forms, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium that can carry the computer program code, etc.

[0117] On the other hand, the present application also provides a computer-readable storage medium.

[0118] In an embodiment of a computer-readable storage medium according to the present application, the computer-readable storage medium can be configured to store a program for executing the digital human video synthesis method of the above-mentioned method embodiment. The program can be loaded and run by a processor to implement the above-mentioned digital human video synthesis method. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present application is a non-transitory computer-readable storage medium.

[0119] On the other hand, the present application also provides a server.

[0120] In an embodiment of a server according to the present application, the server can include at least one processor; and a memory communicatively connected to the at least one processor; wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the method described in any of the above embodiments is implemented. Refer to the attached Figure 8 , Figure 8 It is exemplarily shown in the figure that the memory 11 and the processor 12 are communicatively connected through a bus.

[0121] On the other hand, the present application also provides a digital human video synthesis system.

[0122] Refer to the attached Figure 9 , Figure 9 It is a schematic diagram of the main structure of a digital human video synthesis system according to an embodiment of the present application. As Figure 9 shown, the digital human video synthesis system can include a keyword image generation module 21, an image rendering module 22, and a video synthesis module 23.

[0123] The keyword-based image generation module 21 can be configured to: in response to a user's input operation on the image description text on the interaction interface, obtain the image description text input by the user, where the image description text is used to describe the feature information of the digital human image; extract the first keyword from the image description text, obtain the second keyword semantically similar to the first keyword from the preset keyword library, and input the first keyword and the second keyword into a preset image generation model for image generation to obtain the first digital human image.

[0124] The image rendering module 22 can be configured to: obtain the facial information of the digital human in the first digital human image; in response to a user's input operation on the video description information on the interaction interface, obtain the video description information input by the user, and convert the video description information into multi-frame voice information; generate rendering information corresponding to each frame of voice information respectively according to the facial information and the multi-frame voice information, where the rendering information includes the mouth shape, expression, and head movement of the digital human.

[0125] The video synthesis module 23 can be configured to: generate the second digital human image corresponding to each frame of voice information respectively according to the rendering information corresponding to each frame of voice information and the first digital human image; perform video synthesis on each frame of voice information and the second digital human image corresponding to each frame of voice information to obtain the digital human video.

[0126] For the description of the specific implementation functions of the above keyword-based image generation module 21, image rendering module 22, and video synthesis module 23, reference can be made to steps S101 to S108 in the foregoing method embodiment.

[0127] Refer to the appendix Figure 10 In some embodiments, the digital human video synthesis system may further include an image preprocessing module 24 and a pipeline module 25.

[0128] The image preprocessing module 24 can be configured to: identify the face area of the digital human in the first digital human image; extract the face key points and facial bone points of the face area; obtain the facial information according to the face key points and the facial bone points. For the description of the specific implementation function of the image preprocessing module 24, reference can be made to step S104 in the foregoing method embodiment. The pipeline module 25 can be configured to manage the digital human images, rendering information, and digital human videos generated during the digital human video synthesis process, facilitating subsequent viewing and problem troubleshooting by the user.

[0129] In some embodiments, there are multiple versions of the image generation model, and the digital human video synthesis system can manage different versions of the image generation model. When performing video synthesis, the digital human video synthesis system will use one version of the image generation model to generate the first digital human image, and the version information of this image generation model can be selected by the user.

[0130] The above digital human video synthesis system is used to execute Figure 1-7 the embodiments of the digital human video synthesis method shown. The technical principles, technical problems solved, and technical effects produced by both are similar. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process and related descriptions of the digital human video synthesis system can refer to the content described in the embodiments of the digital human video synthesis method, which will not be elaborated here.

[0131] So far, the technical solution of the present application has been described in conjunction with one embodiment shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Without departing from the principle of the present application, those skilled in the art can make equivalent changes or substitutions to relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present application.

Claims

1. A digital human video synthesis method, characterized in that, The method includes: In response to a user's input operation on the image description text in the interaction interface, obtaining the image description text input by the user, where the image description text is used to describe the feature information of the digital human image; Extracting a first keyword from the image description text, obtaining a second keyword semantically similar to the first keyword from a preset keyword library, and inputting the first keyword and the second keyword into a preset image generation model for image generation to obtain a first digital human image; Obtaining the facial information of the digital human in the first digital human image; In response to a user's input operation on the video description information in the interaction interface, obtaining the video description information input by the user, and converting the video description information into multi-frame voice information; Generating rendering information corresponding to each frame of voice information according to the facial information and the multi-frame voice information, where the rendering information includes the mouth shape, expression, and head movement of the digital human; Generating a second digital human image corresponding to each frame of voice information according to the rendering information corresponding to each frame of voice information and the first digital human image; Performing video synthesis on each frame of voice information and the second digital human image corresponding to each frame of voice information to obtain a digital human video.

2. The method according to claim 1, characterized in that, The obtaining of the second keyword semantically similar to the first keyword from the preset keyword library includes: Obtaining the first word vector feature of the first keyword, calculating the similarity between the first word vector feature and the second word vector feature of the second keyword in the preset keyword library, and obtaining the similarity between the first keyword and the second keyword; Obtaining the second keyword with the similarity greater than a set threshold as the semantically similar second keyword.

3. The method according to claim 1, characterized in that, The obtaining of the facial information of the digital human in the first digital human image includes: Identifying the face region of the digital human in the first digital human image; Extracting the face key points and facial bone points of the face region; Obtaining the facial information according to the face key points and the facial bone points.

4. The method according to claim 1 or 3, characterized in that, Before the obtaining of the facial information of the digital human in the first digital human image, the method further includes: Optimizing the image information of the first digital human image, where the image information includes color, background, and the facial information of the digital human.

5. The method according to claim 1, wherein The generating of the rendering information corresponding to each frame of voice information according to the facial information and the multi-frame voice information includes: Generating the mouth shape of the digital human corresponding to each frame of voice information according to the pronunciation of each frame of voice information and based on the facial information; Using a pre-trained deep learning model to generate the expression and head movement of the digital human corresponding to each frame of voice information according to the facial information and each frame of voice information.

6. The method according to claim 1, wherein The performing of video synthesis on each frame of voice information and the second digital human image corresponding to each frame of voice information includes: Generating subtitles according to the video description information, and synchronizing the subtitles with the mouth shape of the digital human corresponding to each frame of voice information; Video synthesis is performed on the synchronized subtitles, the frame voice information, and the second digital human image corresponding to the frame voice information to obtain a digital human video.

7. The method according to claim 1, characterized in that, The video synthesis of the frame voice information and the second digital human image corresponding to the frame voice information includes: In response to a user's configuration operation on an image element in an interactive interface, an image element to be added and its configuration information are obtained, where the configuration information includes the time and position at which the image element appears in the digital human video; Video synthesis is performed on the image element, the frame voice information, and the second digital human image corresponding to the frame voice information according to the configuration information to obtain a digital human video.

8. The method according to claim 1, wherein The preset keyword library is a keyword library constructed based on a knowledge graph; And / or, the video description information is audio information or text information.

9. A server, characterized in that, It includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the digital human video synthesis method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium storing multiple program codes, characterized in that, The program code is adapted to be loaded and run by a processor to execute the digital human video synthesis method according to any one of claims 1 to 8.