Backboard video generation method for real-time interactive digital human and related device

By acquiring multiple original images from different angles to generate digital human image images, and using video to generate large models to synthesize base video, the problems of high computing resource consumption and high cost in existing technologies are solved. This achieves efficient and low-cost digital human base video generation, ensuring appearance stability and realism, expanding the range of adaptation, and meeting the needs of large-scale customized production.

CN121619475APending Publication Date: 2026-03-06BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
CN202511777119.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-06

Smart Images

  • Figure CN121619475A_ABST
    Figure CN121619475A_ABST
Patent Text Reader

Abstract

The invention provides a backplane video generation method for real-time interactive digital humans and a related device, and relates to the technical field of image generation, in particular to the technical field of artificial intelligence such as human-computer interaction, digital humans, end-cloud integration and large models. The method comprises the following steps: acquiring a plurality of original images presented by the same target person at different angles; generating a digital human image taking a pure color as a background based on the plurality of original images; generating prompt information based on a preset action expression demand and the digital human image; and inputting the prompt information into a preset video generation large model, and generating a digital person bottom plate video which enables a digital person corresponding to the target person to show a target action corresponding to the action expression demand. According to the method, efficient and low-cost generation of the digital human bottom plate video is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of graphics generation technology, specifically to artificial intelligence technologies such as human-computer interaction, digital humans, edge-cloud integration, and large models, and particularly to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating background video for real-time interactive digital humans. Background Technology

[0002] Currently, digital human technology is rapidly developing and widely used in various application scenarios, especially in live streaming, e-commerce, and online customer service. With the increasing prevalence of digital humans, their technological background encompasses complex computing resource and bandwidth requirements.

[0003] Digital human systems, relying on advanced artificial intelligence and natural language processing technologies, can provide real-time interaction with users, enhancing the user experience. During large-scale deployment, digital humans can efficiently simulate the behavior and language of real people, meeting personalized needs in various scenarios. Summary of the Invention

[0004] This disclosure presents a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating baseboard videos of real-time interactive digital humans, achieving efficient and low-cost generation of digital human baseboard videos.

[0005] In a first aspect, embodiments of this disclosure propose a method for generating a base video for a real-time interactive digital human, comprising: acquiring multiple original images of the same target person presented from different angles; generating a digital human image image with a solid color background based on the multiple original images; generating prompt information based on preset action performance requirements and the digital human image image; inputting the prompt information into a preset video generation model to generate a digital human base video that enables the digital human corresponding to the target person to perform the target action corresponding to the action performance requirements.

[0006] Secondly, embodiments of this disclosure propose a device for generating a base video for a real-time interactive digital human, comprising: an original image acquisition unit configured to acquire multiple original images of the same target person presented from different angles; a digital human image acquisition unit configured to generate a digital human image with a solid color background based on the multiple original images; a prompting information generation unit configured to generate prompting information based on preset action performance requirements and the digital human image; and a base video generation unit configured to input the prompting information into a preset video generation model to generate a digital human base video that causes the digital human corresponding to the target person to perform the target action corresponding to the action performance requirements.

[0007] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the method for generating a baseboard video for a real-time interactive digital human as described in the first aspect.

[0008] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer, when executed, to implement the method for generating a baseboard video for a real-time interactive digital human as described in the first aspect.

[0009] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the steps of the method for generating a baseboard video for a real-time interactive digital human as described in the first aspect.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is an exemplary system architecture to which this disclosure can be applied; Figure 2 A flowchart illustrating a method for generating a base video for a real-time interactive digital human, provided as an embodiment of this disclosure; Figure 3 A flowchart illustrating a method for generating a digital human image provided in this embodiment of the disclosure; Figure 4 A flowchart illustrating a method for generating digital human baseboard video using a large control video generation model, as provided in this embodiment of the disclosure; Figure 5 is a flowchart illustrating a method for generating a base video for a real-time interactive digital human in an application scenario, as provided in an embodiment of this disclosure. Figure 6 A structural block diagram of a baseboard video generation device for real-time interactive digital humans provided in this disclosure embodiment; Figure 7 This is a schematic diagram of the structure of an electronic device suitable for performing a method for generating background video for a real-time interactive digital human, provided as an embodiment of the present disclosure. Detailed Implementation

[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0013] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0014] Figure 1 An exemplary system architecture 100 is shown that can be applied to embodiments of the background video generation method, apparatus, electronic device, and computer-readable storage medium for real-time interactive digital humans disclosed herein.

[0015] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0016] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include e-commerce applications, live streaming applications, and instant messaging applications.

[0017] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.

[0018] Server 105 can provide various services through its built-in applications. Taking an e-commerce application that can provide real-time interactive digital human base video generation services as an example, when running this e-commerce application, server 105 can achieve the following effects: First, acquire multiple original images of the same target person presented from different angles; then, based on the multiple original images, generate a digital human image with a solid color background; next, generate prompt information based on preset action performance requirements and the digital human image; finally, input the prompt information into a preset video generation model to generate a digital human base video that makes the digital human corresponding to the target person perform the target action corresponding to the action performance requirements.

[0019] It should be noted that, in addition to being obtained from terminal devices 101, 102, and 103 via network 104, multiple original images and motion representation requirements can also be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (e.g., when starting to process previously retained background video generation tasks for real-time interactive digital humans), it can choose to directly retrieve this data from locally. In this case, the exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.

[0020] Since generating digital human base video requires significant computing resources and power, the base video generation method for real-time interactive digital humans provided in the subsequent embodiments of this disclosure is generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the base video generation device for real-time interactive digital humans is also generally located in the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also perform the aforementioned calculations performed by the server 105 through e-commerce applications installed on them, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the e-commerce application determines that its terminal device has strong computing power and abundant remaining computing resources, it can allow the terminal device to perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the base video generation device for real-time interactive digital humans can also be located in terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.

[0021] It should be understood that Figure 1The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0022] Please refer to Figure 2 , Figure 2 A flowchart of a method for generating a base video for a real-time interactive digital human, provided in this disclosure embodiment, wherein process 200 includes the following steps: Step 201: Obtain multiple original images of the same target person from different angles; This step is intended to be performed by the implementing entity (e.g.) Figure 1 The server 105 shown acquires multiple original images of the same target person presented from different angles. Among them, images captured by the camera from multiple angles around the target person can be used as original images.

[0023] In this embodiment, the executing entity can acquire at least two original images of the target person with different expressions and / or wearing different clothing from different angles; wherein, the different angles include at least: front view, left view and right view.

[0024] Step 202: Generate a digital human image with a solid color background based on multiple original images; Building upon step 201, this step aims to generate a digital human image against a solid-color background from multiple original images, using the aforementioned executing entity. Specifically, the executing entity inputs multiple original images from different angles into a pre-defined algorithm model. The model infers the shape (e.g., body contours, facial features) and appearance attributes of the target person based on the correspondence and parallax information between the multiple original images. Then, it generates a digital human image based on the shape and appearance attributes of the target person and renders the digital human image against a solid-color background. The digital human image is a set of standardized, driverable image assets or data representations that can represent the digital person. It may be represented as a set of rendered human images with alpha channels from key perspectives, or it may be a lightweight, driverable model containing human appearance and shape parameters.

[0025] Step 203: Generate prompt information based on preset action performance requirements and digital human avatar image; In this embodiment, the executing entity generates prompt information based on preset action performance requirements and a digital human avatar image. Action performance requirements refer to the specific behavioral instructions or content themes that the user wants the digital human to perform in the video. Action performance requirements can be determined in the following ways: the executing entity can understand the text information input by the user to determine the action performance requirements; the executing entity can also provide the user with a preset action library, which the user can select from a drop-down menu or tags to determine the action performance requirements; the user can also upload a live video clip, from which the executing entity extracts the action sequence and posture changes as the action performance requirements.

[0026] In this embodiment, the prompt information may include identity locking information, motion parameters, and scene and style context. Identity locking information involves encoding the digital human avatar into a feature vector or embedding that the model can recognize, thereby fixing the appearance of the video's protagonist. Motion parameters resolve the required action into specific, controllable motion parameters, which may include the posture of various body parts, movement trajectory, facial expression coefficients (such as the degree of upward movement of the corners of the mouth, the degree of furrowing of the brows), lip-sync sequence, and the duration and rhythm of the action. Scene and style context can integrate auxiliary information such as the background environment where the action occurs, shot size (such as close-up, half-body), and lighting style to generate a more immersive video.

[0027] Step 204: Input the prompt information into the preset video generation large model to generate a digital human base video that makes the digital human corresponding to the target person perform the target action corresponding to the action performance requirements.

[0028] In this implementation, the executing entity inputs prompts into a preset video generation model to generate a digital human base video that enables the digital human corresponding to the target person to perform the target actions corresponding to the action performance requirements. The preset video generation model is a deep learning model trained on massive amounts of video data, possessing powerful spatiotemporal generation capabilities. It can be used to understand the temporal logic, physical laws, and relationship between the action and the person's posture. The digital human base video is a ready-to-use semi-finished video that can be used as basic material to generate a complete digital human video. Digital human base videos typically feature solid-color backgrounds (such as gray or green), complete actions (such as explanations, demonstrations, or interactions), and visual consistency. A clean background facilitates seamless background replacement or compositing with other visual elements. Complete actions demonstrate the entire process of the target person performing the specified action. Visual consistency means that the person's appearance, lighting, and style remain stable throughout the video, presenting a professional and unified visual effect.

[0029] In this embodiment, the execution subject inputs the prompt information into the preset video to generate a large model. The appearance of the person maintained in the digital human base video is determined by the identity locking information, ensuring that the facial details, hairstyle and clothing are strictly consistent with the digital human image, so as to achieve appearance stability. Then, the model is guided frame by frame to render the person's posture according to the motion parameters, ensuring that every movement of the digital human is consistent with the motion performance requirements, and the transition between movements is smooth and natural, so that the movements look realistic.

[0030] The method for generating a base video for real-time interactive digital humans provided in this embodiment first acquires multiple original images of the same target person presented from different angles. Then, based on the multiple original images, a digital human image with a solid color background is generated. Prompt information is generated based on preset action performance requirements and the digital human image. Finally, the prompt information is input into a preset video generation model to generate a digital human base video that makes the digital human corresponding to the target person perform the target actions corresponding to the action performance requirements. This embodiment utilizes multiple original images from different angles to construct the digital human image, ensuring the stability and realism of the person's appearance and avoiding distortion caused by a single image. Furthermore, by combining preset action performance requirements and prompt information, the digital human base video is automatically synthesized using a video generation model. The entire process does not require real people to appear on camera or green screen recording; users only need to provide image and text input, significantly reducing the production threshold and cost, and achieving efficient and low-cost generation of digital human base videos.

[0031] Based on the above embodiments, to ensure the quality and reliability of the digital human avatar and lay a solid foundation for the stability and realism of subsequent video generation, please refer to... Figure 3 , Figure 3 A flowchart of a method for generating a digital human image provided in this disclosure embodiment, wherein process 300 includes the following steps: Step 301: Perform quality verification on each original image; In this embodiment, the execution entity performs quality checks on each original image. The quality checks are performed using preset algorithm standards to examine each original image, identifying and excluding images that may seriously affect the quality of subsequent modeling. The quality checks include at least one of the following: sharpness, illumination intensity, background complexity, and loss of key facial information. The execution unit determines image sharpness by detecting whether the image is blurred due to inaccurate focus or object movement. Images with sharpness below a preset sharpness threshold cannot provide accurate human contours and texture details, resulting in blurry edges and loss of detail in the generated image. The execution unit determines lighting intensity by evaluating whether the image is too dark or overexposed. Lighting intensity within a preset lighting threshold range can accurately reproduce human colors and three-dimensionality. The execution unit determines background complexity by analyzing whether the background contains too much cluttered, high-frequency texture information (such as dense foliage or complex bookshelves). Background complexity exceeding a preset complexity threshold will interfere with the algorithm's accurate segmentation of the human subject. The execution unit determines the degree of loss of key facial information by detecting whether the facial area is occluded (such as hands, hair, or glasses), whether there is severe distortion, or whether key features are invisible due to excessive angle. Only when the degree of loss of key facial information is less than a preset loss threshold can a digital human with natural expressions and be driven be generated.

[0032] Step 302: For the first original image that passed the quality check, perform incremental fusion processing using the second original image that failed the quality check to obtain the fused image of the target person; In this embodiment, the executing entity performs incremental fusion processing on the first original image that passed quality verification and the second original image that failed quality verification to obtain the fused image of the target person. The first original image is a set of high-quality images that passed quality verification, constituting the core data for constructing the digital human image; the second original image is a set of suboptimal images that did not completely pass quality verification, and these are not completely discarded but downgraded to auxiliary data.

[0033] In this embodiment, the executing entity can utilize local details that may exist in the second original image but are missing or unclear in the first original image to perform incremental fusion processing on the first original image, thereby supplementing and enriching the final generated digital human image. Specifically, the executing entity can first generate a preliminary digital human model or image based on the first original image, then use an algorithm to analyze which regions in the second original image are missing or of low quality in the digital human model or image, and finally extract the qualified local information from the second original image, perform geometric correction and illumination normalization, and then fuse it with the digital human model or image to obtain a fused image of the target person with more complete details and richer information.

[0034] Step 303: Replace the background of the fused image to obtain a digital human image with a green screen as the background; In this embodiment, the executing entity replaces the background of the fused image to obtain a digital human image with a green screen as the background. Replacing the background with a green screen is a standard practice in the film and video production industry, which greatly facilitates subsequent video generation or compositing with other video materials.

[0035] Step 304: Obtain clothing reference images and / or clothing description text; In this embodiment, the executing entity can obtain clothing reference images and / or clothing description text input by the user according to their needs. The clothing reference image can be a photograph of a real model, a flat lay image of a product, or any image containing visual elements of the target garment, used to provide the most specific and detailed visual information, such as precise texture, pattern, color, and fit. The clothing description text is a description of the clothing entered by the user using natural language, such as "a dark blue business suit" or "a casual T-shirt with an abstract pattern."

[0036] Step 305: Using the visual diffusion model, adjust the clothing of the digital human image according to the clothing reference image and / or clothing description text to obtain the digital human image after clothing adjustment; In this embodiment, the executing entity utilizes a visual diffusion model to adjust the clothing of a digital human image based on a clothing reference image and / or clothing description text, resulting in a digital human image with adjusted clothing. The visual diffusion model is a type of generative artificial intelligence that excels at high-fidelity, high-consistency image editing based on understanding instructions. The visual diffusion model incorporates model weight parameters trained specifically for characteristic clothing styles.

[0037] In this embodiment, the visual diffusion model uses a digital human image as a base and reference images and / or clothing description text as control conditions to regenerate the character's clothing area. It is important to note that while altering the clothing, the visual diffusion model must strictly preserve the original digital human's facial features, hairstyle, body shape, and other identity information. Furthermore, the newly generated clothing must naturally adapt to the digital human's current body posture, conform to fabric dynamics, and not appear pasted on or distorted. In addition, the lighting and shadows of the new clothing need to seamlessly blend with the lighting environment of the original digital human image to maintain visual consistency.

[0038] Step 306: Generate prompt information based on the action performance requirements and the digital human image after clothing adjustments; In this implementation, the executing entity generates prompts based on the required action performance and the digital human avatar after costume adjustments. These prompts may include identity locking information, action parameters, and scene and style context. Identity locking information involves encoding the costumed digital human avatar into a feature vector or embedding that the model can recognize, thus fixing the appearance of the video's protagonist. Action parameters parse the required action performance into specific, controllable action parameters, which may include the posture and trajectory of various body parts, facial expression coefficients (such as the degree of upward movement of the corners of the mouth and the degree of furrowing of the brows), lip-sync sequence, and the duration and rhythm of the action. Scene and style context can integrate auxiliary information such as the background environment where the action occurs, shot size (such as close-up, half-body), and lighting style to generate a more immersive video.

[0039] This embodiment provides a method for generating a digital human image through steps 301-303, and a method for adjusting the clothing of the digital human image through steps 304-306. There is a causal and dependent relationship between the two methods. The methods provided in steps 301-303 can exist independently, while the methods provided in steps 304-306 depend on the methods disclosed in steps 301-303. Therefore, steps 301-303 can be applied independently in one embodiment, while steps 301-303 and 304-306 must be applied simultaneously in one embodiment. This embodiment exists only as a preferred embodiment that simultaneously includes two specific implementation methods.

[0040] The method for generating digital human avatars disclosed in this embodiment introduces a quality verification and incremental fusion mechanism. By mining effective information from suboptimal data to enrich details, it generates industry-standard digital human avatars against a green screen background. This significantly improves the robustness and output quality of the digital human avatar generation process, ensuring the stability and professionalism of the final video product's appearance. Furthermore, this embodiment allows users to flexibly and efficiently digitally modify the digital human's clothing based on the generated avatar, breaking through the limitations of the original avatar and greatly expanding the range of roles, scenes, and styles that digital humans can adapt to, meeting the needs of large-scale customized production.

[0041] Based on the above implementation, to enhance the stability, realism, and controllability of digital human videos, the executing entity can generate a corresponding 3D model of the digital human based on multiple original images and a 2D digital human image. Specifically, the executing entity uses neural radiation field technology or 3D Gaussian splashing technology to generate a 3D model of the digital human from multiple original images and a 2D digital human image. Neural radiation field technology uses a large neural network to learn a continuous 3D space, mapping each point in the space to two attributes: volume density (indicating whether the point has an object surface) and color value. By emitting light rays from a specific viewpoint and accumulating the color and density of all points along these light paths, a 2D image from that viewpoint can be synthesized. During training, the neural network learns to make the synthesized image as consistent as possible with the input multiple original images and the 2D digital human image. After training, the neural network itself becomes the 3D model of the digital human. 3D Gaussian splashing technology is a newer, explicit, and efficient 3D scene representation and rendering technique. It uses millions of tiny, learnable 3D Gaussian spheres to explicitly represent a scene, each possessing attributes such as position, size, color, and opacity. During rendering, these Gaussian spheres are projected onto a 2D image plane, overlapping like splashes to form the final image. The training process optimizes the properties of the Gaussian spheres to match the rendered image with the input image. Its core advantages lie in its extremely fast rendering speed and high visual quality, enabling real-time generation of high-fidelity novel views, making it ideal for applications with high efficiency requirements. Digital human 3D avatar models model people from a 3D perspective. No matter how drastic the head-turning or turning movements are required by the subsequent video generation model, the digital human's appearance, hairstyle, and clothing maintain geometric consistency, completely avoiding problems such as facial distortion and hair deformation that may occur with 2D-driven rendering. Furthermore, the 3D model of the digital human can interact physically correctly with virtual light sources, producing realistic shadows and highlights. As the digital human moves in the video, the changes in light and shadow on its body will naturally change accordingly, greatly enhancing the realism. Generating a 3D model of the digital human means that it can be manipulated more precisely, such as changing hair color, adjusting body shape, and even creating 3D animations, laying the foundation for subsequent technological iterations. This embodiment, by constructing a 3D model of the digital human, fundamentally solves the limitations of 2D driving technology when the perspective changes, providing the most solid and reliable data foundation for generating digital human base videos with natural expressions, realistic movements, and stable appearance.

[0042] Based on the above implementation method, in order to improve the naturalness, consistency, and efficiency of the generated video's movements, please refer to... Figure 4 , Figure 4A flowchart of a method for generating digital human base video from a large model of controlled video generation, provided in this embodiment of the disclosure, is included in process 400, which includes the following steps: Step 401: Input the prompt information into the preset video to generate a large model; In this embodiment, the executing entity will input the prompt information into a preset large video generation model.

[0043] Step 401 and as follows Figure 2 The content of step 204 shown is the same as that shown. For the same part, please refer to the corresponding part of the previous embodiment. It will not be repeated here.

[0044] Step 402: The large model for controlling video generation uses retrieval enhancement technology to search for target action templates that match the action performance requirements in a preset action template library; In this embodiment, the execution entity controls the large-scale video generation model to retrieve target action templates matching the action performance requirements from a pre-set action template library using retrieval enhancement technology. These action performance requirements include: any action in a silent state or any action in a conversational state. Any action in a silent state is body language or facial expression without voice accompaniment (such as walking or smiling), while any action in a conversational state involves facial expressions and gestures synchronized with voice (such as lip movements and gesture emphasis). The retrieval may also need to consider the rhythm and timing of the actions to match the prosody of the speech. An action template is a continuous image sequence representing the corresponding action, including body movements (such as arm swings or body rotations) and / or facial expression movements (such as eyebrow raising or mouth corner raising). The action template library is typically constructed from labeled and normalized motion capture data or high-quality video clips to ensure the smoothness and reusability of the action templates.

[0045] In this embodiment, the executing entity can first control the video generation model to retrieve candidate action templates matching the target action corresponding to the action performance requirements from a preset action template library using retrieval enhancement technology; then, control the video generation model to select the target action template that matches the image features in the digital human avatar image from multiple candidate action templates; wherein, image features include: limb image features and facial image features, and facial image features include mouth image features. Retrieval enhancement technology is a technique that combines information retrieval and generative models, aiming to enhance the performance of generation tasks by introducing external information resources. It is commonly used in natural language processing tasks, especially in generative tasks (such as text generation, question answering, dialogue generation, etc.), and can improve the accuracy and information richness of the retrieved action template library.

[0046] Step 403: Control the video generation of the large model based on the digital human image and target action template to output the generated digital human base video.

[0047] In this embodiment, the execution entity controls the video generation model to output a digital human base video based on the digital human image and the target action template. Specifically, the video generation model merges the retrieved target action template (as an action sequence blueprint) with the digital human image (as the source of identity and appearance) to generate the digital human base video. The video generation model uses feature alignment and spatiotemporal transformation techniques to transfer the action sequence from the target action template to the digital human image, ensuring that the digital human maintains a stable appearance (e.g., consistent clothing and hairstyle) and natural transitions during actions. The final output digital human base video is a professional video asset with a solid color background and coherent actions, which can be directly used for subsequent compositing or distribution.

[0048] This embodiment discloses a method for generating digital human base videos from large-scale control video models. By introducing retrieval enhancement technology and a preset motion template library, it efficiently utilizes predefined motion resources, avoiding complex 3D modeling or live recording. This not only reduces computational overhead and production barriers but also ensures the realism and diversity of the generated video's movements, significantly improving the naturalness, consistency, and efficiency of the generated video's movements. Users only need to input simple text and images to quickly obtain high-quality digital human base videos adapted to different roles and scenarios, strongly supporting the implementation needs of large-scale, customized digital content production.

[0049] Based on the above implementation, to enhance interactive functionality, the executing entity can control a large-scale video generation model to generate digital human base videos that interact with various interactive props according to various prop interaction actions, based on a digital human base video, a preset interactive prop library, and a preset prop interaction action template library. The interactive prop library is a collection of 3D models or high-quality image sequences containing various common items. The prop interaction action template library consists of pre-recorded action sequences that depict physically reasonable interactions between the character and these props. The large-scale video generation model retrieves specified prop action templates and interactive action templates from the interactive prop library and the preset prop interaction action template library, respectively. Then, it redirects the digital human's movements in time and space, ensuring that its skeletal motion strictly matches the interactive action templates. Simultaneously, the model needs to render the props in real-time, ensuring precise contact and occlusion relationships between the props and the digital human's hands or body parts (e.g., when fingers hold a pen, the pen barrel should be partially obscured by the fingers). This embodiment's increased interactivity greatly expands the application scenarios of digital humans, transforming them from simple narrators into actual participants capable of demonstration and operation.

[0050] Based on the above implementation, in some optional embodiments of this example, at least one of the following post-processing operations is performed on the digital human base video: color normalization, temporal stabilization filtering, compression, and artifact repair. Color normalization is performed to ensure the consistency of video colors. Since there may be slight inter-frame color fluctuations during the generation process, or the video may need to be composited with different platforms and backgrounds, color normalization adjusts the overall hue, contrast, and saturation of the video to conform to a unified color standard, avoiding flickering or inconsistencies with the background. Temporal stabilization filtering is performed to eliminate unnatural high-frequency jitter or subtle shaking in the video. Although the generation model strives for stability, frame-by-frame generation may still introduce minute, imperceptible positional jumps. Sequential filtering analyzes the global motion of characters and props between consecutive frames and smooths it, resulting in an extremely stable and visually smooth video, especially suitable for explanatory scenes requiring prolonged viewing. Compression is a necessary optimization for network transmission and storage. While ensuring acceptable visual quality, it uses efficient video coding standards to reduce the size of video files, making them easier to distribute, play, and store, meeting the actual needs of internet applications. Artifact removal is the last line of defense in quality control, specifically designed to detect and repair imperfections that may occur during the generation process. These artifacts may include unnatural textures that appear momentarily on faces or clothing, flickering edges, or temporary distortions in localized areas. Through dedicated image processing algorithms or lightweight repair models, these defects can be automatically identified and smoothed, ensuring the purity and professionalism of the video content. This embodiment systematically processes the generated digital human prototype video, transforming the uncertain output sometimes brought about by cutting-edge generative technologies into a stable, reliable, and directly usable industrial product.

[0051] To enhance understanding, this disclosure also provides a specific implementation scheme based on a particular application scenario. Please refer to the example below. Figure 5-1 and Figure 5-2 , Figure 5-1 This is one example of the implementation method. Figure 5-2 This is a specific implementation process for this method.

[0052] This implementation proposes a novel, efficient, and low-cost digital human video creation technology. The technology revolves around the core path of "generating a high-quality background video from multiple images, and then generating a digital human from the background video." It systematically optimizes the entire process of digital human modeling and video generation, achieving high fidelity, high flexibility, and low-cost large-scale production of digital human content without relying on large-scale training resources or dedicated actors.

[0053] The specific implementation steps of this method are as follows: 1. Material Acquisition and Image Enhancement Processing. Users upload images from multiple angles, including front-facing photos, side-view photos, and outfit reference images. The system automatically verifies image clarity, lighting, background complexity, and other indicators to ensure that the images meet subsequent processing standards. For materials with poor aesthetics or cluttered backgrounds, the system uses image fusion and style enhancement technologies to generate high-quality portrait images with green screen backgrounds, fundamentally lowering the barrier to entry for users and addressing the pain point of "appearance anxiety and reluctance to appear on camera."

[0054] 2. Character Image Optimization and Clothing Fusion Modeling. The system uses a visual diffusion model (such as FLUX Kontext) combined with user-provided clothing reference images and text descriptions to automatically generate clothing styles that meet the user's needs. These styles are then fused with the character image, and professional LORA (Low-Rank Adaptation) weights are applied to achieve specific style effects. Through a processing pipeline including artistic enhancement and lighting optimization, high-quality character images are output as input for subsequent video generation.

[0055] 3. Green Screen Background Video Generation. Based on the optimized green screen subject image, the system calls upon a large-scale video generation model and combines it with a built-in large-scale expression and action template library. Using retrieval-enhanced generation technology, it automatically generates a green screen background video that meets the user's requirements. This video supports various subject states, including silent and conversational states, ensuring natural and realistic body movements and emotional expressions. Furthermore, a unified post-processing module performs color normalization, temporal stabilization filtering, and compression artifact repair to ensure consistency and stability of the final product.

[0056] like Figure 5-1 As shown, in this implementation, the user uploaded a total of 5 images from multiple angles, including a front view, a side view, and an outfit reference image. The system then used image enhancement processing, character image optimization, and clothing fusion modeling to finally generate the background green screen task image on the right side of the arrow.

[0057] This proposed digital human video creation technology utilizes multiple original images from different angles to construct a digital human image, ensuring the stability and realism of the human's appearance and avoiding distortion caused by a single image. Furthermore, by combining preset action performance requirements and prompts, it automatically synthesizes a digital human base video using a large video generation model. The entire process requires no live-action appearance or green screen recording; users only need to provide image and text input, significantly reducing the production threshold and cost. This enables efficient and low-cost generation of digital human base videos. Based on this, users can flexibly and efficiently digitally modify the digital human's clothing, breaking through the limitations of the original digital human image and greatly expanding the range of roles, scenes, and styles that digital humans can adapt to, meeting the needs of large-scale customized production. Moreover, by introducing quality verification and incremental fusion mechanisms, it enriches details by mining effective information from suboptimal data, generating industry-standard digital human image images with a green screen background. This significantly improves the robustness and output quality of the digital human image generation process, ensuring the stability and professionalism of the final video product's appearance.

[0058] Further reference Figure 6 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a device for generating background video for real-time interactive digital humans, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0059] like Figure 6 As shown, the real-time interactive digital human base video generation device 600 of this embodiment may include: an original image acquisition unit 601, a digital human image acquisition unit 602, a prompt information generation unit 603, and a base video generation unit 604. The original image acquisition unit 601 is configured to acquire multiple original images of the same target person presented from different angles; the digital human image acquisition unit 602 is configured to generate a digital human image with a solid color background based on the multiple original images; the prompt information generation unit 603 is configured to generate prompt information based on preset action performance requirements and the digital human image; and the base video generation unit 604 is configured to input the prompt information into a preset video generation model to generate a digital human base video that causes the digital human corresponding to the target person to perform the target action corresponding to the action performance requirements.

[0060] In this embodiment, the specific processing and technical effects of the base video generation device 600 for real-time interactive digital humans—including the original image acquisition unit 601, the digital human image acquisition unit 602, the prompt information generation unit 603, and the base video generation unit 604—can be found in the following references. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.

[0061] In some optional implementations of this embodiment, the original image acquisition unit 601 is further configured to acquire at least two original images of the target person with different expressions and / or wearing different clothing from different angles; wherein, the different angles include at least: a frontal view, a left-side view, and a right-side view.

[0062] In some optional implementations of this embodiment, the digital human image acquisition unit 602 includes: a quality verification module configured to perform quality verification on each original image; wherein, the quality verification includes at least one of: sharpness, illumination intensity, background complexity, and loss of facial key information; an incremental fusion module configured to perform incremental fusion processing on the first original image that passed the quality verification and a second original image that failed the quality verification to obtain a fused image of the target person; and a background replacement module configured to replace the background of the fused image to obtain a digital human image with a green screen as the background.

[0063] In some optional implementations of this embodiment, the digital human image acquisition unit 602 further includes: a clothing acquisition module, configured to acquire a clothing reference image and / or clothing description text; and an emphasis adjustment module, configured to use a visual diffusion model to adjust the clothing of the digital human image according to the clothing reference image and / or clothing description text, to obtain a digital human image with adjusted clothing; wherein, the visual diffusion model is loaded with model weight parameters trained for characteristic clothing styles; correspondingly, the prompt information generation unit 603 includes: a prompt information generation module, configured to generate prompt information based on action performance requirements and the digital human image with adjusted clothing.

[0064] In some optional implementations of this embodiment, the digital human image acquisition unit 602 further includes a three-dimensional model generation module, which is configured to generate a three-dimensional image model of the corresponding digital human based on multiple original images and two-dimensional digital human image images.

[0065] In some optional implementations of this embodiment, the three-dimensional model generation module in the digital human image acquisition unit 602 is further configured to generate a three-dimensional image model of the corresponding digital human from multiple original images and two-dimensional digital human image images through neural radiation field technology or three-dimensional Gaussian splashing technology.

[0066] In some optional implementations of this embodiment, the baseboard video generation unit 604 includes: a prompt information input module configured to input prompt information into a preset video generation model; an action template retrieval module configured to control the video generation model to retrieve a target action template matching the action performance requirements from a preset action template library using retrieval enhancement technology; wherein, the action performance requirements include: any action in a silent state or any action in a dialogue state, the action template is a continuous image sequence representing the corresponding action, and the action includes limb movements and / or facial expression movements; the baseboard video generation module is configured to control the video generation model to output the generated digital human baseboard video based on the digital human image and the target action template.

[0067] In some optional implementations of this embodiment, the motion template retrieval module in the base plate video generation unit 604 further includes: a candidate module retrieval module, configured to control the large video generation model to retrieve candidate motion templates that match the target motion corresponding to the motion performance requirements from a preset motion template library through retrieval enhancement technology; and a motion template matching module, configured to control the large video generation model to select a target motion template that matches the image features in the digital human image image from multiple candidate motion templates; wherein, the image features include: limb image features and facial image features, and the facial image features include mouth image features.

[0068] In some optional implementations of this embodiment, the base plate video generation unit 604 further includes: a prop interaction module, configured to control the video generation model to generate digital human base plate videos that interact with various interactive props according to various prop interaction actions based on the digital human base plate video, a preset interactive prop library and a preset prop interaction action template library.

[0069] In some optional implementations of this embodiment, the base video generation device 600 for real-time interactive digital human further includes: a post-processing unit 605, configured to perform at least one of the following post-processing operations on the digital human base video: color normalization, temporal stabilization filtering, compression, and artifact repair.

[0070] This embodiment exists as a device embodiment corresponding to the above method embodiment. The device for generating a base video for real-time interactive digital humans provided in this embodiment constructs a digital human image using multiple original images from different angles. This ensures the stability and realism of the human's appearance, avoiding distortion caused by a single image. Furthermore, combined with preset action performance requirements and prompts, it automatically synthesizes the digital human base video using a large video generation model. The entire process does not require real people to appear on camera or green screen recording; users only need to provide image and text input, significantly reducing the production threshold and cost. This achieves efficient and low-cost generation of digital human base videos. On this basis, it allows users to flexibly and efficiently digitally modify the digital human's clothing, breaking through the limitations of the original digital human image and greatly expanding the range of roles, scenes, and styles that digital humans can adapt to, meeting the needs of large-scale customized production. Moreover, by introducing a quality verification and incremental fusion mechanism, it enriches the details by mining effective information from suboptimal data, generating an industry-standard digital human image with a green screen background. This significantly improves the robustness and output quality of the digital human image generation process, ensuring the appearance stability and professionalism of the final video product.

[0071] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the method for generating a baseboard video for a real-time interactive digital human as described in any of the above embodiments.

[0072] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that, when executed by a computer, enable the method for generating a baseboard video for a real-time interactive digital human as described in any of the above embodiments.

[0073] According to embodiments of this disclosure, this disclosure also provides a computer program product that, when executed by a processor, can implement the method for generating background video for real-time interactive digital humans as described in any of the above embodiments.

[0074] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0075] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded into random access memory (RAM) 703 from storage unit 708. The RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0076] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0077] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as a method for generating background video for a real-time interactive digital human. For example, in some embodiments, the method for generating background video for a real-time interactive digital human can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the method for generating background video for a real-time interactive digital human described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured by any other suitable means (e.g., by means of firmware) to perform a method for generating baseboard video for a real-time interactive digital human.

[0078] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0079] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0080] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0081] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0082] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0083] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0084] According to the technical solution of this disclosure, a digital human image is constructed using multiple original images from different angles. This ensures the stability and realism of the human's appearance, avoiding distortion caused by a single image. Furthermore, by combining preset action performance requirements and prompts, a digital human base video is automatically synthesized using a large video generation model. The entire process requires no live-action appearance or green screen recording; users only need to provide image and text input, significantly reducing the production threshold and cost. This achieves efficient and low-cost generation of digital human base videos. Based on this, users can flexibly and efficiently digitally modify the digital human's clothing, breaking through the limitations of the original digital human image and greatly expanding the range of roles, scenes, and styles that digital humans can adapt to, meeting the needs of large-scale customized production. Moreover, by introducing quality verification and incremental fusion mechanisms, and by mining effective information from suboptimal data to enrich details, an industry-standard digital human image with a green screen background is generated. This significantly improves the robustness and output quality of the digital human image generation process, ensuring the stability and professionalism of the final video product's appearance.

[0085] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0086] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating a base video of a real-time interactive digital person, comprising: obtaining multiple original images of the same target person presented from different angles; generating a digital person image with a solid color background based on the multiple original images; generating prompt information based on a preset action performance requirement and the digital person image; inputting the prompt information into a preset video generation large model to generate a digital person base video that makes the digital person corresponding to the target person perform a target action corresponding to the action performance requirement.

2. The method of claim 1, wherein, The obtaining multiple original images of the same target person presented from different angles comprises: obtaining at least two original images of the target person presented from different angles with different expressions and / or wearing different costumes; wherein the different angles at least include: a front view, a left side view and a right side view.

3. The method of claim 1, wherein, The generating a digital person image with a solid color background based on the multiple original images comprises: performing quality checking on each of the original images; wherein the quality checking includes at least one of: sharpness, illumination intensity, background complexity, and face key information loss degree; performing incremental fusion processing on a first original image that passes the quality checking using a second original image that fails the quality checking to obtain a fused image of the target person; performing background replacement on the fused image to obtain a digital person image with a green screen background.

4. The method of claim 3, further comprising: obtaining a dressing reference image and / or a dressing description text; using a visual diffusion model to dress the digital person image according to the dressing reference image and / or the dressing description text to obtain a digitally dressed digital person image; wherein the visual diffusion model loads model weight parameters trained for a specific dressing style; Correspondingly, the generating prompt information based on a preset action performance requirement and the digital person image comprises: generating the prompt information based on the action performance requirement and the digitally dressed digital person image.

5. The method of claim 3 or 4, further comprising: generating a three-dimensional image model of the corresponding digital person based on the multiple original images and a two-dimensional digital person image.

6. The method of claim 5, wherein, The generating a three-dimensional image model of the corresponding digital person based on the multiple original images and a two-dimensional digital person image comprises: generating a three-dimensional image model of the corresponding digital person based on the multiple original images and a two-dimensional digital person image using a neural radiance field technology or a three-dimensional Gaussian spatter technology.

7. The method of claim 1, wherein, The inputting the prompt information into a preset video generation large model to generate a digital person base video that makes the digital person corresponding to the target person perform a target action corresponding to the action performance requirement comprises: inputting the prompt information into a preset video generation large model; The control of the video generation large model retrieves a target action template matching the action performance requirement in a preset action template library through a retrieval enhancement technology; wherein, the action performance requirement includes: any action in a silent state or any action in a dialogue state, the action template is a continuous image sequence showing the corresponding action, and the action includes a body action and / or a facial expression action; The control of the video generation large model outputs the generated digital human bottom plate video based on the digital human image and the target action template.

8. The method of claim 7, wherein, The control of the video generation large model retrieves a target action template matching the action performance requirement in a preset action template library through a retrieval enhancement technology, including: The control of the video generation large model retrieves a target action template matching the action performance requirement in a preset action template library through a retrieval enhancement technology; The control of the video generation large model selects a target action template matching the image feature in the digital human image from the multiple candidate action templates; wherein, the image feature includes: a body image feature and a facial image feature, and the facial image feature includes a mouth image feature.

9. The method of claim 7 or 8, further comprising: The control of the video generation large model generates a digital human bottom plate video interacting with various interactive props in various prop interaction actions based on the digital human bottom plate video, a preset interactive prop library and a preset prop interaction action template library.

10. The method of claim 1, further comprising: Performing at least one post-processing operation on the digital human bottom plate video: color normalization processing, timing stabilization filtering processing, compression, artifact repair.

11. An apparatus for generating a bottom plate video of a real-time interactive digital human, comprising: An original image acquisition unit configured to acquire multiple original images of the same target person presented at different angles; A digital human image acquisition unit configured to generate a digital human image with a solid color background based on the multiple original images; A prompt information generation unit configured to generate prompt information based on a preset action performance requirement and the digital human image; A bottom plate video generation unit configured to input the prompt information into a preset video generation large model to generate a digital human bottom plate video making a digital human corresponding to the target person perform a target action corresponding to the action performance requirement.

12. An electronic device, comprising: At least one processor; And A memory connected in communication with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the bottom plate video generation method for a real-time interactive digital human of any one of claims 1-10.

13. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the bottom plate video generation method for a real-time interactive digital human of any one of claims 1-10.

14. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method for real-time interactive digital man rig video generation according to any one of claims 1-10.

Citation Information

Patent Citations

  • Image enhancement method and device

    CN112712470A

  • Virtual image generation method and device, virtual image model training method and device and electronic equipment

    CN116385643A

  • Digital human video generation method and device, electronic equipment and storage medium

    CN116528017A

  • Three-dimensional digital human generation and interaction method and system

    CN117496072A

  • Server, terminal, display device and digital human interaction method

    CN117812279A

Cited By

  • AIGC character generation consistency control method based on mud kneading entity constraint

    CN121837470A

  • Aigc character generation consistency control method based on clay entity constraint

    CN121837470B