Video generation method and device, electronic equipment and computer readable storage medium
By generating complete background images and extended background images, the background and subject are decoupled. Combined with multi-stage modeling and caching mechanisms, the problem of inconsistent backgrounds in long video generation is solved, achieving stability of the background area and continuity of the subject's actions, and reducing training and inference costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies suffer from severe background inconsistency when generating long videos, making it difficult to maintain the structural and color continuity of the background area. Furthermore, training and inference costs are high, and changes in the subject's movements can easily cause dynamic disturbances in the background area.
By acquiring reference images, a complete background image, an expanded background image, and a target object image are generated, decoupling the background from the subject. Multi-stage background modeling and a caching pool mechanism are used to ensure background consistency and maintain visual stability during the stitching process.
Significantly reduces background texture and tone drift across segments, achieving stable rendering of background areas under changes in subject movement, reducing training and inference costs, and generating high-quality, continuous, and complete video content.
Smart Images

Figure CN121865059A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a video generation method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] In live streaming applications, virtual elements (such as AI digital humans and virtual cartoons) are widely used to enhance the richness and interactivity of live streaming content, which involves a large number of long video animation generation needs. Currently, mainstream video generation models, such as the Dit (Diffusion Transformer) series of models, typically use a segmented processing approach when generating long videos, breaking the long video into multiple segments, generating them sequentially, and then splicing them together.
[0003] Currently, methods such as sliding window cross-generation are commonly used to improve the consistency between segments. However, these methods rely solely on information sharing in locally overlapping areas, making it difficult to achieve globally stable modeling of the background region, and they also impose significant pressure on training and inference costs. Furthermore, the background and subject generation processes are tightly coupled; changes in subject movement easily trigger dynamic perturbations in the background region, further exacerbating its inconsistency. Therefore, there is an urgent need for a technical solution that can effectively address the background region inconsistency problem in long video generation while maintaining low training and inference costs, and mitigate the impact of color difference, to meet the high-quality video generation requirements of real-world live streaming scenarios. Summary of the Invention
[0004] In view of this, the object of the present invention is to provide a video generation method, apparatus, electronic device and computer-readable storage medium capable of generating long videos with consistent backgrounds.
[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, the present invention provides a video generation method, the method comprising: Obtain the reference image and text prompts for each video segment to be generated; Based on the reference image, a complete background image, an extended background image, and a target object image are obtained; Generate corresponding video clips based on the reference image, the complete background image, the extended background image, the target object image, and the text prompts; The video segments are spliced together to obtain a complete video.
[0006] In an optional implementation, obtaining the complete background image, the extended background image, and the target object image based on the reference image includes: The reference images are segmented based on the target object to obtain the target object image and the incomplete background image; the background regions contained in each reference image remain unchanged during the video generation process; The incomplete background image is filled in to obtain the complete background image; The complete background image is expanded to obtain the expanded background image.
[0007] In an optional implementation, generating a corresponding video segment based on the reference image, the complete background image, the extended background image, the target object image, and the text prompt includes: Latent features are generated based on the complete background image, the extended background image, and the reference image; the latent features are used to characterize the structure, color, and texture of the image; Target object features are generated based on the target object image, and complete background features are generated based on the complete background image; the complete background features are used to characterize the static visual appearance of the background. A prompt feature is generated based on the extended background image and the text prompt words; the prompt feature is used to characterize the fusion information of prompt semantics and visual style; Generate corresponding video clips based on the potential features, the cue features, the target object features, and the complete background features.
[0008] In an optional implementation, generating prompt features based on the extended background image and the text prompt words includes: Generate an expanded background prompt based on the expanded background image; Generate an extended background feature based on the extended background cue words, and generate text features based on the text cue words; The prompt feature is generated based on the background features and the text features.
[0009] In an optional implementation, after generating the prompt features based on the extended background image and the text prompt words, the method further includes: Generate background key-value pairs based on the prompt features and the features extracted from the complete background image; Generate an expanded background prompt based on the expanded background image; If the similarity between the token vector corresponding to the background key-value pair and the token vector corresponding to the extended background prompt word is higher than a first preset threshold, the background key-value pair is saved to the background cache pool.
[0010] In an optional implementation, generating complete background features based on the complete background image includes: Retrieve the previous background key-value pair from the background cache pool; The features extracted based on the complete background image are fused with the previous background key-value pair to obtain the complete background features.
[0011] In an optional implementation, the method further includes: Generate action prompts based on the target object graph, and generate action key-value pairs based on the action prompts; If the similarity between the token vector corresponding to the action key-value pair and the token vector corresponding to the text prompt word is higher than a second preset threshold, the action key-value pair is saved to the action cache pool; the action cache pool is used to retrieve the previous action key-value pair from the action cache pool and fuse it with the features extracted based on the target object graph when generating target object features.
[0012] In a second aspect, the present invention provides a video generation apparatus, the apparatus comprising: The acquisition module is used to acquire the reference image and text prompts corresponding to each video segment to be generated; The generation module is used to obtain a complete background image, an extended background image, and a target object image based on the reference image; and to generate corresponding video clips based on the reference image, the complete background image, the extended background image, the target object image, and the text prompt words. The combination module is used to splice the various video segments to obtain a complete video.
[0013] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing a computer program executable by the processor, the processor executing the computer program to implement the video generation method described in any of the foregoing embodiments.
[0014] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video generation method as described in any of the foregoing embodiments.
[0015] Compared to existing technologies, the video generation method, apparatus, electronic device, and computer-readable storage medium provided in this invention acquire reference images and text prompts, and generate a complete background image, an extended background image, and a target object image based on the reference images. This makes the generation of video segments dependent on a unified and extended background reference, thereby maintaining the structural and color continuity of the background area during segmented generation. Furthermore, by combining multi-stage background modeling with the complete and extended background images, the generated video segments maintain background consistency while also realizing dynamic changes of the target object under rich semantic guidance. The resulting continuous and complete video content, after splicing, maintains the static and stable characteristics of the background, significantly reduces cross-segment background texture and tone drift, and achieves stable presentation of the background area under changes in the subject's actions.
[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A schematic flowchart of a video generation method provided by an embodiment of the present invention is shown.
[0019] Figure 2 A block diagram of a video generation model provided in an embodiment of the present invention is shown.
[0020] Figure 3 A block diagram of a video generation apparatus provided in an embodiment of the present invention is shown.
[0021] Figure 4 A block diagram of an electronic device provided in an embodiment of the present invention is shown.
[0022] Icons: 400 - Video generation device; 401 - Acquisition module; 402 - Generation module; 403 - Combination module; 500 - Electronic device; 510 - Memory; 520 - Processor; 530 - Communication module. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0024] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0025] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0026] In existing technologies, text-driven or image-based video generation can produce high-quality videos within a short timeframe. However, when generating long videos, a common strategy is to split the long video into multiple segments, generate them in segments, and then stitch them together. This leads to the following problems: Poor background consistency: The randomness of different segments, slight camera movement, exposure shift, and style drift can cause shifts in background texture, lighting, and tone, resulting in "jumps" after stitching.
[0027] Subject-background coupling: Most models unify the subject and background in the same diffusion process. The conflict between the subject's actions and the static nature of the background is difficult to maintain across segments, causing the part of the background that is occluded by the subject to be regenerated in each action, resulting in distortion of the background changes in long videos.
[0028] Insufficient background condition injection: Most of the existing image conditions are reference images with obfuscated backgrounds or the first frame. They are injected based on Cross Attention (CA) or ControlNet class constraints, but can only refer to the existing parts and cannot maintain consistency with the regenerated parts. For "static background layouts across segments", they are not persistent, explicit and reusable conditions.
[0029] Based on this, the video generation method, apparatus, electronic device, and computer-readable storage medium provided in this invention acquire reference images and text prompts, and generate a complete background image, an extended background image, and a target object image based on the reference images. This makes the generation of video segments dependent on a unified and extended background reference, thereby maintaining the structural and color continuity of the background area during segmented generation. Furthermore, by combining multi-stage background modeling with the complete and extended background images, the generated video segments maintain background consistency while also realizing dynamic changes of the target object under rich semantic guidance. After splicing, a continuous and complete video content is obtained, maintaining the static steady-state characteristics of the background, significantly reducing cross-segment background texture and tone drift, and achieving stable presentation of the background area under changes in the subject's actions.
[0030] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0031] Please refer to Figure 1 , Figure 1 A schematic flowchart of a video generation method provided by an embodiment of the present invention is shown. The method includes the following steps: Step S10: Obtain the reference image and text prompts corresponding to each video segment to be generated.
[0032] In this embodiment of the invention, by acquiring the reference images and text prompts required to generate each video segment at once, and combining this with a background and subject decoupling mechanism, a long video with background consistency is generated. Before generating each video segment, it is necessary to acquire the reference images and text prompts required to generate that video segment.
[0033] A reference image is an image that includes a dynamic subject (such as a digitized human) and its background. This image can be obtained from a previously generated video clip, or the same reference image can be used for all video clips. This flexibility makes it more adaptable to video generation tasks in different scenarios. It should be understood that the method of obtaining the reference image is determined by the application scenario. For example, in a live broadcast, if the digitized human's movements are relatively minor, a uniform reference image can be used to enhance background consistency; while when the movements are more varied, the reference image may need to be dynamically updated.
[0034] Step S20: Obtain the complete background image, the extended background image, and the target object image based on the reference image.
[0035] Among them, the target object image is an image containing a person segmented from the reference image; the complete background image is an image that includes the background in the reference image and is then completed; and the expanded background image is an image obtained by expanding the complete background image.
[0036] In this embodiment of the invention, the target object map is defined as an image region containing a dynamic subject (e.g., a person) extracted from a reference image using segmentation techniques. Its purpose is to separate the subject from the background, preventing changes in the subject's movement from affecting the consistency of the background during diffusion. A complete background map refers to a complete background image obtained by filling in the missing areas after separating the background region from the reference image.
[0037] The expanded background image is a complete background image that is expanded outwards to accommodate changes in the background range caused by camera movements (such as dance moves), thereby ensuring that the generated video can maintain the consistency of the background when the perspective changes.
[0038] Step S30: Generate corresponding video clips based on the reference image, the complete background image, the extended background image, the target object image, and the text prompts.
[0039] In this embodiment of the invention, the background image and the target object image are processed separately to decouple the static nature of the background from the dynamic nature of the subject, thereby ensuring the richness of action in the video clip while maintaining the stability of the background.
[0040] Step S40: The video segments are spliced together to obtain a complete video.
[0041] In this embodiment of the invention, video segments are spliced together to generate a complete long video. During the splicing process, the boundary areas of the video segments need to be consistent to prevent visual jumps caused by randomness or illumination drift.
[0042] In summary, the video generation method provided by this invention acquires reference images and text prompts, and generates a complete background image, an extended background image, and a target object image based on the reference images. This ensures that the generation of video segments relies on a unified and extended background reference, thereby maintaining the structural and color continuity of the background area during segmented generation. Furthermore, by combining multi-stage background modeling with the complete and extended background images, the generated video segments maintain background consistency while also enabling dynamic changes in the target object under rich semantic guidance. The resulting continuous and complete video content, after splicing, maintains the static stability of the background, significantly reduces cross-segment background texture and tone drift, and achieves stable presentation of the background area under changes in the subject's actions.
[0043] Alternatively, the following is one possible implementation for generating the complete background image, the extended background image, and the target object image. Figure 1 The sub-steps of step S20 may include: Step S200: Based on the segmentation of the reference image of the target object, a target object image and a fragmented background image are obtained; the background region contained in each reference image remains unchanged during the video generation process.
[0044] In this embodiment of the invention, an image segmentation operation is first performed on the reference image. This operation is based on an image semantic segmentation model and is used to extract the region containing the dynamic subject from the reference image as the target object image, while the remaining part excluding the target object region is used as a fragmented background image. For example, the person in a portrait image is cut out to obtain a target object image containing only the person, and a fragmented background image containing the other regions excluding the person.
[0045] It should be understood that the incomplete background image specifically includes the background area in the reference image that is not occluded by the dynamic subject. It usually contains background information that is missing due to subject occlusion and cannot be directly used to generate a consistent background representation.
[0046] It is worth noting that the background areas in these reference images remain unchanged throughout the entire video generation process. That is, no matter how many segments are generated later, they are all constructed based on the same original background and will not be regenerated or have their styles changed as the main subject moves. Step S210: Fill in the incomplete background image to obtain the complete background image.
[0047] In this embodiment of the invention, the filling process can employ an image completion model or context-based inpainting (i.e., filling) technology. By reasonably filling the missing areas in the incomplete background image, the complete background image maintains continuity and consistency in visual structure, providing basic data support for subsequent expansion processing.
[0048] Step S220: Expand the background image to obtain an expanded background image.
[0049] In this embodiment of the invention, background expansion specifically includes an outpainting operation based on an image generation model. This involves expanding the image content outward from the boundary region of the complete background image to simulate background areas that may be exposed when the camera's viewing angle changes. For example, an existing expansion model can be used to expand the complete background image based on an expansion mask region and / or cue information. The expansion mask region is supplemented with colors and textures consistent with the style of the complete background image, and / or image details consistent with the semantics of the cue information, thereby obtaining an expanded background image. This expanded background image is used in subsequent video generation processes to enhance the stability of the background region during camera movement or changes in viewing angle, thereby avoiding abrupt changes or breaks in the background caused by viewing angle shifts.
[0050] It should be noted that when expanding the background, color uniformity and high dynamic range (HDR) compression can be used to ensure color consistency across different ranges.
[0051] As can be seen, this embodiment of the invention decouples the dynamic subject and background in the reference image before generating a long video, and fills the background according to the background consistency requirements based on the decoupled background, so that the complete background image maintains continuity and consistency in visual structure. Finally, it expands the complete background image according to the requirement of color uniformity to obtain an expanded background image with consistent background. Before the video segment is generated, a background representation with complete structure and expansion capability is constructed in advance, thereby effectively improving the structural stability and visual consistency of the background area during long video generation, and providing reliable basic data support for background consistency control in subsequent video segment generation.
[0052] Alternatively, the following is one possible implementation for generating the complete background image, the extended background image, and the target object image. Figure 1 The sub-steps of step S30 may include: Step S300: Generate latent features based on the complete background image, the extended background image, and the reference image; the latent features are used to characterize the structure, color, and texture of the image.
[0053] In this embodiment of the invention, the complete background image, the extended background image, and the reference image are encoded by the encoder of the VAE model, converted into feature representations in the latent space, and then noise is added to the obtained feature representations to obtain latent features. The latent features carry the structural, color, and texture information of the image, providing a low-dimensional abstraction of visual content for the generation of video clips.
[0054] It should be noted that, to ensure consistency in noise scheduling, the same noise can be used when adding noise, or the same mean and variance can be used to sample the noise.
[0055] Step S310: Generate target object features based on the target object image and generate complete background features based on the complete background image; the complete background features are used to characterize the static visual appearance of the background.
[0056] In one embodiment of the invention, as a possible implementation, an image encoder is used to extract high-level semantic features of the target object image, and the features extracted from the target object image are directly used as the target object features. An image encoder is also used to extract high-level semantic features of the complete background image, specifically the target object features, and the features extracted based on the complete background image are directly used as the complete background features. Here, the target object features are used to express the appearance and posture information of the dynamic subject, while the complete background features are used to express the static visual features of the background.
[0057] Step S320: Generate prompt features based on the extended background image and text prompt words; prompt features are used to characterize the fusion information of prompt semantics and visual style.
[0058] In this embodiment of the invention, a prompt feature is generated based on the extended background image and text prompt words through a multimodal fusion mechanism. The prompt feature not only contains the semantic information of the text description, but also integrates the visual style information of the extended background image, thereby generating dynamic content that conforms to the text description and is consistent with the background style.
[0059] Step S360: Generate corresponding video segments based on potential features, cue features, target object features, and complete background features.
[0060] In this embodiment of the invention, latent features provide the foundation for image structure, cue features control the semantics and style of generated content, target object features ensure the coherence of dynamic subject (e.g., digital human) actions, and complete background features are used to maintain the stability of the background region over time. Through the synergistic effect of multimodal features, the visual consistency and stability of the background region in long videos are significantly improved while ensuring the quality of video clips.
[0061] As can be seen, the embodiments of the present invention extract features from the complete background image, the extended background image, and the reference image to generate latent features, enabling the video generation process to build a foundation based on the original visual content of the images, thereby ensuring the structural accuracy of the generated content; by extracting target object features and complete background features from the target object image and the complete background image respectively, independent modeling of the subject and background information is achieved, ensuring that changes in the subject's actions do not interfere with the stability of the background area; by combining the extended background image and text prompts to generate prompt features, the generation process can simultaneously take into account the semantic guidance of the text and the consistency control of the background visual style.
[0062] Finally, by fusing latent features, cue features, target object features and complete background features to generate video clips, the generated results significantly improve the visual consistency and stability of the background area in the time dimension while maintaining the richness of the main action, thus effectively solving the problems of background color shift, texture breakage and style drift in the generation of long video segments.
[0063] Optionally, regarding how to generate the prompt features, the following is a possible implementation. The sub-steps of step S320 may include: Step S320-1: Generate background prompts based on the background image.
[0064] In this embodiment of the invention, a multimodal model is used to perform content recognition on the extended background image and generate a natural language description that matches its visual content, namely, extended background cue words. The extended background cue words are used to express information such as scene semantics, color style, and spatial layout contained in the extended background image.
[0065] Step S320-2: Generate extended background features based on extended background cue words, and generate text features based on text cue words.
[0066] In this embodiment of the invention, an extended background cue words are converted into high-dimensional vector representations, i.e., extended background features, using a text encoder to characterize the visual semantics of the background image in the latent space. Similarly, text cue words are converted into text features using a text encoder to express the user's intended action or style description as depicted in the generated video. These two features respectively carry background visual information and user intent.
[0067] Step S320-3: Generate prompt features based on the extended background features and text features.
[0068] In this embodiment of the invention, the extended background features and text features are multimodally fused by means of feature splicing, weighted fusion and other methods to generate the final prompt features. By combining visual background semantics and text instructions, more precise semantic control and visual consistency constraints on the generated content are achieved, thereby improving the stability and coherence of the background area when the long video is generated in segments.
[0069] As can be seen, this embodiment of the invention generates extended background prompts based on the extended background image and extracts extended background features based on these prompts, enabling the structured expression of the visual semantic information of the background. Simultaneously, text features are generated based on the text prompts, allowing the model to accurately understand the user's specified actions or style intentions. Finally, by fusing extended background features and text features to generate prompt features, the collaborative guidance of text semantics and background visual style is achieved, thereby enhancing the semantic consistency and visual stability of the generated video clips in the background area and ensuring that the background content remains coherent and controllable during the dynamic generation process of long videos.
[0070] Optionally, regarding how to construct the background cache pool, the following is a possible implementation. After step S320, the following may also be included: Step S330: Generate background key-value pairs based on the prompt features and features extracted from the complete background image.
[0071] In this embodiment of the invention, to ensure that the background remains consistent across different segments of the generated long video, after generating the cue features, a set of background key-value pairs is generated by combining the obtained cue features with image features extracted from the complete background image. The cue features are generated jointly by the extended background image and text cue words, containing semantic information about the extended background description and the action intent. The image features extracted from the complete background image can characterize the structural and textural features of the background region in the reference image.
[0072] It should be understood that background key-value pairs are control structures used to stabilize background representation. The "key" refers to the fused feature generated based on cue features and features extracted from the complete background image, containing static information such as background color, layout, and texture. It is equivalent to a recognition label that tells the model which background it is currently facing. The "value" refers to the reusable computational result corresponding to this background, that is, the key vector and value vector generated in the cross-attention layer of the diffusion model.
[0073] Step S340: Generate background prompts based on the background image.
[0074] In this embodiment of the invention, a multimodal model is used to extract extended background cue words from the extended background image. An extended background cue word is a natural language text or semantic vector describing the background content. For example, an extended background cue word could be "a sunny living room with a sofa, television, and windows in the background."
[0075] Step S350: If the similarity between the token vector corresponding to the background key-value pair and the token vector corresponding to the extended background prompt is higher than the first preset threshold, the background key-value pair is saved to the background cache pool.
[0076] In this embodiment of the invention, a corresponding token vector is generated based on the background key-value pair, a corresponding token vector is generated based on the extended background prompt word, and the similarity between the token vector corresponding to the background key-value pair and the token vector corresponding to the extended background prompt word is calculated. If the similarity is higher than a first preset threshold (e.g., 90%), the background key-value pair is saved to the background cache pool.
[0077] It should be noted that the token vector here refers to the vectorized representation of text features in the model, and similarity calculation can be implemented using methods such as cosine similarity. This judgment mechanism is used to filter out key-value pairs that are highly similar to the current background semantics, ensuring that only highly representative background features are retained in the background cache pool, thereby providing a stable background control signal for the generation of subsequent video segments.
[0078] As can be seen, this embodiment of the invention constructs a memory of background information through a background cache pool, which can be reused in the same long video, reducing the time spent generating long videos while ensuring background consistency across long videos. Optionally, regarding how to utilize the background cache pool to generate complete background features, the following is a possible implementation method. The sub-step of generating complete background features in step S310 may further include: Step S310-1: Retrieve the previous background key-value pair from the background cache pool.
[0079] In this embodiment of the invention, to ensure that the background remains consistent across generated video segments, the background key-value pairs previously stored in the background cache pool are fully utilized when processing the current segment. During the generation of complete background features, the previously stored background key-value pair is first retrieved from the background cache pool. The background cache pool is built incrementally and stores background key-value pairs that semantically match the extended background cue words; it can be understood as a stable memory of the static background.
[0080] Step S310-2: The features extracted based on the complete background image are fused with the previous background key-value pair to obtain the complete background features.
[0081] Next, the features extracted from the complete background image are fused with the previous background key-value pair. The fusion methods include, but are not limited to, weighted combination, splicing, and attention mechanisms. The goal is to ensure that the background representation of the current frame retains the original bytes while continuing the style and structure of the previous segment.
[0082] In this way, the complete background features used in the current frame not only depend on the current input reference image, but also inherit consistent background information from previous segments. This effectively reduces issues such as abrupt background changes, color shifts, or texture breaks between different segments, even when generating long videos in segments.
[0083] For example, in practical applications, when generating a live video of an AI digital human dancing, the background is a living room scene containing a sofa, window, and bookshelves. Although the digital human's movements are constantly changing, the texture, color, and layout of the background area remain consistent in each frame of the video because the background key-value pairs are continuously cached and merged. This avoids the background "jumping" phenomenon commonly seen in traditional segmented generation methods.
[0084] As can be seen, by reusing the previous background key-value pair in the background cache pool and fusing it with the features extracted from the current complete background image, the embodiments of the present invention can effectively enhance the stability and continuity of background features in cross-segment generation without recalculating all background information, significantly improve the temporal consistency and visual coherence of the background region in long videos, and thus achieve high-quality, low-stitching-distortion long video generation effects.
[0085] Optionally, regarding how to construct the action cache pool, the following is one possible implementation. The video generation method may also include: Action prompts are generated based on the target object graph, and action key-value pairs are generated based on the action prompts. If the similarity between the token vector corresponding to the action key-value pair and the token vector corresponding to the text prompt is higher than a second preset threshold, the action key-value pair is saved to the action cache pool. The action cache pool is used to retrieve the previous action key-value pair from the action cache pool and fuse it with the features extracted based on the target object graph when generating target object features.
[0086] In this embodiment of the invention, when generating long videos, in addition to maintaining a stable background as much as possible, it is also necessary to ensure that the main actions are coherent and natural. In order to maintain the coherence of the main actions, a multimodal model is used to extract actions from the target object graph to obtain action cue words. The action cue words are used to express the action features presented by the current target object graph, providing a semantic basis for the subsequent generation of action key-value pairs, such as raising the right hand or taking a step forward.
[0087] Action key-value pairs are control structures used to maintain the continuity of the subject's actions. The "key" is the semantic feature corresponding to the action cue word generated from the target object graph, such as waving or making a heart shape. The "value" is the dynamic response parameter required for the corresponding action during the denoising process, which determines the way the limb movement trajectory and posture changes are expressed.
[0088] Token vectors are generated based on action key-value pairs and text prompts, and the similarity between the token vectors corresponding to the action key-value pairs and the token vectors corresponding to the text prompts is calculated. If the calculated similarity is higher than a second preset threshold, the action key-value pairs are saved to the action cache pool. This judgment mechanism is used to filter out historical action features that highly match the current action intent, ensuring that only highly representative action features are retained in the cache pool, thereby providing stable action control signals for the generation of subsequent segments.
[0089] The action cache pool is used to fuse features corresponding to the target object graph when generating video clips. Specifically, during the generation of the current video clip, the previous action key-value pair is retrieved from the action cache pool and fused with the features extracted from the current target object graph to obtain the target object features. These features are then combined with latent features, cue features, and complete background features to generate the video clip. This aims to combine historical action information with current action information to ensure the continuity and consistency of character movements in the generated footage.
[0090] Understandably, the action key-value pair caching mechanism essentially constructs a "memory" of action information, allowing the generation of new segments to reference action features from historical segments, thereby effectively improving the stability and efficiency of long video generation. The action key-value pair caching mechanism complements the background key-value pair caching mechanism, constructing a two-level cache that operates on the main subject and background areas respectively. This achieves layered control over the generated video content, and the action key-value pair and background key-value pair sharing mechanism maintains consistency among the generated multiple video segments.
[0091] For example, in practical applications, when generating a live video of an AI digital human dancing, although the reference image of each frame may remain unchanged or be inherited from the previous frame, the motion key-value pairs are continuously cached and fused for use. Therefore, the digital human's dance movements can maintain a smooth transition between different segments, avoiding the common problems of motion abruptness or inconsistent posture in traditional generation methods.
[0092] As can be seen, by introducing a mechanism for generating, filtering, and caching action key-value pairs, this embodiment of the invention enables the reuse of historical action key-value pairs within the same video, reducing inference time. By using historical action key-value pairs from the action cache pool for the current video segment, the continuity of character actions can be effectively preserved during the continuous generation of video segments, avoiding motion distortion caused by randomness or differences in conditions between segments.
[0093] One possible approach is to pre-train a video generation model, inputting reference images and text prompts from the user, so that the model can generate a long video based on the reference images and text prompts.
[0094] by Figure 2 For example, video generation models include, but are not limited to, image segmentation models, subject image encoders, background image encoders, text encoders, and video clip generation models. Among these, video clip generation models include variational autoencoders and diffusion models; variational autoencoders include encoders and decoders.
[0095] The reference image and text prompts input by the user are fed into the video generation model. First, the reference image is segmented using a segmentation model to obtain a target object image and a fragmented background image. The target object image is then encoded using a subject image encoder to obtain target object features. Finally, the fragmented background image is filled in to obtain a complete background image, which is then encoded using a background image encoder to obtain complete background features.
[0096] Next, the complete background image is expanded to obtain an expanded background image, and a multimodal model is used to generate expanded background cue words based on the expanded background image. A text encoder is then used to process the expanded background cue words and the text cue words corresponding to the first video segment to obtain cue features.
[0097] Subsequently, the cue features, complete background features, target object features, reference image, complete background map, and extended background map are input into the video segment generation model. This allows the variational autoencoder's encoder to encode based on the reference image, complete background map, and extended background map, and to superimpose noise to generate latent features. The latent features, cue features, complete background features, and target object features are then injected into a diffusion model. This model processes these features and inputs the resulting vectors into the variational autoencoder's decoder to obtain the first video segment. Injection methods include, but are not limited to, cross-attention and contextualization.
[0098] It should be understood that, in order to inject better into the background region, the background key-value pairs corresponding to the previous block layer and the background region will be fixed with a certain probability in the early steps of denoising.
[0099] Similarly, the reference image and text prompt for the next video segment to be generated are obtained sequentially. The reference image can be obtained from the previous video segment or a reference image input by the user. The video generation model is used to generate each video segment in sequence and then the video segments are spliced together to obtain the complete video.
[0100] It should be noted that in existing technologies, when generating video clips using a diffusion model, the features corresponding to the reference image are injected into the diffusion model in a single injection. To improve the background consistency of long videos, this invention creatively pre-segments the people and background in the reference image and injects the complete background features and target object features into the diffusion model in two injections.
[0101] Considering that the original structure of the diffusion model only has one injection module, the background injection and character injection will reuse the same injection structure. Referring to ControlNet's zero initialization, the parameters of this module will be fully learned during the training phase, instead of just training the rank parameters.
[0102] When training the video generation model, the parameters of the video segment generation model are iteratively updated using reconstruction loss, background consistency loss, and attention consistency loss to minimize the model training cost, ultimately resulting in the trained video generation model. The reconstruction loss is based on diffusion inversion or conventional conditional reconstruction, calculating reconstruction loss (e.g., L1 / Charbonnier) and perceptual loss (e.g., LPIPS) in the background mask region for the predicted frame (i.e., the decoder output) and the target frame (i.e., the original video pointer).
[0103] Background consistency loss is achieved by calculating color-histogram matching loss and style statistics (e.g., AdaIN mean / variance) consistency loss across frames or segments in the background region to suppress color drift. Attention consistency loss applies Kullback-Leibler (KL) consistency regularization to the cross-attention weight distribution of the background cue word token vector across adjacent frames to ensure stable background attention.
[0104] As can be seen, the embodiments of the present invention, through zero-initialization LoRA overlay and selective hierarchical injection, reduce the number of training parameters and data requirements of the video generation model to a controllable range, quickly adapt to different live broadcast backgrounds, reduce training costs, and effectively solve the problems of high training costs, large amounts of data and graphics card resources required to avoid overfitting on small datasets in the prior art when fine-tuning the model for background consistency.
[0105] Based on the same inventive concept, the basic principle and technical effects of the video generation device provided in this embodiment are the same as those in the above embodiments. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above embodiments.
[0106] Please refer to Figure 3 , Figure 3 This is a block diagram of a video generation apparatus 400 provided in an embodiment of the present invention. The video generation apparatus 400 includes an acquisition module 401, a generation module 402, and a combination module 403.
[0107] The acquisition module 401 is used to acquire the reference image and text prompt for each video segment to be generated; the reference image is acquired from the previously generated video segment or all video segments use the same reference image.
[0108] The generation module 402 is used to obtain a complete background image, an extended background image, and a target object image based on a reference image; and to generate corresponding video clips based on the reference image, the complete background image, the extended background image, the target object image, and text prompts.
[0109] The combination module 403 is used to splice together the various video segments to obtain a complete video.
[0110] In summary, the video generation apparatus provided in this embodiment of the invention acquires reference images and text prompts, and generates a complete background image, an extended background image, and a target object image based on the reference images. This ensures that the generation of video segments relies on a unified and extended background reference, thereby maintaining the structural and color continuity of the background area during segmented generation. Furthermore, by combining multi-stage background modeling with the complete and extended background images, the generated video segments maintain background consistency while also enabling dynamic changes in the target object under rich semantic guidance. The resulting continuous and complete video content, after splicing, maintains the static stability of the background, significantly reduces cross-segment background texture and tone drift, and achieves stable presentation of the background area under changes in the subject's actions.
[0111] Optionally, the generation module 402 is specifically used to segment reference images based on the target object to obtain a target object image and a fragmented background image; the background regions contained in each reference image remain unchanged during the video generation process; the fragmented background image is filled to obtain a complete background image; and the complete background image is expanded to obtain an expanded background image.
[0112] Optionally, the generation module 402 is specifically used to generate latent features based on the complete background image, the extended background image, and the reference image; the latent features are used to characterize the structure, color, and texture of the image; generate target object features based on the target object image, and generate complete background features based on the complete background image; the complete background features are used to characterize the static visual of the background; generate cue features based on the extended background image and the text cue words; the cue features are used to characterize the fusion information of cue semantics and visual style; and generate corresponding video segments based on the latent features, cue features, target object features, and complete background features.
[0113] Optionally, the generation module 402 is specifically used to generate an extended background prompt word based on the extended background image; generate an extended background feature based on the extended background prompt word, and generate text features based on the text prompt word; and generate prompt features based on the extended background feature and the text feature.
[0114] Optionally, the generation module 402 is further configured to generate background key-value pairs based on the prompt features and features extracted based on the complete background image; generate extended background prompt words based on the extended background image; and if the similarity between the token vector corresponding to the background key-value pair and the token vector corresponding to the extended background prompt word is higher than a first preset threshold, save the background key-value pair to the background cache pool.
[0115] Optionally, the generation module 402 is specifically used to obtain the previous background key-value pair from the background cache pool; and to fuse the features extracted based on the complete background image with the previous background key-value pair to obtain the complete background features.
[0116] Optionally, the generation module 402 is further configured to generate action prompt words based on the target object graph, and generate action key-value pairs based on the action prompts; if the similarity between the token vector corresponding to the action key-value pair and the token vector corresponding to the text prompt word is higher than a second preset threshold, the action key-value pair is saved to the action cache pool; the action cache pool is used to retrieve the previous action key-value pair from the action cache pool and fuse the features extracted based on the target object graph when generating target object features.
[0117] Please refer to Figure 4 This is a block diagram illustrating an electronic device 500 provided in an embodiment of the present invention. The electronic device 500 includes, but is not limited to, a personal computer (PC), a handheld computer (PDA), a laptop computer, a tablet computer, and a server. The electronic device 500 includes a memory 510, a processor 520, and a communication module 530. The memory 510, processor 520, and communication module 530 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0118] The memory 510 is used to store programs or data. The memory 510 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0119] The processor 520 is used to read / write data or programs stored in the memory 510 and perform corresponding functions. For example, when a computer program stored in the memory 510 is executed by the processor 520, the video generation method disclosed in the above embodiments can be implemented.
[0120] The communication module 530 is used to establish a communication connection between the electronic device 500 and other communication terminals via a network, and to send and receive data via the network.
[0121] It should be understood that, Figure 4The structure shown is only a schematic diagram of the electronic device 500. The electronic device 500 may also include components that are larger than those shown. Figure 4 The more or fewer components shown, or having the same Figure 4 The different configurations shown. Figure 4 The components shown can be implemented using hardware, software, or a combination thereof.
[0122] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor 520, implements the video generation method disclosed in the above embodiments.
[0123] This invention also provides a program product that, when executed by processor 520, implements the video generation method disclosed in the above embodiments.
[0124] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0125] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0126] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0127] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A video generation method, characterized in that, The method includes: Obtain the reference image and text prompts for each video segment to be generated; Based on the reference image, a complete background image, an extended background image, and a target object image are obtained; Generate corresponding video clips based on the reference image, the complete background image, the extended background image, the target object image, and the text prompts; The video segments are spliced together to obtain a complete video.
2. The video generation method according to claim 1, characterized in that, The process of obtaining the complete background image, the extended background image, and the target object image based on the reference image includes: The reference images are segmented based on the target object to obtain the target object image and the incomplete background image; the background regions contained in each reference image remain unchanged during the video generation process; The incomplete background image is filled in to obtain the complete background image; The complete background image is expanded to obtain the expanded background image.
3. The video generation method according to claim 1, characterized in that, The step of generating corresponding video clips based on the reference image, the complete background image, the extended background image, the target object image, and the text prompts includes: Latent features are generated based on the complete background image, the extended background image, and the reference image; the latent features are used to characterize the structure, color, and texture of the image; Target object features are generated based on the target object image, and complete background features are generated based on the complete background image; the complete background features are used to characterize the static visual appearance of the background. A prompt feature is generated based on the extended background image and the text prompt words; the prompt feature is used to characterize the fusion information of prompt semantics and visual style; Generate corresponding video clips based on the potential features, the cue features, the target object features, and the complete background features.
4. The video generation method according to claim 3, characterized in that, The step of generating prompt features based on the extended background image and the text prompt words includes: Generate an expanded background prompt based on the expanded background image; Generate an extended background feature based on the extended background cue words, and generate text features based on the text cue words; The prompt feature is generated based on the background features and the text features.
5. The video generation method according to any one of claims 3-4, characterized in that, After generating prompt features based on the extended background image and the text prompt words, the method further includes: Generate background key-value pairs based on the prompt features and the features extracted from the complete background image; Generate an expanded background prompt based on the expanded background image; If the similarity between the token vector corresponding to the background key-value pair and the token vector corresponding to the extended background prompt word is higher than a first preset threshold, the background key-value pair is saved to the background cache pool.
6. The video generation method according to claim 5, characterized in that, The step of generating complete background features based on the complete background image includes: Retrieve the previous background key-value pair from the background cache pool; The features extracted based on the complete background image are fused with the previous background key-value pair to obtain the complete background features.
7. The video generation method according to claim 3, characterized in that, The method further includes: Generate action prompts based on the target object graph, and generate action key-value pairs based on the action prompts; If the similarity between the token vector corresponding to the action key-value pair and the token vector corresponding to the text prompt word is higher than a second preset threshold, the action key-value pair is saved to the action cache pool; the action cache pool is used to retrieve the previous action key-value pair from the action cache pool and fuse it with the features extracted based on the target object graph when generating target object features.
8. A video generation apparatus, characterized in that, The device includes: The acquisition module is used to acquire the reference image and text prompts corresponding to each video segment to be generated; The generation module is used to obtain a complete background image, an extended background image, and a target object image based on the reference image; and to generate corresponding video clips based on the reference image, the complete background image, the extended background image, the target object image, and the text prompt words. The combination module is used to splice the various video segments to obtain a complete video.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program executable by the processor, the processor being able to execute the computer program to implement the video generation method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the video generation method as described in any one of claims 1-7.