A video processing method and device, electronic equipment and storage medium

By aggregating historical frame feature information of video frames during the video generation process, the problem of insufficient visual consistency in long video generation is solved, achieving high-quality long video generation and improving user experience.

CN122120573APending Publication Date: 2026-05-29BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2026-03-02
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, the fine-grained control of the video content output by the model is insufficient when generating long videos, and it is impossible to maintain visual consistency in multiple frames, resulting in poor generation effects and difficulty in meeting the requirements of high visual fidelity.

Method used

By acquiring object images and video generation instruction text, and using a video script generation model to aggregate video script information into script image feature information based on historical frame feature information of each video frame, the video script information is input into the video generation model for video generation, thereby achieving fine-grained control and maintaining the visual consistency and plot coherence of multiple frames in a long video.

Benefits of technology

It improves the quality of long video generation, maintains visual consistency and narrative coherence across multiple frames, and enhances the user's video production experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120573A_ABST
    Figure CN122120573A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video processing method and device, electronic equipment and storage medium. The method comprises: obtaining an object image and a video generation instruction text; the video generation instruction text indicates to generate a video with a coherent plot associated with a target object, and the object image is a reference image containing the target object; inputting the object image and the video generation instruction text into a video script generation model; in the process of generating a video script based on the object image and the video generation instruction text, aggregating video script information of each video frame into script image feature information based on historical frame feature information corresponding to each video frame, to obtain script image feature information corresponding to a plurality of video frames; inputting the object image and the script image feature information corresponding to the plurality of video frames into a video generation model to generate a video, to obtain a video containing generated images corresponding to the plurality of video frames. The embodiments of the present disclosure can maintain the fine-grained control effect of long video generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a video processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the continuous development of artificial intelligence technology, AI has made significant progress in generating high-quality short videos and single images. However, in practical applications (such as film production and e-commerce advertising), it is often necessary to generate long video content with a coherent plot.

[0003] In related technologies, large language models are usually used to generate storyboards, and then the storyboards are input into text-to-image models for image generation. However, the video content output by the model often suffers from insufficient fine-grained control and cannot maintain visual consistency across multiple frames. For example, it cannot maintain specific character features and environmental details, resulting in poor long video generation effects and making it difficult to meet the requirements of long video generation scenarios with extremely high visual fidelity. Summary of the Invention

[0004] This disclosure provides a video processing method, apparatus, electronic device, and storage medium to at least solve the technical problems in related technologies, such as insufficient fine-grained control of video content output by models, inability to maintain visual consistency across multiple frames of video content, resulting in poor long video generation effects and difficulty in meeting the requirements of long video generation scenarios with extremely high visual fidelity. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a video processing method is provided, comprising: Obtain an object image and video generation instruction text; the video generation instruction text is used to instruct the generation of a video with a coherent plot associated with the target object, and the object image is a reference image containing the target object; The object image and the video generation instruction text are input into the video script generation model. During the process of generating the video script based on the object image and the video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame, so as to obtain script image feature information corresponding to multiple video frames. The object image and the script image feature information corresponding to the multiple video frames are input into the video generation model to generate a target generated video containing the generated images corresponding to the multiple video frames.

[0005] In an optional embodiment, during the process of generating a video script based on the object image and the video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame, resulting in script image feature information corresponding to multiple video frames, including: For the current video frame among the plurality of video frames, a video script is generated for the current video frame based on the object image, the video generation instruction text, and the video script information corresponding to the historical video frames, to obtain the video script information corresponding to the current video frame; Image feature aggregation is performed based on the object image, the video generation instruction text, the video script information corresponding to the historical video frames, the script image feature information corresponding to the historical video frames, and the video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame; The historical video frames are the video frames that precede the current video frame among the plurality of video frames.

[0006] In an optional embodiment, the step of generating a video script for the current video frame based on the object image, the video generation instruction text, and the video script information corresponding to historical video frames, to obtain the video script information corresponding to the current video frame, includes: Causal self-attention processing is performed on the object image, the video generation instruction text, and the video script information corresponding to the historical video frames to obtain the script text feature information corresponding to the current video frame; Based on the script text feature information corresponding to the current video frame, a video script is generated to obtain the video script information corresponding to the current video frame. The step of aggregating image features based on the object image, the video generation instruction text, the video script information corresponding to the historical video frames, the script image feature information corresponding to the historical video frames, and the video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame includes: Based on pre-trained image query feature information, causal attention processing is performed on the object image, the video generation instruction text, the video script information corresponding to the historical video frames, the script image feature information corresponding to the historical video frames, and the video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame.

[0007] In an optional embodiment, the video generation model includes: an image generation model, wherein inputting the object image and script image feature information corresponding to the plurality of video frames into the video generation model to generate a target generated video containing generated images corresponding to the plurality of video frames includes: The object image and the script image feature information corresponding to the plurality of video frames are input into the image generation model. Based on the object image, the script image feature information corresponding to the current video frame in the plurality of video frames, and the generated images corresponding to the historical video frames, an image is generated to obtain the generated image corresponding to the current video frame. The target generated video is obtained based on the generated images corresponding to the multiple video frames; The historical video frames are the video frames that precede the current video frame among the plurality of video frames.

[0008] In an optional embodiment, the step of generating an image based on the object image, the script image feature information corresponding to the current video frame among the plurality of video frames, and the generated images corresponding to historical video frames to obtain the generated image corresponding to the current video frame includes: Based on the distance between the historical video frame and the current video frame, the image feature information of the generated image corresponding to the historical video frame is attenuated and compressed to obtain the historical image feature information corresponding to the current video frame. Image generation is performed based on the object image feature information of the object image, the script image feature information corresponding to the current video frame, and the historical image feature information to obtain the generated image corresponding to the current video frame.

[0009] In an optional embodiment, the step of attenuating and compressing the image feature information of the generated image corresponding to the historical video frame based on the distance between the historical video frame and the current video frame to obtain the historical image feature information corresponding to the current video frame includes: After the generated image corresponding to the previous video frame of the current video frame is generated, the historical image feature information corresponding to the previous video frame is extracted from the feature cache library; The historical image feature information corresponding to the previous video frame and the image feature information of the generated image corresponding to the previous video frame are compressed to obtain the historical image feature information corresponding to the current video frame. In the feature cache library, the historical image feature information corresponding to the previous video frame is updated to the historical image feature information corresponding to the current video frame.

[0010] In an optional embodiment, the video generation model includes a feature alignment model and an image generation model. The object image and the video generation instruction text are input into the video script generation model. During the video script generation process based on the object image and the video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame. The resulting script image feature information corresponding to multiple video frames includes: The object image and the video generation instruction text are input into the video script generation model. During the process of generating the video script based on the first image feature information of the object image and the video generation instruction text, the video script information of each generated video frame and the historical frame feature information corresponding to each video frame are aggregated into script image feature information based on the pre-trained image query feature information to obtain the script image feature information corresponding to the multiple video frames. The step of inputting the object image and the script image feature information corresponding to the plurality of video frames into the video generation model to generate a target generated video containing the generated images corresponding to the plurality of video frames includes: Obtain the second image feature information of the object image; The script image feature information corresponding to the multiple video frames is input into the feature alignment model for feature alignment, thereby obtaining the aligned image feature information corresponding to the multiple video frames. The second image feature information and the aligned image feature information corresponding to the plurality of video frames are input into the image generation model to generate an image, thereby obtaining the generated image corresponding to the plurality of video frames; The target generated video is obtained based on the generated images corresponding to the multiple video frames.

[0011] In an optional embodiment, the method further includes: The sample instruction text, the sample object image corresponding to the sample instruction text, the preset video script information of multiple video frames corresponding to the sample instruction text, and the preset generated image of multiple video frames corresponding to the preset video script information are obtained; the sample instruction text is used to instruct the generation of a video with a coherent plot associated with the sample object, and the sample object image is a sample reference image containing the sample object. Based on the sample image feature information corresponding to the sample object image, the sample instruction text, and the preset video script information of the multiple video frames, the video script generation model to be trained is trained to obtain the video script generation model. Based on the preset video script information and the preset generated image of each video frame, feature space alignment training is performed on the image query feature information to be trained and the feature alignment model to be trained to obtain the pre-trained image query feature information and the feature alignment model. Based on the sample object image, the sample script image feature information corresponding to the preset video script information, and the preset generated images of the multiple video frames, the image generation model to be trained is trained to obtain the image generation model.

[0012] In an optional embodiment, the method further includes: Obtain the sample object description information corresponding to the sample object; Based on the sample object description information, a reference script is generated to obtain reference script information; Based on the reference script information, a reference image is generated to obtain the sample object image; Based on the sample object image and the sample object description information, a video script is generated to obtain the preset video script information of the multiple video frames; Image generation is performed based on the sample object image and the preset video script information of the multiple video frames to obtain the preset generated image of the multiple video frames; Based on the preset video script information and the preset generated image, the instruction settings are performed to obtain the sample instruction text.

[0013] According to a second aspect of the present disclosure, a video processing apparatus is provided, comprising: The instruction acquisition module is configured to acquire an object image and video generation instruction text; the video generation instruction text is used to instruct the generation of a video with a coherent plot associated with the target object, and the object image is a reference image containing the target object; The video script generation module is configured to input the object image and the video generation instruction text into the video script generation model. During the process of generating the video script based on the object image and the video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame, thereby obtaining script image feature information corresponding to multiple video frames. The video generation module is configured to input the object image and the script image feature information corresponding to the multiple video frames into the video generation model to generate a target generated video containing the generated images corresponding to the multiple video frames.

[0014] In an optional embodiment, the video script generation module includes: The script generation unit is configured to perform video script generation on the current video frame among the plurality of video frames, based on the object image, the video generation instruction text, and video script information corresponding to the historical video frames, to obtain the video script information corresponding to the current video frame. The image feature aggregation unit is configured to perform image feature aggregation based on the object image, the video generation instruction text, the video script information corresponding to the historical video frames, the script image feature information corresponding to the historical video frames, and the video script information corresponding to the current video frame, to obtain the script image feature information corresponding to the current video frame; The historical video frames are the video frames that precede the current video frame among the plurality of video frames.

[0015] In an optional embodiment, the script generation unit is further configured to perform causal self-attention processing on the object image, the video generation instruction text, and the video script information corresponding to the historical video frames to obtain script text feature information corresponding to the current video frame; and generate a video script based on the script text feature information corresponding to the current video frame to obtain video script information corresponding to the current video frame. The image feature aggregation unit is further configured to perform pre-trained image query feature information, and to perform causal attention processing on the object image, the video generation instruction text, the video script information corresponding to the historical video frames, the script image feature information corresponding to the historical video frames, and the video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame.

[0016] In an optional embodiment, the video generation model includes: an image generation model, and the video generation module includes: The image generation unit is configured to input the object image and the script image feature information corresponding to the plurality of video frames into the image generation model, and generate an image based on the object image, the script image feature information corresponding to the current video frame in the plurality of video frames and the generated images corresponding to the historical video frames, to obtain the generated image corresponding to the current video frame; The video synthesis unit is configured to generate images based on the multiple video frames to obtain the target generated video; The historical video frames are the video frames that precede the current video frame among the plurality of video frames.

[0017] In an optional embodiment, the image generation unit is further configured to perform attenuation compression on the image feature information of the generated image corresponding to the historical video frame based on the spacing between the historical video frame and the current video frame, to obtain the historical image feature information corresponding to the current video frame; and to generate an image based on the object image feature information of the object image, the script image feature information corresponding to the current video frame, and the historical image feature information, to obtain the generated image corresponding to the current video frame.

[0018] In an optional embodiment, the image generation unit is further configured to: after the generated image corresponding to the previous video frame is generated, extract historical image feature information corresponding to the previous video frame from the feature cache; compress the historical image feature information corresponding to the previous video frame and the image feature information of the generated image corresponding to the previous video frame to obtain historical image feature information corresponding to the current video frame; and update the historical image feature information corresponding to the previous video frame to the historical image feature information corresponding to the current video frame in the feature cache.

[0019] In an optional embodiment, the video generation model includes: a feature alignment model and an image generation model; The video script generation module is further configured to acquire first image feature information of the object image; input the first image feature information and the video generation instruction text into the video script generation model; during the video script generation process based on the first image feature information and the video generation instruction text, based on pre-trained image query feature information, aggregate the video script information of each generated video frame and the historical frame feature information corresponding to each video frame into script image feature information to obtain the script image feature information corresponding to the multiple video frames; The video generation module is further configured to: acquire second image feature information of the object image; input script image feature information corresponding to the plurality of video frames into the feature alignment model for feature alignment to obtain aligned image feature information corresponding to the plurality of video frames; input the second image feature information and the aligned image feature information corresponding to the plurality of video frames into the image generation model for image generation to obtain generated images corresponding to the plurality of video frames; and obtain the target generated video based on the generated images corresponding to the plurality of video frames.

[0020] In an optional embodiment, the apparatus further includes: The sample acquisition module is configured to acquire sample instruction text, a sample object image corresponding to the sample instruction text, preset video script information of multiple video frames corresponding to the sample instruction text, and preset generated images of multiple video frames corresponding to the preset video script information; the sample instruction text is used to instruct the generation of a video with a coherent plot associated with the sample object, and the sample object image is a sample reference image containing the sample object. The first training module is configured to execute a video script generation training on the video script generation model to be trained based on the sample image feature information corresponding to the sample object image, the sample instruction text, and the preset video script information of the multiple video frames, so as to obtain the video script generation model. The second training module is configured to execute a preset video script based on each video frame and a preset generated image for each video frame, and perform feature space alignment training on the image query feature information to be trained and the feature alignment model to be trained, so as to obtain the pre-trained image query feature information and the feature alignment model. The third training module is configured to execute image generation training on the image generation model to be trained based on the sample object image, the sample script image feature information corresponding to the preset video script information, and the preset generated images of the multiple video frames, so as to obtain the image generation model.

[0021] In an optional embodiment, the sample acquisition module is further configured to perform: Obtain the sample object description information corresponding to the sample object; Based on the sample object description information, a reference script is generated to obtain reference script information; Based on the reference script information, a reference image is generated to obtain the sample object image; Based on the sample object image and the sample object description information, a video script is generated to obtain the preset video script information of the multiple video frames; Image generation is performed based on the sample object image and the preset video script information of the multiple video frames to obtain the preset generated image of the multiple video frames; Based on the preset video script information and the preset generated image, the instruction settings are performed to obtain the sample instruction text.

[0022] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement a method as described in any one of the video processing methods of the present disclosure.

[0023] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the methods described in the video processing methods of the present disclosure.

[0024] According to a fifth aspect of the present disclosure, a computer program product including instructions is provided that, when run on a computer, causes the computer to perform any of the video processing methods described in the embodiments of the present disclosure.

[0025] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: In a long video generation scenario, this application acquires an object image and video generation instruction text. The video generation instruction text instructs the generation of a video with a coherent plot associated with a target object. The object image is a reference image containing the target object. The object image and video generation instruction text are then input into a video script generation model. During video script generation based on the object image and video generation instruction text, the video script information for each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame. This yields script image feature information corresponding to multiple video frames. Based on automatically planning a coherent video script, each video... The historical frame feature information corresponding to each frame and the text features of the video script for each video frame are aggregated into script image feature information at the image dimension. This improves the fusion effect of multimodal information and the semantic coherence of script image feature information, providing macro-level guidance for visual consistency in subsequent image generation. Then, the object image and the script image feature information corresponding to multiple video frames are input into the video generation model for video generation, resulting in a target generated video containing generated images corresponding to multiple video frames. This enables fine-grained control of the long video generation process, maintains the visual consistency and plot coherence of multiple frames in the long video, effectively improves the generation quality of long videos, and thus enhances the user's video production experience.

[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0028] Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment; Figure 2 This is a flowchart illustrating a video processing method according to an exemplary embodiment; Figure 3 This is a flowchart illustrating another video processing method according to an exemplary embodiment; Figure 4a This is a schematic diagram illustrating an attention mechanism mask according to an exemplary embodiment; Figure 4b This is a structural diagram of a video script generation model according to an exemplary embodiment; Figure 5 This is a flowchart illustrating another video processing method according to an exemplary embodiment; Figure 6 This is a structural diagram of an image generation model according to an exemplary embodiment; Figure 7 This is a flowchart illustrating another video processing method according to an exemplary embodiment; Figure 8 This is a flowchart illustrating a model training process according to an exemplary embodiment; Figure 9 This is a structural diagram of a video processing model according to an exemplary embodiment; Figure 10 This is a block diagram of a video processing apparatus according to an exemplary embodiment; Figure 11 This is a block diagram illustrating an electronic device for video processing according to an exemplary embodiment. Detailed Implementation

[0029] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0030] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar different contents and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0032] First, some nouns or terms that appear in the description of the embodiments of this disclosure shall be interpreted as follows: Large Language Model (LLM): Also known as a large model, natural language model, or large-scale language model, it refers to a natural language processing model with a large number of parameters and training data. The training process of a large language model typically employs unsupervised learning, that is, training the model using a large-scale text corpus to learn the probability distribution and rules of language. During training, the large language model usually uses a language model as the objective function, optimizing the model parameters by maximizing the predicted probability of the next word.

[0033] Multimodal Large Language Model (MLLM): Based on LLM, it integrates media data from other modalities (such as images, videos, audio, etc.), enabling the model to process information from different modalities simultaneously, better understand and express semantics, and thus improve the effectiveness and accuracy of applications.

[0034] Transformer: An encoder-decoder architecture based on an attention mechanism.

[0035] Token: In LLM, a token represents the smallest unit of information that the model can understand and generate. The definition of a token may vary in different contexts, but they are generally the basic units for multimodal data analysis and processing. Tokens are assigned numerical values ​​or identifiers, arranged in sequences or vectors, and are input into or output from the model; they are the linguistic building blocks of the model.

[0036] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment, which may include a terminal 100 and a server 200.

[0037] In an optional embodiment, terminal 100 can be used to provide services such as video processing to any user. Specifically, terminal 100 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices, or software running on the aforementioned electronic devices, such as applications. Optionally, the operating system running on the electronic device can be, but is not limited to, Android, iOS, Linux, Windows, etc.

[0038] In an optional embodiment, server 200 can provide background services to terminal 100. Specifically, server 200 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services.

[0039] In addition, it should be noted that, Figure 1 The example shown is merely one application environment provided by this disclosure. In practical applications, other application environments may also be included, such as more terminals.

[0040] In the embodiments described in this specification, the terminal 100 and the server 200 can be directly or indirectly connected through wired or wireless communication, and this disclosure does not impose any restrictions.

[0041] Figure 2 This is a flowchart illustrating a video processing method according to an exemplary embodiment, such as... Figure 2 As shown, the method may include the following steps: In step S201, an object image and a video generation instruction text are obtained; the video generation instruction text is used to instruct the generation of a video with a coherent plot associated with the target object, and the object image is a reference image containing the target object.

[0042] In one specific embodiment, the video generation instruction text can be a description of an instruction to generate a video with a coherent plot associated with a target object. Specifically, the association of the target object with a video with a coherent plot can mean that the target object is present throughout the complete storyline of the video. The target object can include any type of real biological object, such as an adult or a pet cat. The target object can also include virtual biological objects, such as cartoon characters or cartoon animals. The target object can also include any type of non-biological object, such as a product. The embodiments of this application do not limit the specific type of the target object, and those skilled in the art can select it according to actual needs. For example, when the target object is a product, the video with a coherent plot associated with the target object can be a product recommendation advertisement; when the target object is a person, the video with a coherent plot associated with the target object can be a story film showcasing the person.

[0043] In one specific embodiment, the object image can be a reference image containing the target object. For example, when the target object is a product, the object image can be a product image or a product display image, etc.; when the target object is a person, the object image can be a person's headshot or a full-body photo of a person, etc.

[0044] In step S202, the object image and video generation instruction text are input into the video script generation model. During the process of generating the video script based on the object image and video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame, thereby obtaining script image feature information corresponding to multiple video frames.

[0045] In one specific embodiment, the video script generation model can be used to aggregate the video script information of each generated video frame into script image feature information based on the historical frame feature information corresponding to each video frame during the process of generating video scripts based on object images and video generation instruction text, thereby obtaining script image feature information corresponding to multiple video frames. Specifically, the model structure of the video script generation model can adopt a multimodal large model.

[0046] In a specific embodiment, the video script information for each video frame can be the descriptive text of the corresponding video frame. Specifically, the video script information can describe the video frame from descriptive dimensions such as scene composition, object actions, emotional tone, lighting requirements, color matching, prop elements, shot size, and camera angle. The script image feature information corresponding to each video frame can be used to characterize the global image features of the video script information in the image feature space. The script image feature information corresponding to each video frame can be feature information obtained by aggregating the historical frame feature information and the video script information corresponding to each video frame based on the image feature space. The historical frame feature information corresponding to each video frame can include: the video script information corresponding to the historical video frames before each video frame and the script image feature information corresponding to the historical video frames before each video frame. Illustratively, the script image feature information can be represented as a script image token sequence.

[0047] In an optional embodiment, the video generation instruction text can be used to instruct the generation of a video with a coherent plot associated with the target object. The video may include several target video frames. Accordingly, the above-mentioned inputting the object image and video generation instruction text into the video script generation model, and in the process of generating the video script based on the object image and video generation instruction text, aggregating the video script information of each generated video frame into script image feature information based on the historical frame feature information corresponding to each video frame, to obtain script image feature information corresponding to multiple video frames may include: inputting the object image and video generation instruction text into the video script generation model, and in the process of generating the video script based on the object image and video generation instruction text, aggregating the video script information of each generated video frame into script image feature information based on the historical frame feature information corresponding to each video frame, to obtain script image feature information corresponding to multiple video frames.

[0048] In a specific embodiment, such as Figure 3 As shown, in the process of generating video scripts based on object images and video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame. The resulting script image feature information corresponding to multiple video frames may include: In step S301, for the current video frame among multiple video frames, a video script is generated for the current video frame based on the object image, video generation instruction text and video script information corresponding to historical video frames, so as to obtain the video script information corresponding to the current video frame. In step S302, image features are aggregated based on the object image, video generation instruction text, video script information corresponding to historical video frames, script image feature information corresponding to historical video frames, and video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame. Among them, historical video frames are video frames that precede the current video frame from among multiple video frames.

[0049] Specifically, if the current video frame is the first video frame among multiple video frames, since there are no historical video frames for the first video frame, a script can be generated for the first video frame based on the object image and video generation instruction text to obtain the script information corresponding to the first video frame. Then, image feature aggregation is performed based on the object image, video generation instruction text, and the video script information corresponding to the first video frame to obtain the script image feature information corresponding to the first video frame. If the current video frame is not the first video frame among multiple video frames, a video script is generated for the current video frame based on the object image, video generation instruction text, and the video script information corresponding to historical video frames to obtain the video script information corresponding to the current video frame. Then, image feature aggregation is performed based on the object image, video generation instruction text, the video script information corresponding to historical video frames, the script image feature information corresponding to historical video frames, and the video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame.

[0050] For example, taking four video frames as an example, with the current video frame being the third video frame, the historical video frames are the first and second video frames. Accordingly, a video script can be generated for the third video frame based on the object image, video generation instruction text, video script information corresponding to the first and second video frames, to obtain the video script information corresponding to the third video frame. Then, image feature aggregation is performed based on the object image, video generation instruction text, video script information corresponding to the first and second video frames, script image feature information corresponding to the first and second video frames, and the video script information corresponding to the third video frame to obtain the script image feature information corresponding to the third video frame.

[0051] In the above embodiments, for the current video frame among multiple video frames, a video script is generated based on the object image, video generation instruction text, and video script information corresponding to historical video frames, to obtain the video script information corresponding to the current video frame. Then, image feature aggregation is performed based on the object image, video generation instruction text, video script information corresponding to historical video frames, script image feature information corresponding to historical video frames, and video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame. This can balance the causality of video script generation and the globality of image feature aggregation. While ensuring the narrative logic of the video script information, it allows the script image feature information to fully absorb global information, thereby providing macro-level guidance for visual consistency in subsequent image generation.

[0052] In a specific embodiment, the above-mentioned video script generation based on the object image, video generation instruction text, and video script information corresponding to historical video frames, to obtain the video script information corresponding to the current video frame, includes: In step S3011 (not shown in the figure), causal self-attention processing is performed on the object image, video generation instruction text, and video script information corresponding to historical video frames to obtain the script text feature information corresponding to the current video frame.

[0053] Specifically, script text feature information is used to characterize the feature information of the current video frame in the text feature space. The script text feature information corresponding to the current video frame can be obtained by performing causal self-attention processing on the object image, video generation instruction text, and video script information corresponding to historical video frames.

[0054] Specifically, causal self-attention processing can refer to self-attention processing based on causal masks, and correspondingly, such as Figure 4aAs shown, the first image feature information of the object image, the instruction text feature information of the video generation instruction text, and the script text feature information of the video script information corresponding to the historical video frames can be concatenated to obtain the first feature sequence. Then, the first feature sequence is subjected to self-attention processing based on causal masking to obtain the script text feature information corresponding to the current video frame.

[0055] In step S3012 (not shown in the figure), a video script is generated based on the script text feature information corresponding to the current video frame to obtain the video script information corresponding to the current video frame.

[0056] Specifically, text decoding can be performed based on the script text feature information corresponding to the current video frame to obtain the video script information corresponding to the current video frame.

[0057] The above-mentioned image feature aggregation based on the object image, video generation instruction text, video script information corresponding to historical video frames, script image feature information corresponding to historical video frames, and video script information corresponding to the current video frame yields the following script image feature information corresponding to the current video frame: In step S3021 (not shown in the figure), based on the pre-trained image query feature information, causal attention processing is performed on the object image, video generation instruction text, video script information corresponding to historical video frames, script image feature information corresponding to historical video frames, and video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame.

[0058] Specifically, pre-trained image query feature information can be pre-trained query feature information used for feature queries in the image feature space; such as Figure 4a As shown, in the process of causal attention processing, the pre-trained image query feature information can be used as the query vector to aggregate features of all historical contexts, including object images, video generation instruction text, video script information corresponding to historical video frames, script image feature information corresponding to historical video frames, and video script information corresponding to the current video frame, to obtain the script image feature information corresponding to the current video frame.

[0059] In a specific embodiment, such as Figure 4b As shown, the video script generation model can adopt a multimodal large model structure. The video script generation model may include: a first image encoder and a large language model. The large language model may include: a text encoding layer, an attention processing layer, and a text decoding layer. The above-mentioned causal self-attention processing of the object image, video generation instruction text, and video script information corresponding to historical video frames, to obtain the script text feature information corresponding to the current video frame, may include: The object image is input into the first image encoder for image feature extraction to obtain the first image feature information; the video generation instruction text is input into the text encoding layer for text encoding processing to obtain the instruction text feature information; the first image feature information of the object image, the instruction text feature information of the video generation instruction text, and the script text feature information corresponding to the historical video frames are concatenated to obtain the first feature sequence; the first feature sequence is input into the attention processing layer for self-attention processing based on causal masking to obtain the script text feature information corresponding to the current video frame. The above-mentioned video script generation based on the script text feature information corresponding to the current video frame can include: Input the script text feature information corresponding to the current video frame into the text decoding layer for text decoding processing to obtain the video script information corresponding to the current video frame. The aforementioned pre-trained image query feature information, through causal attention processing of the object image, video generation instruction text, video script information corresponding to historical video frames, script image feature information corresponding to historical video frames, and video script information corresponding to the current video frame, yields the script image feature information corresponding to the current video frame, which may include: The first image feature information of the object image, the instruction text feature information of the video generation instruction text, the script text feature information corresponding to the historical video frames, the script image feature information corresponding to the historical video frames, and the script text feature information corresponding to the current video frame are concatenated to obtain the second feature sequence. The pre-trained image query feature information and the second feature sequence are input into the attention processing layer. Using the pre-trained image query feature information as the query vector, causal attention learning is performed on the second feature sequence to obtain the script image feature information corresponding to the current video frame.

[0060] In the above embodiments, causal self-attention processing is performed on the object image, video generation instruction text, and video script information corresponding to historical video frames to obtain script text feature information corresponding to the current video frame. Then, video script is generated based on the script text feature information corresponding to the current video frame to obtain video script information corresponding to the current video frame. Based on pre-trained image query feature information, causal attention processing is performed on the object image, video generation instruction text, video script information corresponding to historical video frames, script image feature information corresponding to historical video frames, and video script information corresponding to the current video frame to obtain script image feature information corresponding to the current video frame. This can effectively balance the causality of video script generation and the globality of image feature aggregation.

[0061] In step S203, the object image and script image feature information corresponding to multiple video frames are input into the video generation model to generate video, thereby obtaining a target generated video containing generated images corresponding to multiple video frames.

[0062] In one specific embodiment, the video generation model can be used to generate video based on the feature information of the object image and the script images corresponding to multiple video frames, resulting in a target generated video containing generated images corresponding to multiple video frames. Specifically, the video generation model can employ a text-to-image model, such as a diffusion model or a DIT (DiffusionTransformer).

[0063] Specifically, the target generated video can be a target video generated based on the object image and the video generation instruction text; the number of frames of the generated images contained in the target generated video is consistent with the number of target frames indicated by the video generation instruction text.

[0064] In one specific embodiment, the video generation model includes: an image generation model, such as... Figure 5 As shown, the above-mentioned input of the object image and script image feature information corresponding to multiple video frames into the video generation model for video generation results in a target generated video containing generated images corresponding to multiple video frames, including: In step S501, the object image and the script image feature information corresponding to multiple video frames are input into the image generation model. Based on the object image, the script image feature information corresponding to the current video frame in multiple video frames, and the generated images corresponding to historical video frames, the generated image corresponding to the current video frame is generated. In step S502, a target generated video is obtained based on the generated images corresponding to multiple video frames; wherein, historical video frames are video frames that precede the current video frame among the multiple video frames.

[0065] Specifically, if the current video frame is the first video frame among multiple video frames, since there are no historical video frames for the first video frame, image generation can be performed based on the object image and the script image feature information corresponding to the first video frame to obtain the generated image corresponding to the first video frame; if the current video frame is not the first video frame among multiple video frames, image generation can be performed based on the object image, the script image feature information corresponding to the current video frame, and the generated images corresponding to historical video frames to obtain the generated image corresponding to the current video frame.

[0066] For example, taking four video frames as a whole, with the current video frame being the third video frame, the historical video frames are the first and second video frames. Accordingly, an image can be generated based on the object image, the script image feature information corresponding to the third video frame, the generated image corresponding to the first video frame, and the generated image corresponding to the second video frame to obtain the generated image corresponding to the third video frame.

[0067] In the above embodiments, the object image and script image feature information corresponding to multiple video frames are input into the image generation model. Based on the object image, the script image feature information corresponding to the current video frame in multiple video frames, and the generated images corresponding to historical video frames, image generation is performed to obtain the generated image corresponding to the current video frame. By using the generated images corresponding to historical video frames as constraints for the current video frame, precise control of image generation for the current video frame and pixel-level visual consistency maintenance are achieved.

[0068] In a specific embodiment, the above-mentioned image generation based on the object image, the script image feature information corresponding to the current video frame in multiple video frames, and the generated images corresponding to historical video frames, to obtain the generated image corresponding to the current video frame includes: In step S5011 (not shown in the figure), based on the distance between the historical video frame and the current video frame, the image feature information of the generated image corresponding to the historical video frame is attenuated and compressed to obtain the historical image feature information corresponding to the current video frame.

[0069] Specifically, the distance between historical video frames and the current video frame can be the difference in the number of frames between them. Based on the distance between historical video frames and the current video frame, the image feature information of the generated image corresponding to the historical video frame is attenuated and compressed, so that the historical image feature information of the historical video frame that is further away from the current video frame retains more macroscopic information.

[0070] In an optional embodiment, the above-mentioned attenuation compression of the image feature information of the generated image corresponding to the historical video frame based on the distance between the historical video frame and the current video frame, to obtain the historical image feature information corresponding to the current video frame, includes: In step S1, after the generated image corresponding to the previous video frame of the current video frame is generated, the historical image feature information corresponding to the previous video frame is extracted from the feature cache library.

[0071] Specifically, the feature cache library is used to cache the image feature information of the generated images corresponding to the generated video frames.

[0072] In step S2, the historical image feature information corresponding to the previous video frame and the image feature information of the generated image corresponding to the previous video frame are compressed to obtain the historical image feature information corresponding to the current video frame.

[0073] Specifically, the historical image feature information corresponding to the current video frame can be the image feature information obtained by compressing the historical image feature information corresponding to the previous video frame and the image feature information of the generated image corresponding to the previous video frame. The historical image feature information corresponding to the current video frame is used to characterize the compressed feature information of the historical generated images before the current video frame.

[0074] In an optional embodiment, the historical image feature information corresponding to the previous video frame and the image feature information of the generated image corresponding to the previous video frame can be compressed using average pooling to obtain the historical image feature information corresponding to the current video frame. This reduces computational costs while allowing greater focus on recently generated images. Specifically, average pooling compression can introduce a decay factor. This causes the feature length of the generated image feature information corresponding to the historical video frame (i.e., the nk-th video frame) with a distance of k from the current video frame (the nth video frame) to change from the original length I to... Correspondingly, the sum of the feature lengths of the image feature information of the generated images corresponding to all historical video frames of the current video frame can be as follows:

[0075] As can be seen from the above, by using average pooling compression, the feature length of the image feature information of the generated image corresponding to the historical video frame decreases geometrically, ensuring that the sum of the feature lengths of the image feature information of the generated image corresponding to the historical video frame is bounded, thereby achieving effective compression of image features.

[0076] In step S3, the historical image feature information corresponding to the previous video frame is updated to the historical image feature information corresponding to the current video frame in the feature cache.

[0077] In the above embodiments, after the generated image corresponding to the previous video frame is generated, the historical image feature information corresponding to the previous video frame is extracted from the feature cache library. The historical image feature information corresponding to the previous video frame and the image feature information of the generated image corresponding to the previous video frame are compressed to obtain the historical image feature information corresponding to the current video frame. In the feature cache library, the historical image feature information corresponding to the previous video frame is updated to the historical image feature information corresponding to the current video frame. This can reduce the computational cost of the image generation process while paying more attention to recently generated images and maintaining good visual consistency.

[0078] In step S5012 (not shown in the figure), an image is generated based on the object image feature information of the object image, the script image feature information corresponding to the current video frame, and the historical image feature information to obtain the generated image corresponding to the current video frame.

[0079] Specifically, the object image feature information, the script image feature information corresponding to the current video frame, and the historical image feature information are concatenated to obtain concatenated features. Based on the concatenated features, an image is generated to obtain the generated image corresponding to the current video frame.

[0080] In an optional embodiment, the number of historical video frames involved in the historical image feature information can be controlled, that is, only the image feature information corresponding to the generated images of historical video frames k frames before the current video frame is retained, thereby further controlling the computational cost of the image generation process.

[0081] In the above embodiments, based on the distance between historical video frames and the current video frame, the image feature information of the generated image corresponding to the historical video frame is attenuated and compressed to obtain the historical image feature information corresponding to the current video frame. Then, based on the object image feature information of the object image, the script image feature information corresponding to the current video frame, and the historical image feature information, an image is generated to obtain the generated image corresponding to the current video frame. By using the object image feature information of the object image, the script image feature information corresponding to the current video frame, and the historical image feature information as constraints for the current video frame, precise control of the image generation of the current video frame and pixel-level visual consistency maintenance are achieved.

[0082] In the above embodiments, in the long video generation scenario, an object image and video generation instruction text are obtained. The video generation instruction text is used to instruct the generation of a video with a coherent plot associated with the target object. The object image is a reference image containing the target object. Then, the object image and video generation instruction text are input into the video script generation model. During the video script generation process based on the object image and video generation instruction text, based on the historical frame feature information corresponding to each video frame, the video script information of each generated video frame is aggregated into script image feature information, resulting in script image feature information corresponding to multiple video frames. This allows for the automatic planning of a coherent video script, and the aggregation of each... The historical frame feature information corresponding to each video frame and the text features of the video script for each video frame are aggregated into script image feature information at the image dimension. This improves the fusion effect of multimodal information and the semantic coherence of script image feature information, providing macro-level guidance for visual consistency in subsequent image generation. Then, the object image and the script image feature information corresponding to multiple video frames are input into the video generation model to generate the video, resulting in a target generated video containing generated images corresponding to multiple video frames. This enables fine-grained control of the long video generation process, maintains the visual consistency and plot coherence of multiple frames in the long video, effectively improves the generation quality of long videos, and thus enhances the user's video production experience.

[0083] In one specific embodiment, the video generation model includes: a feature alignment model and an image generation model, such as... Figure 6 As shown, the image generation model may include: a second image encoder, a feature cache library, a feature fusion layer, and an image generation layer.

[0084] In a specific embodiment, such as Figure 7As shown, the above-mentioned object image and video generation instruction text are input into the video script generation model. During the video script generation process based on the object image and video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame. The resulting script image feature information corresponding to multiple video frames includes: In step S701, the object image and video generation instruction text are input into the video script generation model. During the process of generating the video script based on the first image feature information of the object image and the video generation instruction text, the video script information of each generated video frame and the historical frame feature information corresponding to each video frame are aggregated into script image feature information based on the pre-trained image query feature information, thereby obtaining script image feature information corresponding to multiple video frames.

[0085] For details regarding step S701, please refer to the detailed content of steps S3011 to S3021 mentioned above, which will not be repeated here.

[0086] The above-mentioned input of the object image and script image feature information corresponding to multiple video frames into the video generation model for video generation results in a target generated video containing generated images corresponding to multiple video frames, including: In step S702, the second image feature information of the object image is obtained.

[0087] Specifically, the first image feature information of the object image can be obtained by image feature encoding by the first image encoder, and the second image feature information of the object image can be obtained by image feature encoding by the second image encoder. The first image encoder and the second image encoder can adopt the same encoder structure or different encoder structures. This application does not impose any special restrictions on this. For illustration, the second image encoder can adopt VAE.

[0088] In step S703, the script image feature information corresponding to multiple video frames is input into the feature alignment model for feature alignment to obtain the aligned image feature information corresponding to multiple video frames.

[0089] Specifically, the script image feature information corresponding to each video frame in multiple video frames can be input into the feature alignment model for feature alignment to obtain the aligned image feature information corresponding to each video frame. In a specific embodiment, the feature alignment model can be any artificial capability model with feature semantic alignment capability in the prior art, and this application does not impose any special restrictions on it.

[0090] In step S704, the second image feature information and the aligned image feature information corresponding to multiple video frames are input into the image generation model to generate images, thereby obtaining generated images corresponding to multiple video frames.

[0091] In step S705, the target generated video is obtained based on the generated images corresponding to multiple video frames.

[0092] In the above embodiments, the feature alignment model is used to align the script image feature information output by the video script generation model with the semantic space of the image generation model, so as to maintain semantic consistency.

[0093] In a specific embodiment, such as Figure 8 As shown, the above method also includes: In step S801, the sample instruction text, the sample object image corresponding to the sample instruction text, the preset video script information of multiple video frames corresponding to the sample instruction text, and the preset generated image of multiple video frames corresponding to the preset video script information are obtained; the sample instruction text is used to instruct the generation of a video with a coherent plot associated with the sample object, and the sample object image is a sample reference image containing the sample object.

[0094] Specifically, the preset video script information for multiple video frames corresponding to the sample instruction text can be video script tags for multiple video frames; the preset generated images for multiple video frames corresponding to the preset video script information can be generated image tags for multiple video frames.

[0095] In one specific embodiment, the above method further includes: In step S8011 (not shown in the figure), the sample object description information corresponding to the sample object is obtained.

[0096] Specifically, the sample object description information can be used to describe the object characteristics of the sample object. For example, the sample object description information corresponding to the sample object can be extracted from the sample object maintenance platform. For instance, taking the sample object as an e-commerce product, the object maintenance platform can be the e-commerce platform, and the sample product description information corresponding to the sample e-commerce product can be obtained from the e-commerce platform.

[0097] In step S8012 (not shown in the figure), a reference script is generated based on the sample object description information to obtain the reference script information.

[0098] Specifically, the sample object description information can be input into the pre-trained first large language model to generate a reference script, thus obtaining reference script information. Illustratively, the pre-trained large language model can be any existing artificial intelligence model with script generation capabilities; this application does not impose any particular limitations on it.

[0099] In step S8013 (not shown in the figure), a reference image is generated based on the reference script information to obtain a sample object image.

[0100] Specifically, the reference script information can be input into the pre-trained first raw image model to generate reference images and obtain sample object images.

[0101] In step S8014 (not shown in the figure), a video script is generated based on the sample object image and sample object description information to obtain preset video script information for multiple video frames.

[0102] Specifically, the sample object image and sample object description information can be input into the pre-trained second language model to generate video scripts, resulting in preset video script information for multiple video frames. Illustratively, the first and second language models can use the same model architecture or different model structures; this application does not impose any particular limitations on this.

[0103] In step S8015 (not shown in the figure), image generation is performed based on the sample object image and the preset video script information of multiple video frames to obtain the preset generated images of multiple video frames.

[0104] Specifically, the sample object image and preset video script information of multiple video frames can be input into the pre-trained second text-based image model to generate images, resulting in preset generated images of multiple video frames. Illustratively, the first and second text-based image models can use the same model architecture or different model structures; this application does not impose any particular limitations on this.

[0105] In step S8016 (not shown in the figure), instruction settings are performed based on preset video script information and preset generated images to obtain sample instruction text.

[0106] Specifically, preset video script information and preset generated images can be input into the large language model for instruction settings, or instructions can be set manually based on the preset video script information and preset generated images.

[0107] The above embodiments can effectively improve the construction quality of long video training data, effectively maintain the cross-frame consistency of sample video frames, and thus improve the training effect of subsequent video script generation models and video generation models.

[0108] In step S802, based on the sample image feature information corresponding to the sample object image, the sample instruction text, and the preset video script information of multiple video frames, the video script generation model to be trained is trained to obtain the video script generation model.

[0109] Specifically, in the first training phase, the large language model in the video script generation model is trained and the first image encoder in the video script generation model is frozen. In a specific embodiment, the training objective of the first training phase can be expressed as the following formula: , in, This represents the predicted video script information corresponding to the j-th video frame in a script sequence of video length NT. Accordingly, the training objective of the first training phase can be to minimize the NTP (Next-Token Prediction) loss of the predicted video script information. Optionally, the loss function of the first training phase can be the cross-entropy loss function.

[0110] In step S803, based on the preset video script information and the preset generated image of each video frame, feature space alignment training is performed on the image query feature information to be trained and the feature alignment model to be trained to obtain the pre-trained image query feature information and feature alignment model.

[0111] Specifically, in the second training phase, feature space alignment training is performed on the image query feature information and feature alignment model to be trained. Since this feature alignment model is used to connect the video script generation model and the image generation model, it aligns the semantic space of the image query feature information with that of the subsequent image generation model. In a specific embodiment, text-image matching data can be constructed first based on the preset video script information and the preset generated image of each video frame. Based on the text-image matching data, the image query feature information and the feature alignment model to be trained are pre-trained. Then, the preset video script information and the preset generated image of each video frame are interleaved to obtain an interleaved text-image sequence. The image query feature information to be trained is trained to learn the global feature aggregation capability for the image feature space through the interleaved text-image sequence. Accordingly, the training objective of the second training phase can be expressed as the following formula: , in, Represents a latent variable with noise. It is a preset of the real image features of the generated image. It's Gaussian noise. Represents the predicted vector field, The image query feature information for the nth video frame is used to train the image. Optionally, the loss function in the second training stage can be the stream matching loss function.

[0112] In step S804, based on the sample object image, the sample script image feature information corresponding to the preset video script information, and the preset generated images of multiple video frames, the image generation model to be trained is trained to obtain the image generation model.

[0113] Specifically, in the third training phase, the image feature information corresponding to the sample object image, the sample script image feature information corresponding to the preset video script information of the nth video frame, and the image feature information corresponding to the preset generated images of the historical video frames of the nth video frame are concatenated to obtain the sample image generation conditions corresponding to the nth video frame. Based on the sample image generation conditions, the image generation model to be trained is trained to generate visually consistent images. In a specific embodiment, the training objective of the third training phase can be expressed as the following formula: , in, This represents the conditions for generating the sample image corresponding to the nth video frame. Optionally, the loss function in the third training phase can be the stream matching loss function.

[0114] In the above embodiments, a three-stage model training strategy is adopted, in which video script generation training is performed in the first stage, feature space alignment training is performed in the second stage, and fine-grained consistent image generation training is performed in the third stage. This strategy can maintain good training results under the condition of limited computing resources.

[0115] For example, taking the long video generation scenario as an e-commerce advertising video generation scenario, the target object can be a new hiking backpack, the object image can be a product image of the new hiking backpack, and the video generation instruction text can be "Generate a storyboard showing the use of this hiking backpack in an outdoor hiking scene, containing 4 video frames"; correspondingly, such as Figure 9As shown, the object image and video generation instruction text are input into the video script generation model to generate the video script. The video script generation model outputs the video script information corresponding to the first video frame: "A man stands at the foot of a mountain with a backpack, looking up at the peak.", the script image feature information q1 corresponding to the first video frame, the video script information corresponding to the second video frame: "A man walks on a rugged mountain path, with his backpack close to his back.", the script image feature information q2 corresponding to the second video frame, the video script information corresponding to the third video frame, the script image feature information q3 corresponding to the third video frame, the video script information corresponding to the fourth video frame, and the script image feature information q4 corresponding to the fourth video frame. Then, q1 and the image feature information fcond corresponding to the product image are input into the video generation model to generate the image, obtaining the generated image corresponding to the first video frame. The image feature information f1 of the generated image corresponding to the first video frame is stored in the dynamic cache library, and q2, fcond, and f1 are used as the image generation conditions for the second video frame, input into the video generation model to generate the image, obtaining the generated image corresponding to the second video frame. At this time, the model... The model acquires both the semantics of the video script (walking) and the visual features of the generated image corresponding to the first video frame (the man's clothing and backpack style), ensuring that the second video frame maintains character consistency with the first. Then, the image feature information f2 of the generated image corresponding to the second video frame is stored in a dynamic cache, f1 is compressed and attenuated, and q3, fcond, f2, and the first-compressed f1 are used as the image generation conditions for the third video frame. This data is input into the video generation model to generate the image corresponding to the third video frame. Next, the image feature information f3 of the generated image corresponding to the third video frame is stored in a dynamic cache, f2 is compressed and attenuated once, and f1 is compressed and attenuated a second time. q4, fcond, f3, the first-compressed f2, and the second-compressed f1 are used as the image generation conditions for the fourth video frame. This data is input into the video generation model to generate the image corresponding to the fourth video frame. Based on the generated images corresponding to the four video frames (a set of four keyframe images with a unified visual style, consistent character, and conforming to the hiking narrative), an e-commerce advertising video for the new hiking backpack is obtained.

[0116] As can be seen from the technical solutions provided in the embodiments of this specification above, in a long video generation scenario, an object image and video generation instruction text are obtained. The video generation instruction text is used to instruct the generation of a video with a coherent plot associated with the target object. The object image is a reference image containing the target object. Then, the object image and video generation instruction text are input into a video script generation model. During the video script generation process based on the object image and video generation instruction text, based on the historical frame feature information corresponding to each video frame, the video script information of each generated video frame is aggregated into script image feature information, resulting in script image feature information corresponding to multiple video frames. Based on the automatic planning of a coherent video script, the historical frame feature information corresponding to each video frame and the text features of the video script of each video frame are aggregated into an image-dimensional script. Image feature information enhances the fusion effect of multimodal information and the semantic coherence of script image feature information, providing macro-level guidance for visual consistency in subsequent image generation. Then, the object image and script image feature information corresponding to multiple video frames are input into the video generation model for video generation, resulting in a target generated video containing generated images corresponding to multiple video frames. This enables fine-grained control over the long video generation process, maintaining visual consistency and narrative coherence across multiple frames in a long video, effectively improving the quality of long video generation and thus enhancing the user's video production experience. Furthermore, a three-stage model training strategy—training video script generation in the first stage, feature space alignment in the second stage, and fine-grained consistent image generation in the third stage—maintains good training results even with limited computational resources.

[0117] Figure 10 This is a block diagram of a video processing apparatus according to an exemplary embodiment. (Refer to...) Figure 10 The device includes: The instruction acquisition module 1010 is configured to acquire object images and video generation instruction text; the video generation instruction text is used to instruct the generation of a video with a coherent plot associated with the target object, and the object image is a reference image containing the target object; The video script generation module 1020 is configured to input object images and video generation instruction text into the video script generation model. During the process of generating video scripts based on object images and video generation instruction text, based on the historical frame feature information corresponding to each video frame, the video script information of each generated video frame is aggregated into script image feature information to obtain script image feature information corresponding to multiple video frames. The video generation module 1030 is configured to input the object image and script image feature information corresponding to multiple video frames into the video generation model to generate a target generated video containing generated images corresponding to multiple video frames.

[0118] In an optional embodiment, the video script generation module 1020 includes: The script generation unit is configured to perform video script generation for the current video frame among multiple video frames, based on the object image, video generation instruction text and video script information corresponding to historical video frames, to obtain the video script information corresponding to the current video frame. The image feature aggregation unit is configured to perform image feature aggregation based on the object image, video generation instruction text, video script information corresponding to historical video frames, script image feature information corresponding to historical video frames, and video script information corresponding to the current video frame, to obtain the script image feature information corresponding to the current video frame. Among them, historical video frames are video frames that precede the current video frame from among multiple video frames.

[0119] In an optional embodiment, the script generation unit is further configured to perform causal self-attention processing on the object image, video generation instruction text, and video script information corresponding to historical video frames to obtain script text feature information corresponding to the current video frame; and generate a video script based on the script text feature information corresponding to the current video frame to obtain video script information corresponding to the current video frame. The image feature aggregation unit is also configured to perform pre-trained image query feature information, and to perform causal attention processing on the object image, video generation instruction text, video script information corresponding to historical video frames, script image feature information corresponding to historical video frames, and video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame.

[0120] In an optional embodiment, the video generation model includes: an image generation model, and the video generation module 1030 includes: The image generation unit is configured to input the object image and the script image feature information corresponding to multiple video frames into the image generation model, and generate an image based on the object image, the script image feature information corresponding to the current video frame in the multiple video frames and the generated images corresponding to the historical video frames, to obtain the generated image corresponding to the current video frame. The video synthesis unit is configured to generate images based on multiple video frames to obtain the target generated video. Among them, historical video frames are video frames that precede the current video frame from among multiple video frames.

[0121] In an optional embodiment, the image generation unit is further configured to perform attenuation compression on the image feature information of the generated image corresponding to the historical video frame based on the spacing between the historical video frame and the current video frame, to obtain the historical image feature information corresponding to the current video frame; and to perform image generation based on the object image feature information of the object image, the script image feature information corresponding to the current video frame, and the historical image feature information, to obtain the generated image corresponding to the current video frame.

[0122] In an optional embodiment, the image generation unit is further configured to, after the generation of the generated image corresponding to the previous video frame is completed, extract historical image feature information corresponding to the previous video frame from the feature cache; compress the historical image feature information corresponding to the previous video frame and the image feature information of the generated image corresponding to the previous video frame to obtain historical image feature information corresponding to the current video frame; and update the historical image feature information corresponding to the previous video frame to the historical image feature information corresponding to the current video frame in the feature cache.

[0123] In an optional embodiment, the video generation model described above includes: a feature alignment model and an image generation model; The aforementioned video script generation module 1020 is also configured to input the object image and video generation instruction text into the video script generation model. During the process of generating the video script based on the first image feature information of the object image and the video generation instruction text, based on the pre-trained image query feature information, the video script information of each generated video frame and the historical frame feature information corresponding to each video frame are aggregated into script image feature information to obtain script image feature information corresponding to multiple video frames. The aforementioned video generation module 1030 is further configured to: acquire second image feature information of the target image; input script image feature information corresponding to multiple video frames into a feature alignment model for feature alignment to obtain aligned image feature information corresponding to multiple video frames; input the second image feature information and the aligned image feature information corresponding to multiple video frames into an image generation model for image generation to obtain generated images corresponding to multiple video frames; and obtain the target generated video based on the generated images corresponding to multiple video frames.

[0124] In an optional embodiment, the above-described apparatus further includes: The sample acquisition module is configured to acquire sample instruction text, sample object image corresponding to the sample instruction text, preset video script information of multiple video frames corresponding to the sample instruction text, and preset generated image of multiple video frames corresponding to the preset video script information; the sample instruction text is used to instruct the generation of a video with a coherent plot associated with the sample object, and the sample object image is a sample reference image containing the sample object. The first training module is configured to execute preset video script information based on sample image feature information, sample instruction text and multiple video frames corresponding to the sample object image, to train the video script generation model to be trained, and obtain the video script generation model. The second training module is configured to execute a feature space alignment training on the image query feature information and the feature alignment model to be trained based on the preset video script information and the preset generated image of each video frame, so as to obtain the pre-trained image query feature information and the feature alignment model. The third training module is configured to execute image generation training on the image generation model to be trained based on the sample object image, the sample script image feature information corresponding to the preset video script information, and the preset generated images of multiple video frames, so as to obtain the image generation model.

[0125] In an optional embodiment, the sample acquisition module is also configured to perform: Retrieve the sample object description information corresponding to the sample object; Reference script information is obtained by generating a reference script based on the sample object description information. Based on the reference script information, a reference image is generated to obtain the sample object image; Video scripts are generated based on sample object images and sample object description information to obtain preset video script information for multiple video frames. Image generation is performed based on sample object images and preset video script information of multiple video frames to obtain preset generated images of multiple video frames. Based on the preset video script information and preset generated image, the instruction settings are performed to obtain sample instruction text.

[0126] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0127] Figure 11 This is a block diagram illustrating an electronic device for video processing according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 11 As shown, it may include RF (Radio Frequency) circuitry 1110, a memory 1120 including one or more computer-readable storage media, an input unit 1130, a display unit 1140, a sensor 1150, an audio circuit 1160, a WiFi (Wireless Fidelity) module 1170, a processor 1180 including one or more processing cores, and a power supply 1190, among other components. Those skilled in the art will understand that... Figure 11 The terminal structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: RF circuit 1110 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 1180 for processing; additionally, it transmits uplink data to the base station. Typically, RF circuit 1110 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, an LNA (Low Noise Amplifier), a duplexer, etc. Furthermore, RF circuit 1110 can also communicate wirelessly with networks and other terminals. Wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System for Mobile communication), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, SMS (Short Messaging Service), etc.

[0128] The memory 1120 can be used to store software programs and modules. The processor 1180 executes various functional applications and data processing by running the software programs and modules stored in the memory 1120. The memory 1120 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for the functions, etc.; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 1120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 1120 may also include a memory controller to provide access to the memory 1120 by the processor 1180 and the input unit 1130.

[0129] Input unit 1130 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, input unit 1130 may include touch-sensitive surface 1131 and other input devices 1132. Touch-sensitive surface 1131, also known as a touch display screen or touchpad, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 1131), and drive corresponding connection devices according to a pre-set program. Optionally, touch-sensitive surface 1131 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 1180, and can receive and execute commands from processor 1180. In addition, the touch-sensitive surface 1131 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. Besides the touch-sensitive surface 1131, the input unit 1130 may also include other input devices 1132. Specifically, other input devices 1132 may include, but are not limited to, one or more of the following: a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick.

[0130] Display unit 1140 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the terminal. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 1140 may include display panel 1141, which may optionally be configured as an LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or similar display panel. Further, touch-sensitive surface 1131 may cover display panel 1141. When touch-sensitive surface 1131 detects a touch operation on or near it, it transmits the information to processor 1180 to determine the type of touch event. Subsequently, processor 1180 provides corresponding visual output on display panel 1141 according to the type of touch event. Touch-sensitive surface 1131 and display panel 1141 can be two independent components to implement input and output functions. However, in some embodiments, touch-sensitive surface 1131 and display panel 1141 can be integrated to achieve input and output functions.

[0131] The terminal may also include at least one sensor 1150, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 1141 according to the ambient light level, and the proximity sensor can turn off the display panel 1141 and / or the backlight when the terminal is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that identify the terminal's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, tapping), etc. Other sensors that may be configured on the terminal, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0132] Audio circuitry 1160, speaker 1161, and microphone 1162 provide an audio interface between the user and the terminal. Audio circuitry 1160 converts received audio data into electrical signals, which are then transmitted to speaker 1161, where they are converted into sound signals for output. Conversely, microphone 1162 converts collected sound signals into electrical signals, which are received by audio circuitry 1160, converted back into audio data, and then processed by processor 1180 before being transmitted via RF circuitry 1110 to, for example, another terminal, or output to memory 1120 for further processing. Audio circuitry 1160 may also include an earphone jack to facilitate communication between a peripheral headset and the terminal.

[0133] WiFi is a short-range wireless transmission technology. This terminal, through the WiFi module 1170, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 11 WiFi module 1170 is shown, but it is understood that it is not an essential component of the terminal and can be omitted as needed without changing the nature of the invention.

[0134] The processor 1180 is the control center of the terminal, connecting various parts of the terminal via various interfaces and lines. It executes software programs and / or modules stored in the memory 1120, and calls data stored in the memory 1120, to perform various functions and process data, thereby enabling overall monitoring of the terminal. Optionally, the processor 1180 may include one or more processing cores; preferably, the processor 1180 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1180.

[0135] The terminal also includes a power supply 1190 (such as a battery) to power various components. Preferably, the power supply can be logically connected to the processor 1180 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 1190 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0136] Although not shown, the terminal may also include a camera, Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the terminal is a touch screen display, and the terminal also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors of the instructions in the method embodiment of the present invention.

[0137] In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the video processing method as described in the embodiments of this disclosure.

[0138] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the video processing method of the present disclosure embodiments.

[0139] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the video processing method of the present disclosure embodiments.

[0140] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0141] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0142] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A video processing method, characterized in that, include: Obtain the object image and video to generate instruction text; The video generation instruction text is used to instruct the generation of a video with a coherent plot associated with the target object, and the object image is a reference image containing the target object; The object image and the video generation instruction text are input into the video script generation model. During the process of generating the video script based on the object image and the video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame, so as to obtain script image feature information corresponding to multiple video frames. The object image and the script image feature information corresponding to the multiple video frames are input into the video generation model to generate a target generated video containing the generated images corresponding to the multiple video frames.

2. The video processing method according to claim 1, characterized in that, In the process of generating a video script based on the object image and the video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame, resulting in script image feature information corresponding to multiple video frames, including: For the current video frame among the plurality of video frames, a video script is generated for the current video frame based on the object image, the video generation instruction text, and the video script information corresponding to the historical video frames, to obtain the video script information corresponding to the current video frame; Image feature aggregation is performed based on the object image, the video generation instruction text, the video script information corresponding to the historical video frames, the script image feature information corresponding to the historical video frames, and the video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame; The historical video frames are the video frames that precede the current video frame among the plurality of video frames.

3. The video processing method according to claim 2, characterized in that, The step of generating a video script for the current video frame based on the object image, the video generation instruction text, and the video script information corresponding to historical video frames, to obtain the video script information corresponding to the current video frame, includes: Causal self-attention processing is performed on the object image, the video generation instruction text, and the video script information corresponding to the historical video frames to obtain the script text feature information corresponding to the current video frame; Based on the script text feature information corresponding to the current video frame, a video script is generated to obtain the video script information corresponding to the current video frame. The step of aggregating image features based on the object image, the video generation instruction text, the video script information corresponding to the historical video frames, the script image feature information corresponding to the historical video frames, and the video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame includes: Based on pre-trained image query feature information, causal attention processing is performed on the object image, the video generation instruction text, the video script information corresponding to the historical video frames, the script image feature information corresponding to the historical video frames, and the video script information corresponding to the current video frame to obtain the script image feature information corresponding to the current video frame.

4. The video processing method according to claim 1, characterized in that, The video generation model includes an image generation model, wherein inputting the object image and script image feature information corresponding to the plurality of video frames into the video generation model to generate a target generated video containing the generated images corresponding to the plurality of video frames includes: The object image and the script image feature information corresponding to the plurality of video frames are input into the image generation model. Based on the object image, the script image feature information corresponding to the current video frame in the plurality of video frames, and the generated images corresponding to the historical video frames, an image is generated to obtain the generated image corresponding to the current video frame. The target generated video is obtained based on the generated images corresponding to the multiple video frames; The historical video frames are the video frames that precede the current video frame among the plurality of video frames.

5. The video processing method according to claim 4, characterized in that, The step of generating an image based on the object image, the script image feature information corresponding to the current video frame in the plurality of video frames, and the generated images corresponding to historical video frames to obtain the generated image corresponding to the current video frame includes: Based on the distance between the historical video frame and the current video frame, the image feature information of the generated image corresponding to the historical video frame is attenuated and compressed to obtain the historical image feature information corresponding to the current video frame. Image generation is performed based on the object image feature information of the object image, the script image feature information corresponding to the current video frame, and the historical image feature information to obtain the generated image corresponding to the current video frame.

6. The video processing method according to claim 5, characterized in that, The step of attenuating and compressing the image feature information of the generated image corresponding to the historical video frame based on the distance between the historical video frame and the current video frame to obtain the historical image feature information corresponding to the current video frame includes: After the generated image corresponding to the previous video frame of the current video frame is generated, the historical image feature information corresponding to the previous video frame is extracted from the feature cache library; The historical image feature information corresponding to the previous video frame and the image feature information of the generated image corresponding to the previous video frame are compressed to obtain the historical image feature information corresponding to the current video frame. In the feature cache library, the historical image feature information corresponding to the previous video frame is updated to the historical image feature information corresponding to the current video frame.

7. The video processing method according to any one of claims 1 to 6, characterized in that, The video generation model includes a feature alignment model and an image generation model. The object image and the video generation instruction text are input into the video script generation model. During the video script generation process based on the object image and the video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame. This results in script image feature information corresponding to multiple video frames, including: The object image and the video generation instruction text are input into the video script generation model. During the process of generating the video script based on the first image feature information of the object image and the video generation instruction text, the video script information of each generated video frame and the historical frame feature information corresponding to each video frame are aggregated into script image feature information based on the pre-trained image query feature information to obtain the script image feature information corresponding to the multiple video frames. The step of inputting the object image and the script image feature information corresponding to the plurality of video frames into the video generation model to generate a target generated video containing the generated images corresponding to the plurality of video frames includes: Obtain the second image feature information of the object image; The script image feature information corresponding to the multiple video frames is input into the feature alignment model for feature alignment, thereby obtaining the aligned image feature information corresponding to the multiple video frames. The second image feature information and the aligned image feature information corresponding to the plurality of video frames are input into the image generation model to generate an image, thereby obtaining the generated image corresponding to the plurality of video frames; The target generated video is obtained based on the generated images corresponding to the multiple video frames.

8. The video processing method according to claim 7, characterized in that, The method further includes: The sample instruction text, the sample object image corresponding to the sample instruction text, the preset video script information of multiple video frames corresponding to the sample instruction text, and the preset generated image of multiple video frames corresponding to the preset video script information are obtained; the sample instruction text is used to instruct the generation of a video with a coherent plot associated with the sample object, and the sample object image is a sample reference image containing the sample object. Based on the sample image feature information corresponding to the sample object image, the sample instruction text, and the preset video script information of the multiple video frames, the video script generation model to be trained is trained to obtain the video script generation model. Based on the preset video script information and the preset generated image of each video frame, feature space alignment training is performed on the image query feature information to be trained and the feature alignment model to be trained to obtain the pre-trained image query feature information and the feature alignment model. Based on the sample object image, the sample script image feature information corresponding to the preset video script information, and the preset generated images of the multiple video frames, the image generation model to be trained is trained to obtain the image generation model.

9. The video processing method according to claim 8, characterized in that, The method further includes: Obtain the sample object description information corresponding to the sample object; Based on the sample object description information, a reference script is generated to obtain reference script information; Based on the reference script information, a reference image is generated to obtain the sample object image; Based on the sample object image and the sample object description information, a video script is generated to obtain the preset video script information of the multiple video frames; Image generation is performed based on the sample object image and the preset video script information of the multiple video frames to obtain the preset generated image of the multiple video frames; Based on the preset video script information and the preset generated image, the instruction settings are performed to obtain the sample instruction text.

10. A video processing apparatus, characterized in that, The device includes: The instruction acquisition module is configured to acquire an object image and video generation instruction text; the video generation instruction text is used to instruct the generation of a video with a coherent plot associated with the target object, and the object image is a reference image containing the target object; The video script generation module is configured to input the object image and the video generation instruction text into the video script generation model. During the process of generating the video script based on the object image and the video generation instruction text, the video script information of each generated video frame is aggregated into script image feature information based on the historical frame feature information corresponding to each video frame, thereby obtaining script image feature information corresponding to multiple video frames. The video generation module is configured to input the object image and the script image feature information corresponding to the multiple video frames into the video generation model to generate a target generated video containing the generated images corresponding to the multiple video frames.

11. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the video processing method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the video processing method as described in any one of claims 1 to 9.

13. A computer program product, characterized in that, It includes computer program instructions that, when executed by a computer's processor, cause the computer to perform the video processing method as described in any one of claims 1 to 9.