Video generation method and apparatus, electronic device, and medium

CN121711542BActive Publication Date: 2026-09-22BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511835441.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-09-22
Estimated Expiration
2045-12-05

AI Technical Summary

Benefits of technology

[0012]根据本公开的另一方面,提供了一种计算机程序产品,包括计算机程序,其中,所述计算机程序在被处理器执行时实现上述方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121711542B_ABST
    Figure CN121711542B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video generation method and device, electronic equipment and medium, relates to the technical field of artificial intelligence, and particularly relates to the technical field of video generation, video processing and the like. The implementation scheme is: obtaining input information of a user, including video duration information, description information of a target object and reference information; analyzing the input information to generate a video script; generating a first frame preview of a video based on the video script and a target image; in response to a confirmation instruction of the user on the first frame preview of the video, determining a total number of video clips contained in a to-be-generated video; using a picture generation model to generate a plurality of key frames on a time axis based on the first frame preview of the video and the total number of video clips; for each pair of key frames adjacent on the time axis in the plurality of key frames, using a video generation model to generate a connected video clip based on the pair of key frames; and splicing the connected video clips corresponding to each pair of adjacent key frames in the order of the time axis to generate a target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of video generation and video processing, specifically to a video generation method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] With the rapid development of mobile internet technology and the continuous innovation of e-commerce models, the way products are displayed and marketed is undergoing a profound transformation from static images and text to dynamic short videos and live streaming. Under this trend, high-quality, high-frequency product demonstration videos have become a core medium for merchants to attract user attention, convey product information, and improve conversion rates.

[0004] In recent years, artificial intelligence technologies, represented by deep learning, have made groundbreaking progress in the field of multimedia content generation. In particular, text understanding technology based on large language models and image / video generation technology based on diffusion models have provided new possibilities for the automated production of video content. As an emerging interactive medium, digital humans, due to their advantages such as being online 24 / 7 and having customizable appearances, are gradually being applied in scenarios such as e-commerce sales, news broadcasting, and online education, becoming an important means of replacing live-action filming.

[0005] However, in e-commerce and other vertical application scenarios, video generation is not merely a simple compositing of images. It involves a precise understanding of user marketing intentions, a rigorous reproduction of product visual features, and the logical arrangement of interactions between people and products. Maintaining high production efficiency while ensuring consistent visual style, stable product display, and coherent script logic over long periods is a crucial challenge for the industrial application of AI content generation technology. Therefore, proposing a method that can accurately capture user needs through natural language dialogue, automatically plan storyboards, and stably and controllably generate high-quality videos containing complex interactions has significant application value and practical significance for improving content production efficiency and lowering production barriers.

[0006] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0007] This disclosure provides a video generation method, apparatus, electronic device, computer-readable storage medium, and computer program product.

[0008] According to one aspect of this disclosure, a video generation method is provided, comprising: acquiring user input information, the input information including video duration information of the video to be generated, description information of a target object, and reference information, wherein the reference information is used to provide a target image of the target object; parsing the input information to generate a video script, wherein the video script includes a video theme field and a display method field of the target object; generating a preview image of the first frame of the video based on the video script and the target image; responding to a user's confirmation instruction for the preview image of the first frame of the video, determining the total number of video segments contained in the video to be generated based on the video script and the video duration information; generating multiple keyframes on a timeline using a raw image model based on the preview image of the first frame of the video and the total number of video segments; generating connecting video segments based on each pair of keyframes that are adjacent on the timeline using a video generation model; and splicing the connecting video segments corresponding to each pair of adjacent keyframes in timeline order to generate a target video.

[0009] According to another aspect of this disclosure, a video generation apparatus is provided, comprising: an acquisition module configured to acquire user input information, the input information including video duration information of the video to be generated, description information of a target object, and reference information, wherein the reference information is used to provide a target image of the target object; a parsing module configured to parse the input information to generate a video script, wherein the video script includes a video theme field and a display method field of the target object; a first generation module configured to generate a first frame preview image of the video based on the video script and the target image; and a determination module configured to respond to the input information. The system comprises: a first generation module, configured to receive a user confirmation instruction for the first frame preview of the video; a second generation module, configured to determine the total number of video segments contained in the video to be generated based on the video script and the total number of video segments; a third generation module, configured to generate a series of video segments based on each pair of adjacent keyframes on the timeline using a video generation model; and a stitching module, configured to stitch the corresponding video segments of each pair of adjacent keyframes in chronological order to generate the target video.

[0010] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods described above.

[0011] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the above-described method.

[0012] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the above-described method.

[0013] According to one or more embodiments of this disclosure, a video generation method is provided. This method achieves WYSIWYG pre-verification by introducing an intermediate state of generating a video script and generating a preview image of the first frame before video generation, and requiring the user to confirm the first frame preview image. The energy-intensive video generation step is only initiated after the user approves the first frame, thus avoiding blind generation. Furthermore, an anchored generation strategy is adopted, using the confirmed first frame preview image as a reference to generate subsequent timeline keyframes, and generating seamless video segments based on adjacent keyframes. This segmented generation method ensures that the visual features of the entire long video are always anchored to the user-confirmed first frame, solving the style drift and logic breakdown problems that are prone to occur in long video generation.

[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0015] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0016] Figure 1 This is a schematic diagram illustrating an example system in which various methods described herein may be implemented according to exemplary embodiments;

[0017] Figure 2 A flowchart of a video generation method according to an embodiment of the present disclosure is shown; Figure 3 A flowchart illustrating a portion of the process of a video generation method according to an embodiment of the present disclosure is shown; Figure 4 A schematic diagram illustrating a portion of the flow of a video generation method according to an embodiment of the present disclosure is shown; Figure 5A and Figure 5B A schematic diagram illustrating a portion of the flow of a video generation method according to an embodiment of the present disclosure is shown; Figure 6A and Figure 6B A schematic diagram illustrating a portion of the flow of a video generation method according to an embodiment of the present disclosure is shown; Figure 7 A schematic diagram illustrating timing planning in video generation according to an embodiment of the present disclosure is shown; Figure 8 A structural block diagram of a video generation apparatus according to an embodiment of the present disclosure is shown; and Figure 9 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0018] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0019] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0020] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0021] In related technologies, an end-to-end model is typically used to directly generate videos based on input prompts. In this model, users cannot predict the visual effect before the generated result is available. If the generated character, scene, or product form does not meet expectations, a complete regeneration is required, resulting in a significant waste of computing resources, a long production cycle, and difficulty in accurately controlling the visual style consistency of long videos.

[0022] To address the aforementioned issues, this disclosure provides a video generation method that introduces an intermediate state—generating a video script and generating a preview image of the first frame—before video generation, requiring user confirmation of the first frame preview image, thus achieving WYSIWYG pre-verification. The energy-intensive video generation step is only initiated after the user approves the first frame, avoiding blind generation. Furthermore, an anchored generation strategy is employed, using the confirmed first frame preview image as a baseline to generate subsequent timeline keyframes, and generating seamless video segments based on adjacent keyframes. This segmented generation method ensures that the visual features of the entire long video are always anchored to the user-confirmed first frame, resolving the style drift and logic breakdown problems that easily occur in long video generation.

[0023] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0024] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0025] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of video generation methods.

[0026] In some embodiments, server 120 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.

[0027] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the methods described herein, and is not intended to be limiting.

[0028] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to execute video generation methods. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0029] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0030] Network 110 can be any type of network well known to those skilled in the art, and can support data communication using any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0031] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0032] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0033] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0034] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0035] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0036] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0037] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0038] It should be noted that the video generation method provided in this application has broad application prospects and versatility. Firstly, the technical solution of this application is particularly suitable for the e-commerce field. In this field, merchants have a huge demand for high-quality, high-frequency product display videos. Using the method of this application, realistic digital human anchors can be generated, enabling the anchor to vividly explain and demonstrate the products, i.e., the target audience. For example, when generating e-commerce short videos or live stream clips, this solution can precisely control the anchor's posture, image, and the way products are stacked and displayed. It can even generate complex interactive actions such as the anchor picking up products to demonstrate details, thereby effectively improving product exposure and conversion rates.

[0039] However, it should be understood that the application scenarios of this application are not limited to this. The video generation method proposed in this application is also applicable to other scenarios that require the generation of high-quality video content. For example: in online education scenarios, it is used to generate teaching videos of virtual teachers, where the "target object" can be teaching aids, experimental equipment, or books, and the virtual teacher can hold the teaching aids to demonstrate and explain; in news media scenarios, it is used to generate broadcast videos of virtual news anchors, where the "target object" can be a news display board, a handheld card, or an interview microphone; in entertainment and social scenarios, it is used to generate interactive videos of virtual idols, where the "target object" can be everyday props, gifts, or musical instruments; in corporate training scenarios, it is used to generate product operation demonstration videos, where the "target object" can be specific mechanical parts or operating tools.

[0040] To enable those skilled in the art to more clearly and intuitively understand the technical solution and implementation details of the present invention, the following description of specific embodiments will primarily use the generation of product display videos in the e-commerce field as a typical application scenario for illustration. However, this exemplary description should not be construed as limiting the scope of protection of the present invention, and the technical solution of the present invention can be fully applied to other fields mentioned above based on the same technical principles.

[0041] Figure 2 A flowchart of a video generation method according to an embodiment of the present disclosure is shown.

[0042] like Figure 2 As shown, the video generation method 200 includes: Step S201: Obtain user input information, which includes video duration information of the video to be generated, description information of the target object, and reference information, wherein the reference information is used to provide the target image of the target object; Step S202: Parse the input information to generate a video script, wherein the video script includes a video theme field and a display method field for the target object; Step S203: Based on the video script and the target image, generate a preview image of the first frame of the video; Step S204: In response to the user's confirmation instruction on the first frame preview of the video, determine the total number of video segments contained in the video to be generated based on the video script and the video duration information; Step S205: Using the raw image model, generate multiple keyframes on the timeline based on the preview image of the first frame of the video and the total number of video segments; Step S206: For each pair of keyframes that are adjacent on the time axis among the plurality of keyframes, generate a continuous video segment based on the pair of keyframes using a video generation model; and Step S207: Stitch together the video segments corresponding to each pair of adjacent keyframes in chronological order to generate the target video.

[0043] Understandably, the video generation method 200 is executed in response to a user's video generation request to generate a target video containing the target object. Therefore, in step S201, the user's input information is first received. The input information can be multimodal, including but not limited to a text description of the target object, such as "generate a video introducing this coffee machine"; reference materials, such as when the target object is a product, the user can upload the main image link or image file of the product as reference material for the target object; and the user's video length requirement for the generated video, for example, 30 seconds.

[0044] In step S202, the input information is parsed to output a video script for generating the target video. This script not only includes the video theme but also specifies how the target object will be displayed in the target video to be generated. For example, whether the target object needs to be stacked in the frame or whether a person in the video needs to hold and display the target object.

[0045] For example, a large model, such as a large language model (LLM), can be invoked to parse the multimodal input information of the user input, thereby obtaining information such as the video theme of the target video to be generated and how the target object is displayed in the target video, which is used as part of the video script to clarify the display content and presentation format of the final generated target video.

[0046] In step S203, based on the target image provided by the video script and reference materials, an image generation and compositing technique is used to generate a static preview image of the first frame of the video. When method 200 is applied to a product display scenario in the e-commerce field, the preview image of the first frame of the video can be used to display the anchor's image, the scene environment, and the approximate position of the target object in the frame for users to preview.

[0047] In step S204, a preview image can be displayed to the user. This preview image is only used as the basis for subsequent generation after the user sends a confirmation command for the first frame preview image. The total number of video segments to be generated is then calculated based on the user-set duration. It is understood that the total number of video segments can be calculated based on a preset slicing rule, such as one segment every 5 seconds, combined with the total video duration. For example, when the user sets the video duration to 30 seconds, the total number of video segments can be determined to be 6 based on the slicing rule of one segment every 5 seconds.

[0048] In step S205, using the raw image model with the user-confirmed first frame preview as a strong reference, keyframes are generated at each node of the timeline. For example, keyframes can be generated at the connection node between two adjacent segments, as in the example above, i.e., a keyframe is generated every 5 seconds starting from 0s.

[0049] In step S206, a transition segment is generated. Specifically, for every two adjacent keyframes, the video generation model is invoked to generate a dynamic video segment that transitions from the previous keyframe to the next keyframe.

[0050] Finally, in step S207, all the generated video clips are spliced ​​together strictly in chronological order to generate the target video.

[0051] As can be seen, video generation method 200, by introducing a video script to pre-plan the target video to be generated and incorporating an intermediate state of first-frame preview, breaks the traditional black-box mode of AIGC input-output. This allows users to verify the results before consuming computing power to generate the full video, significantly reducing trial-and-error costs. Simultaneously, it adopts an anchored generation strategy, where all subsequent keyframes are generated based on the user-confirmed first frame, and video segments are transitions generated based on adjacent keyframes. This segmented generation architecture effectively prevents style drift and logic breakdown issues common in long video generation, significantly improving the consistency and smoothness of the generated video and enhancing the user's viewing experience.

[0052] Furthermore, in the process of generating the video script, in order to achieve accurate planning of the display method of the target object and reduce the professional threshold for users when writing prompts, this disclosure provides a specific implementation scheme for intelligent reasoning based on the attributes of the target object, so that the method 200 can automatically analyze the category characteristics of the target object and match the camera language that best matches the physical attributes and display logic of the item, such as stacking or handheld, thereby generating a more reasonable and professional video script.

[0053] According to some embodiments, step S202 includes: determining the category attribute of the target object based on the description information; determining the display method field based on the category attribute and the target image, including: determining that the display method field includes stacked display in response to determining that the category attribute belongs to daily necessities or small appliances; determining that the display method field does not include stacked display in response to determining that the category attribute belongs to large home appliances or virtual goods; and determining that the display method field includes handheld display in response to determining that the category attribute belongs to a regular rigid body class and the size meets the predetermined size requirements.

[0054] For example, a large language model or classification model can be used to perform semantic analysis on the descriptive information of the target object input by the user and / or image analysis on the target image pointed to by the reference material, in order to determine the category attributes of the target object. For example, it can be identified that "facial cleanser" belongs to "daily necessities", "double-door refrigerator" belongs to "large household appliances", "thermos cup" belongs to "regular rigid body", and so on.

[0055] If the target object is determined to be a daily necessity or a small appliance, meaning it is small in size and suitable for desktop placement, step S202 automatically configures the "Display Method Field" in the video script to include stacked display, i.e., the product is placed on the table in front of the host. If the target object is determined to be a large home appliance or a virtual product, meaning it is too large or has no physical form, the "Display Method Field" is configured not to include stacked display. If the target object is determined to be a regular rigid body and its size meets the predetermined size requirements, meaning it is not easily deformable and can be held with one or two hands, the "Display Method Field" is configured to include handheld display.

[0056] Therefore, without requiring users to possess professional directing skills, Method 200 can automatically plan the most appropriate camera language based on the physical attributes of the products. For example, small items are suitable for being piled on a table, while large items are suitable for background display. Furthermore, through rule constraints, it can avoid generating logical errors and prevent the model from generating scenes that violate common sense or e-commerce standards, such as "holding a refrigerator" or "a table piled with cars."

[0057] After clarifying the presentation method of the target audience, in order to build a complete and expressive video script, it is also necessary to refine the character images, specific postures, and scene environment in the video. Since user input information is usually multimodal and contains implicit needs, such as a user uploading a specific background image or text implying requirements for the anchor's style, it is also necessary to further conduct in-depth intent mining of the input information to accurately capture the user's personalized needs and determine the generation strategy of each key field in the script accordingly.

[0058] According to some embodiments, step S202 further includes: using a large model to perform intent parsing on the input information to obtain intent tags corresponding to the input information, wherein the intent tags represent one or more of the following: whether the input information contains a scene reference image, whether the input information specifies the posture of a person in the target video, whether the input information contains image description information of the person, and whether the input information contains scene environment description information; determining an information source for generating the video script based on the intent tags, wherein the information source includes one or more of the information in the input information; generating the video script based on the information source and the category attribute, wherein the video script further includes: a posture field representing the posture of a person in the video to be generated, an image field representing the image of the person, and a scene field representing the scene environment of the video to be generated.

[0059] In practical applications, to meet users' more refined and personalized customization needs, in addition to the aforementioned necessary information—video duration, target object description, and reference information—users can also input various optional information. Specifically, this optional information includes, but is not limited to, scene reference diagrams, image reference diagrams, and detailed text instructions containing specific constraints. In this embodiment, the process of using a large model, such as a large language model, to parse the input information and generate a video script employs a hierarchical decision-making mechanism that prioritizes user intent.

[0060] The large language model performs multi-dimensional intent labeling on user input. For example, it may determine whether the user uploaded a specific background image to identify the scene reference image label; whether the user requested the anchor's posture to identify the posture-specific label; or whether the user described the anchor's gender and appearance to identify the persona label, thus obtaining a multi-dimensional intent label.

[0061] The source of each element in the generated script is then determined based on the tags. For example, if a scene reference image is detected in the user's input, the "scene field" in the video script directly extracts features from the reference image; otherwise, it is inferred from the category attributes of the target object. Thus, based on the above logic, the model can output a complete video script containing posture fields (such as "sitting posture"), image fields (such as "professional woman"), and scene fields (such as "bright live broadcast room") for the final video generation.

[0062] This allows the system to handle ambiguous or missing user input and generate complete video scripts through the complementarity of multimodal information. Furthermore, it transforms unstructured user requirements into structured control fields, providing precise prompt parameters for subsequent image generation models and improving the accuracy of the generated videos.

[0063] The following section will provide a more detailed description of the user intent-based hierarchical decision-making mechanism adopted in Method 200.

[0064] First, regarding the processing of scene reference images, when generating the "scene field" of the video script, the system prioritizes checking if a scene reference image exists in the input information. If it does, the visual features of the image are extracted and written into the video script as the highest priority constraint. When generating prompts, the large model directly describes the background as the specific scene uploaded by the user, such as a living room with a Christmas atmosphere, rather than generalizing and inferring based on product attributes. If no scene reference image exists, the default inference logic is executed, which automatically matches the appropriate scene based on the attributes of the target object. For example, if the product is identified as a pot cleaner, a kitchen scene is automatically planned.

[0065] Secondly, regarding the processing of image reference images, the "image field" in the video script is also primarily generated based on these reference images. When a user uploads a photo of a specific person, such as a brand ambassador or a screenshot of an online influencer, the features of the person in the reference image are extracted and converted into structured image descriptions to ensure that the generated digital avatar anchor visually closely resembles the reference image. If no image reference image is provided, it will be inferred from the user's text description or based on product attributes; for example, if the product is men's facial cleanser, the inferred image will be a clean-cut male.

[0066] Furthermore, when processing specific text instructions, the large language model identifies explicit constraint labels during intent tagging. When a user inputs clear constraints, such as "the host must stand while explaining," the decision logic prioritizes user intent over default rules. For example, while the default rule might be "use a seated posture if there are items piled up," if the user explicitly requests a standing posture, a "standing posture" video script will be generated, and the composition will be automatically adjusted when generating the first frame preview (e.g., using a half-body shot or raising the product stacking platform) to simultaneously satisfy both the standing posture and product stacking requirements.

[0067] Furthermore, regarding the image sources in the reference information, Method 200 supports directly uploading locally taken high-resolution product images, and also supports automatically capturing external data as source material when the user only provides a link, such as the main image of an e-commerce product detail page. Through the above processing of optional information, Method 200 allows users to determine the visual style of the video by uploading reference images, avoiding the randomness brought about by pure text generation. It satisfies users' strong consistency customization needs for specific scenes and images, while providing automated default inference to lower the usage threshold for ordinary users, achieving an effective balance between automation and customization.

[0068] After generating a video script containing detailed fields (such as scene, image, posture, and display method) through the above steps, the overall plan for generating the target video is obtained. To provide users with intuitive and verifiable visual feedback before consuming significant computing power to generate the full video, method 200 enters the first frame generation stage. Specifically, when the display method field in the video script is determined to include stacked display—that is, when products need to be displayed on a table in front of the anchor—to ensure that the appearance of the product (target object) in the image is completely consistent with the real product, this embodiment of the disclosure adopts a layered generation followed by layer fusion technique, rather than directly using an end-to-end text-to-image model for one-time generation. This avoids the product detail distortion or text garbling problems that may be caused by generative models.

[0069] Figure 3 A schematic diagram illustrating a portion of the flow of a video generation method according to an embodiment of the present disclosure is shown. Figure 3 Steps S301-S304 shown are performed in response to determining that the display mode field contains a stacked display, specifically, as Figure 3 As shown, step S203 includes: Step S301: In response to determining the display mode field indicating stacked display, perform subject cutout processing on the target object in the target image to obtain a transparent image of the target object; Step S302: Using the pose field, image field, and scene field in the video script, generate a real-scene portrait image containing the person; Step S303: Generate a stacked display effect image of the target object using the transparency image of the target object; and Step S304: Merge the stacked display effect image with the real-scene portrait image to obtain the first frame preview image of the video.

[0070] In step S301, the user-provided reference material (such as the main product image) is first preprocessed. Advanced image segmentation algorithms or salient object detection models can be used to automatically identify the outline of the target object in the image, accurately isolate it from the original background, and generate a transparent image of the target object. This operation ensures that the product image displayed in the subsequent video is the target object, preserving all visual details of the product and laying the foundation for subsequent compositing.

[0071] Then, based on the visual parameters determined in the video script, a high-quality base image is generated using a text-based image model. For example, based on the posture field (sitting posture), image field (30-year-old professional woman), and scene field (warm home background), a real-life photo of the anchor sitting at a table is generated. It should be noted that in the real-life portrait image generated at this stage, the table in front of the anchor may be empty, or may only contain general decorations and may not yet include the specific target object.

[0072] To simulate a real live-streaming e-commerce scenario, single product images need to be transformed into a display effect with spatial depth and richness. Step S303 uses the transparent image obtained in step S301 to generate a staggered stacking effect of products in space through image processing algorithms. For example, multiple products can be copied and arranged according to perspective, or a combination of products and free gifts can be generated to form an independent stacked display effect. During this process, a quality inspection model can also be used to screen the stacking effect to ensure that the arrangement meets aesthetic standards and that the occlusion relationships are reasonable.

[0073] Finally, step S304 uses the stacked display effect image as a foreground layer, overlaying it onto the desktop area of ​​the real-life portrait image. To ensure a natural blend, shadows, brightness, and edge feathering can be applied to the foreground stacked product layer based on the background lighting conditions and perspective angle, making the products appear as if they were actually placed on the table in front of the host, thus synthesizing the final first frame preview image of the video.

[0074] Through the layered processing described in S301 to S304, a high-quality WYSIWYG preview is achieved. Compared to traditional full-image generation methods, this solution not only restores the true appearance details of the product, but also allows users to make targeted modifications to unsatisfactory parts during the preview stage. For example, they can modify only the anchor's image while keeping the product display, or simply adjust the product placement, greatly improving the controllability of the generated results and user satisfaction.

[0075] Once the user confirms the first frame preview of the video is correct, this preview becomes the visual anchor and quality benchmark for subsequent video generation. However, when using a generative AI model to extend the generation of keyframes for subsequent time points (such as the 10th second and the 20th second) based on this first frame, a common challenge arises: the generative model may incorrectly alter the shapes of objects in the background that should remain stationary while maintaining the changes in the character's movements. For example, a product on a table might suddenly deform, change color, or jump in position. To ensure that the appearance of the core product remains stable and consistent throughout the entire video playback, this embodiment introduces a consistency check mechanism based on visual feature comparison during the keyframe generation stage.

[0076] According to some embodiments, step S205 includes: in response to determining that the display method field indicates stacked display, using the preview image of the first frame of the video as a reference image, generating multiple candidate keyframes on the timeline based on the total number of video segments using the raw image model; performing a consistency check on the multiple candidate keyframes, wherein the consistency check includes calculating the similarity between the target object stacked in each of the multiple candidate keyframes and the target object stacked in the preview image of the first frame of the video; in response to the similarity corresponding to each of the multiple candidate keyframes not being lower than a preset threshold, determining the multiple candidate keyframes as the multiple keyframes; in response to the existence of candidate keyframes with similarity lower than the preset threshold among the multiple candidate keyframes, regenerating multiple candidate frames using the raw image model for consistency check.

[0077] In this embodiment, the number and location of keyframes to be generated on the timeline are first determined based on the calculated total number of video segments. For each target time point, an image generation model is invoked, and the preview image of the first frame of the video is passed in as a strong reference condition to generate a candidate keyframe corresponding to that time point. This image should maintain continuity with the first frame in terms of composition and human features, but the human posture may change naturally over time.

[0078] Understandably, the quality of generated candidate frames cannot be judged solely by visual inspection; a consistency check must be performed. Specifically, this check involves: first, using object detection or segmentation algorithms to accurately locate and extract image features belonging to the stacked display area in the candidate keyframes; simultaneously, extracting target object features from the same area in the first preview frame of the video, serving as a baseline. Next, the similarity between the two is calculated. This similarity calculation can be pixel-level (e.g., SSIM structural similarity) or semantic-level (e.g., CLIP Score), aiming to quantitatively assess whether the goods in the generated image have undergone abnormal changes.

[0079] If the calculated similarity score is higher than the set stringent threshold, it means that the product shape in the generated image is highly consistent with the first frame and no obvious deformation has occurred. At this point, the candidate frame is deemed qualified and officially established as the keyframe for that time point, for use in the generation of subsequent video segments.

[0080] Conversely, if the similarity is below the threshold, it indicates that the candidate keyframes generated by the generative model are of poor quality, resulting in defects such as distortion, blurring, or replacement of background products. In this case, a retry mechanism is automatically triggered, discarding the current unqualified frame, adjusting the random seed or fine-tuning the parameters, and calling the raw image model again to regenerate new candidate frames, and re-entering the consistency check process. This process may be repeated multiple times (e.g., up to 5 retries) until a qualified image is generated or a fallback strategy is triggered, such as forced texture repair, which directly reuses the first preview frame of the video as the keyframe of the current time node, thereby maximizing the rigor and high quality of the final video image, ensuring the stability of the appearance features of the target object in the final video image, and avoiding flickering or deformation of the target object in the video background due to the randomness of the model.

[0081] After establishing a series of keyframes on the timeline through the aforementioned consistency check mechanism, these keyframes constitute the static skeleton of the target video. To make the image dynamic and ensure a smooth and natural transition of video content from one keyframe to the next, dynamic content needs to be generated to fill the gaps between adjacent keyframes. However, existing graph-based video models often struggle to precisely control the start and end frames simultaneously in a single generation, and the generation time is limited. To address this technical bottleneck, this disclosure proposes a two-stage transition generation strategy, achieving precise directional transition between the first and last frames through a step-by-step approximation approach.

[0082] According to some embodiments, step S206 includes: for each pair of keyframes that are adjacent on the time axis among the plurality of keyframes, generating a first transition video segment based on the pair of keyframes using the video generation model; using the last frame of the first transition video segment as the first image and the second keyframe of the two adjacent keyframes as the last image, generating a second transition video segment using the video generation model; and splicing the first transition video segment and the second transition video segment to obtain a connecting video segment corresponding to the pair of keyframes.

[0083] For each pair of keyframes adjacent on the timeline, for ease of description, they will be referred to hereafter as the previous keyframe A and the next keyframe B. Method 200 does not directly attempt to generate the video connecting A and B at once, but breaks it down into two sub-steps: First, based on the pair of keyframes, a first transitional video segment is generated using a video generation model. Specifically, keyframe A is used as the input image to the video generation model to generate a short video clip (Part 1). In this stage, the model mainly focuses on the dynamic extension starting from A. The generated video tail frame (A') may have some differences from keyframe B in composition or character pose, and has not yet fully returned to the state of keyframe B. Next, the tail frame of the first transitional video segment is used as the first image, and the next keyframe of the two adjacent keyframes is used as the tail image to generate a second transitional video segment using the video generation model. The last frame A' of Part 1 is extracted and used as the new starting image, while keyframe B is used as the forced constraint ending image, and the video generation model is called again to generate the second video segment (Part 2). In this step, to ensure a smooth transition, fixed text prompts such as "smooth transition" and "natural connection" can be used. This process is equivalent to forcibly pulling the video frame from state A' back to state B. Finally, the first and second transition video segments are spliced ​​together to obtain the connecting video clip corresponding to the pair of keyframes. By connecting Part 1 and Part 2 end-to-end, a complete transition video that accurately connects to keyframe A and ultimately returns to keyframe B can be successfully constructed, effectively eliminating the common scene jump problem in long video generation.

[0084] By performing the above operation on each pair of adjacent keyframes, a long video clip with a smooth transition from the first keyframe to the last keyframe can be obtained, which is the base video.

[0085] At this point, the display method field in the video script will be used for judgment. If this field indicates that it does not include handheld display, such as only requiring the anchor to give a verbal explanation at the table without picking up the target object, then the basic base video generated above is the final target video, and the generation process ends here.

[0086] However, if the display method field indicates that it includes handheld display, such as requiring the host to pick up the target object for detailed demonstration, then while the current video has a coherent background and the host's state, it lacks crucial interactive actions. Therefore, the subsequent steps will continue, entering the independent generation and logical insertion process of the product-held video clip, in order to overlay high-quality dynamic interactive content on top of the aforementioned base video.

[0087] According to some embodiments, in response to determining that the display mode field includes handheld display, the method 200 further includes: generating a handheld action prompt word according to the video script; selecting a first frame image from the generated plurality of keyframes as a starting reference frame; and generating an initial handheld display segment based on the handheld action prompt word and the starting reference frame using the video generation model; and obtaining the last frame of the initial handheld display segment and using it as the starting frame of the next handheld display video segment, so as to generate a plurality of handheld display video segments that are sequentially connected in time using the video generation model.

[0088] First, generate cue words for the handheld action based on the video script. Using a structured cue word template, product information is filled in to generate precise descriptions such as "The host naturally picks up [target object] and smiles as they show details to the camera," ensuring that the generated action conforms to physical laws and marketing logic. Second, select the first frame from the generated keyframes as the starting reference frame. To ensure that the image of the person, clothing, and lighting in the product-holding segment are consistent with the main video, an image needs to be extracted from the existing video stream, such as the frame preceding the time point when the action is planned to be inserted in the aforementioned connecting video segment, as a reference baseline.

[0089] Next, an initial handheld display segment is generated using a video generation model based on handheld action cues and a starting reference frame. This step calls the video generation model, focusing on generating the action of "picking up". Furthermore, to overcome the limitations of existing models' single-generation duration (e.g., only 3-5 seconds) and achieve continuous narration lasting tens of seconds, this embodiment employs a relay generation mechanism: the last frame of the initial handheld display segment is obtained and used as the starting frame of the next handheld display video segment. This mechanism is used to repeatedly call the video generation model, using the end point of the previous segment as the starting point of the next, thus generating multiple handheld display video segments that are continuous in time and have coherent actions. For example, a series of coherent actions such as "picking up", "displaying", "rotating the product", and "putting down" can be generated consecutively, meeting the needs of product narration of any length.

[0090] At this point, we have obtained the non-contained transitional segments representing the main background storyline and the featured segment showcasing the highlights and interactions. The final step is to merge these two parts of the footage according to the timeline plan in the video script, ensuring the correctness of the timeline logic while eliminating any abruptness at the joints between different clips, ultimately producing the finished video.

[0091] According to some embodiments, step S208 includes: splicing the connecting video segments corresponding to each pair of adjacent keyframes in chronological order to obtain a basic video sequence; inserting the plurality of handheld display video segments into the corresponding positions of the basic video sequence according to preset timing logic; and generating a transition frame at the connection between the basic video sequence and the plurality of handheld display video segments to form the target video.

[0092] First, the video segments corresponding to each pair of adjacent keyframes are spliced ​​together according to the timeline to obtain the basic video sequence. All Part1+Part2 segments generated in the previous steps are then connected sequentially to form a complete base video of the anchor explaining the product without holding it. Second, according to a preset timing logic, the multiple handheld display video segments are inserted into the corresponding positions of the basic video sequence. The script's plan, such as "product display from second 15 to second 30," is read, and the product-holding segments generated in the previous steps are precisely placed within the corresponding time periods of the basic video sequence. Finally, transition frames are generated at the junctions of the basic video sequence and the multiple handheld display video segments to form the target video.

[0093] Since "not holding an item" and "holding an item" are two independent generation paths, direct splicing may cause abrupt changes in the position of the person's hands, such as the hand being on the table in one frame and suddenly raised in mid-air in the next. To address this, video frame interpolation algorithms or short video generation models are used at the insertion point to generate several smooth transition frames, simulating the natural trajectory of the arm raising and lowering, thereby eliminating editing artifacts and synthesizing a smooth and natural final target video.

[0094] Figure 4 , Figure 5A and Figure 5B as well as Figure 6A and Figure 6B Schematic diagrams of portions of the video generation method according to embodiments of the present disclosure are shown. Figure 7 A schematic diagram illustrating timing planning in video generation according to an embodiment of the present disclosure is shown, and the applicant will describe the entire process of the video generation method in conjunction with these figures.

[0095] The first step is the initial stage of video generation, namely information preprocessing and video script generation, such as... Figure 4As shown in "Phase 1" on the left. First, multimodal input from the user is acquired. This input may include product links, reference images (images or scenes), video duration (e.g., 180 seconds), and text descriptions. A large language model is used to deconstruct and deeply understand the intent behind this input. For example, if a user uploads a product image, the model is invoked to perform subject cutout, generating a "transparent product image"; if a user provides a product link, the data is read to generate structured product information. Based on this, according to different intent tagging results, the filtered information sources (such as reference images, product understanding information, and user descriptions) are fed into preset structured prompts to generate a detailed video script (i.e.,...). Figure 4 (See the overall generation plan description shown). For example, the generated video script is presented as structured data (such as JSON format), which includes not only macro-level themes but also micro-level visual control parameters. For example, the video script may include: "Video Theme" as "A woman aged 30-40 explains a vegetable and fruit garden pot bottom cleaner in the kitchen"; "Image Description" as "The host is a woman around 30-40 years old, with long black hair, exquisite makeup, and wearing a black top"; "Scene Description" as "A warm and comforting home kitchen, with a medium close-up shot. The background features white cabinets, a kitchen sink, and green plants, while the foreground is a white table, creating an overall atmosphere close to everyday use"; "Product Description" as "Bottled and canned cleaner"; "Host Posture" as "The host sits at the table in the scene"; "Explanation Style" as "Enthusiastic, friendly, easy to understand, and detailed in demonstration"; "Whether to Display Products" as "Products need to be displayed on the table"; "Whether to Hold Products" as "Interspersed with product-holding explanation segments"; and "Video Length" as "180". This script serves as a blueprint, defining the tone for subsequent generation. For example, the "Whether it's a stacked item" field, as part of the display method field, determines whether product cutout compositing is needed when generating the first frame; the "Whether it's a held item" field, as part of the display method field, determines whether a separate handheld action clip needs to be generated for subsequent video generation.

[0096] Subsequently, the process enters the stage of generating and confirming the first frame preview image, such as... Figure 5A As shown in "Phase 2," after obtaining the video script, the video is not generated immediately. Instead, a "first frame preview" is first created for user confirmation. Specifically, the video generation method reads the "product stacking" field from the video script. If it is "yes" (e.g., "products need to be displayed on the table" in the example above), the generation flow for products stacking is entered: table material is randomly selected from the material library, combined with the previously cut-out transparent product image, to generate a "product stacking texture." At the same time, based on the "image description" and "scene description" in the video script, a realistic portrait image of the anchor sitting in front of the table that meets the requirements is generated (e.g., ...). Figure 5B(As shown on the left). Finally, the "item texture" and the "real-world portrait image" are combined to obtain the final preview image of the first frame of the video (as shown on the left). Figure 5B (As shown on the right). This image visually illustrates the combined effect of the person, the target object, and the scene.

[0097] Once the user confirms that the first frame preview is correct, the actual video generation stage begins, namely global information extraction and process judgment, such as... Figure 6A As shown in "Phase 3". First, the "First Image Generation Plan Description" confirmed by the user is input into the LLM model again for global information extraction and process judgment to determine the specific generation path and parameters. The LLM model performs the following judgments: First, it calculates the total number of segments by dividing the video duration (180 seconds) by the single segment duration (15 seconds), which is 12. Then, it determines that the initial state is "sitting" based on the "anchor posture" in the script containing "sitting". Next, it determines that the "whether to hold an item" in the script is "interspersed with item-holding explanation segments". Finally, based on the comprehensive judgment, since the "initial state" is "sitting" and "whether to hold an item" is "yes", the path of interspersed item-holding base plate generation is selected. The model finally outputs global control information containing the above information.

[0098] Based on the aforementioned global information, multiple keyframes on the timeline are generated using a raw image model, and a consistency check is performed. Based on the characteristics of the first frame preview, prompts for multiple keyframes are automatically generated. For example: "Realistic style. A woman in a pink cheongsam maintains a smile, her gaze gentle and focused forward. Her hands, naturally open (palms up, arms slightly raised), slowly retract and gently overlap in front of her chest (fingers relaxed and not stiff), the movement fluid and elegant. Keep the camera fixed; the person, background (landscape painting, greenery), and clothing remain unchanged; fingers are not distorted; the gaze is focused and not wandering; blinking is natural." After generating each non-product keyframe, a consistency check is performed. Specifically, through an image reasoning model, the generated keyframes are compared with the first frame preview to check for positional shifts, shape distortions, or quantity changes in the items on the foreground table (i.e., the pile of items). If the check fails, it will automatically retry (up to 5 times); if it still fails, a fallback logic is triggered, directly reusing the first frame image as the keyframe, thus ensuring the absolute stability of the background products.

[0099] Next, the connecting video clips and the holding clips are generated. To generate a smooth background, the N keyframes are split into N adjacent groups (e.g., [Frame 1, Frame 2]). For each group, a two-stage generation method is used: first, based on the adjacent keyframes, the first video segment (Part 1) is generated using a graph-generated video model, and its last frame is extracted; then, using the last frame of Part 1 as the first image and Frame 2 as the last image, combined with fixed text prompts (e.g., "From..."), the video clips are generated. Figure 1 Natural transition to Figure 2The process begins with the host facing the camera directly, sharing knowledge with a natural sitting posture. This generates a 5-second transition video (Part 2). Parts 1 and 2 are then combined to form a complete 15-second video. For the "product holding" requirement planned in the script, separate product holding segments are generated. First, specific product holding action prompts are generated based on structured prompts. For example, receiving the "product object description" (e.g., "a bottle of wine with a yellow label") from the video script, it is filled into a preset template, generating the following prompt: "Have the model in the frame naturally pick up a bottle of wine with a yellow label in front of them and smile as they introduce it. The action should be smooth, like in a live-streaming sales scene, and the model should keep their eyes on the camera. Keep the camera fixed." Next, the initial product holding video is generated using this prompt and the starting reference frame. To achieve a longer product holding explanation duration, a loop generation strategy can be used, taking the last frame of the previous product holding video as the starting frame of the next segment, iterating repeatedly to generate multiple consecutive product holding videos.

[0100] Finally, perform the final logical concatenation, such as... Figure 6B As shown. The generated "non-product connection videos" are spliced ​​together sequentially to form the base video. Then, according to the script's timing plan (e.g.... Figure 7 As shown, the generated "holding action video clip" is inserted into the corresponding position on the base plate. A transition frame is generated at the insertion point to smoothly connect the actions.

[0101] Through the above-described entire process, this application not only achieves end-to-end automated production from user intent to finished video, but more importantly, it solves the problem of uncontrollable video generation. By using intermediate state verification of video scripts and first frame previews, users can lock the overall tone of the video before consuming a large amount of computing power. Through a consistency check strategy, it ensures that the core product display in the video remains stable and clear, solving the problem of background product distortion in existing technologies. Furthermore, through a splicing strategy that combines global information extraction and base insertion, it not only ensures the visual style of videos that are several minutes long, but also achieves the natural integration of complex interactive actions (such as picking up products), greatly improving the viewing experience and conversion rate of the video.

[0102] According to another aspect of this disclosure, a video generation apparatus is also provided.

[0103] like Figure 8As shown, the video generation device 800 includes: an acquisition module 801 configured to acquire user input information, the input information including video duration information of the video to be generated, description information of the target object, and reference information, wherein the reference information is used to provide a target image of the target object; a parsing module 802 configured to parse the input information to generate a video script, wherein the video script includes a video theme field and a display method field of the target object; a first generation module 803 configured to generate a preview image of the first frame of the video based on the video script and the target image; and a determination module 804 configured to respond to the user's input information. The confirmation instruction for the first frame preview of the video determines the total number of video segments contained in the video to be generated based on the video script and the video duration information; the second generation module 805 is configured to use a raw image model to generate multiple keyframes on the timeline based on the first frame preview of the video and the total number of video segments; the third generation module 806 is configured to use a video generation model to generate connected video segments based on each pair of keyframes that are adjacent on the timeline among the multiple keyframes; and the splicing module 807 is configured to splice the connected video segments corresponding to each pair of adjacent keyframes in timeline order to generate the target video.

[0104] The acquisition module 801 first receives the user's input information, which can be multimodal, including but not limited to a text description of the target object, such as "generate a video introducing this coffee machine"; reference materials, such as when the target object is a product, the user can upload the main image link or image file of the product as reference material for the target object; and the user's requirements for the length of the video to be generated, for example, 30 seconds.

[0105] The parsing module 802 parses the input information to output a video script for generating the target video. This script not only includes the video theme but also specifies how the target object will be displayed in the target video to be generated. For example, whether the target object needs to be stacked in the frame or whether a person in the video needs to hold and display the target object.

[0106] For example, a large model, such as a large language model (LLM), can be invoked to parse the multimodal input information of the user input, thereby obtaining information such as the video theme of the target video to be generated and how the target object is displayed in the target video, which is used as part of the video script to clarify the display content and presentation format of the final generated target video.

[0107] The first generation module 803 generates a static preview image of the first frame of a video based on the target image provided by the video script and reference materials, using image generation and compositing technology. When the video generation device 800 is applied to product display scenarios in the e-commerce field, the preview image of the first frame of the video can be used to display the anchor's image, the scene environment, and the approximate position of the target object in the frame for users to preview.

[0108] The determination module 804 can display a preview image to the user, and only after the user sends a confirmation command for the first frame preview image will that preview image be used as the basis for subsequent generation, and the total number of video segments to be generated will be calculated based on the user-set duration. It is understood that the total number of video segments can be calculated based on a preset slicing rule, such as one segment every 5 seconds, combined with the total video duration. For example, when the user sets the video duration to 30 seconds, according to the slicing rule of one segment every 5 seconds, the total number of video segments can be determined to be 6.

[0109] The second generation module 805 uses the raw image model with the user-confirmed first frame preview image as a strong reference to generate keyframes at various nodes on the timeline. For example, keyframes can be generated at the connection node between two adjacent segments, as shown in the example above, i.e., a keyframe is generated every 5 seconds starting from 0s.

[0110] The third generation module 806 generates connecting segments. Specifically, for every two adjacent keyframes, it calls the video generation model to generate dynamic video segments that transition from the previous keyframe to the next keyframe.

[0111] Finally, the splicing module 807 splices all the generated video clips strictly in chronological order to generate the target video.

[0112] As can be seen, the video generation device 800, by introducing a video script to pre-plan the target video to be generated and incorporating an intermediate state of first-frame preview, breaks the traditional black-box mode of AIGC input-output. This allows users to verify the results before consuming computing power to generate the full video, significantly reducing trial-and-error costs. Simultaneously, it employs an anchored generation strategy, where all subsequent keyframes are generated based on the user-confirmed first frame, and video segments are transitions generated based on adjacent keyframes. This segmented generation architecture effectively prevents style drift and logical breakdown issues common in long video generation, significantly improving the consistency and smoothness of the generated video and enhancing the user's viewing experience.

[0113] It is understood that the various components of the video generation device 800 can be configured to execute the various method steps defined in method 200 in order to achieve efficient task scheduling, which the applicant will not elaborate on here.

[0114] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a video generation method.

[0115] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause the computer to perform a video generation method.

[0116] According to another aspect of this disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements a video generation method when executed by a processor.

[0117] like Figure 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of the electronic device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0118] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, output unit 907, storage unit 908, and communication unit 909. Input unit 906 can be any type of device capable of inputting information to electronic device 900. Input unit 906 can receive input digital or character information and generate key signal input related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 907 can be any type of device capable of presenting information, and can include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 908 can include, but is not limited to, disk and optical disk. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth. TM Devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.

[0119] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as video generation methods. For example, in some embodiments, the video generation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the video generation method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the video generation method by any other suitable means (e.g., by means of firmware).

[0120] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0122] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0124] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0125] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0126] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0127] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A video generation method, the method comprising: Obtain user input information, which includes video duration information of the video to be generated, description information of the target object, and reference information, wherein the reference information is used to provide a target image of the target object; The input information is parsed to generate a video script, wherein the video script includes a video theme field and a display method field for the target object; Based on the video script and the target image, generate a preview image of the first frame of the video; In response to the user's confirmation instruction for the first frame preview of the video, the total number of video segments contained in the video to be generated is determined based on the video script and the video duration information; Based on the first frame preview image of the video and the total number of video segments, a raw image model is used to generate multiple keyframes on the timeline; For each pair of keyframes that are adjacent on the timeline among the plurality of keyframes, a video generation model is used to generate sequential video segments based on that pair of keyframes; and The target video is generated by splicing together the corresponding video segments of each pair of adjacent keyframes in chronological order.

2. The method according to claim 1, wherein, The step of parsing the input information to generate a video script includes: Based on the description information, the category attribute of the target object is determined; The display method field is determined based on the product category attribute and the target image, including: In response to determining that the category attribute belongs to daily necessities or small appliances, it is determined that the display method field includes stacked display; In response to determining that the category attribute belongs to large home appliances or virtual goods, it is determined that the display method field does not include stacked display; and In response to determining that the category attribute belongs to the regular rigid body class and the size meets the predetermined size requirements, it is determined that the display method field includes handheld display.

3. The method according to claim 2, wherein, The step of parsing the input information to generate a video script also includes: The large model is used to perform intent parsing on the input information to obtain the intent label corresponding to the input information. The intent label represents one or more of the following: whether the input information contains a scene reference image, whether the input information specifies the posture of the person in the video to be generated, whether the input information contains the image description information of the person, and whether the input information contains scene environment description information. Based on the intent tag, an information source for generating the video script is determined, wherein the information source includes one or more pieces of information from the input information; Based on the information source and the category attribute, the video script is generated, wherein the video script further includes: a posture field representing the posture of the person in the video to be generated, an image field representing the image of the person, and a scene field representing the scene environment of the video to be generated.

4. The method according to claim 3, wherein, The process of generating the first frame preview image of the video based on the video script and the target image includes: In response to the determination of the display mode field indicating stacked display, the target object in the target image is subjected to subject cutout processing to obtain a transparent image of the target object; Using the pose field, image field, and scene field in the video script, a real-scene portrait image containing the person is generated; Generate a stacked display effect image of the target object using the transparency image of the target object; and The stacked display effect image and the real-life portrait image are merged into layers to obtain the first frame preview image of the video.

5. The method according to claim 4, wherein, The step of generating multiple keyframes on the timeline using a raw image model based on the first frame preview image of the video and the total number of video segments includes: In response to determining that the display mode field indicates stacked display, using the first frame preview of the video as a reference image, the raw image model generates multiple candidate keyframes on the timeline based on the total number of video segments. A consistency check is performed on the plurality of candidate keyframes, wherein the consistency check includes calculating the similarity between the target object displayed in a stacked manner in each of the plurality of candidate keyframes and the target object displayed in a stacked manner in the preview image of the first frame of the video; In response to the fact that the similarity of each candidate keyframe in the plurality of candidate keyframes is not lower than a preset threshold, the plurality of candidate keyframes are determined as the plurality of keyframes; In response to the presence of candidate keyframes with similarity lower than the preset threshold among the multiple candidate keyframes, multiple candidate frames are regenerated using the raw image model for consistency checking.

6. The method according to any one of claims 1-5, wherein, The step of generating sequential video segments based on each pair of adjacent keyframes on the time axis using a video generation model includes: For each pair of keyframes that are adjacent on the time axis among the plurality of keyframes, Based on the keyframes, the first transition video segment is generated using the video generation model. The second transition video segment is generated using the video generation model by taking the last frame of the first transition video segment as the first image and the second key frame of two adjacent key frames as the last image. The first transition video segment and the second transition video segment are spliced ​​together to obtain a connecting video segment corresponding to the pair of keyframes.

7. The method according to any one of claims 2-5, wherein, In response to determining that the display method field includes handheld display, the method further includes: Generate handheld action prompts based on the video script; Select the first frame image as the starting reference frame from the generated plurality of keyframes; and Using the video generation model, an initial handheld display segment is generated based on the handheld action cue words and the starting reference frame; and The last frame of the initial handheld display segment is obtained and used as the starting frame of the next handheld display video segment, so as to generate multiple handheld display video segments that are consecutive in time using the video generation model.

8. The method according to claim 7, wherein, The step of splicing together the corresponding video segments of every two adjacent keyframes in chronological order to generate the target video includes: The video segments corresponding to each pair of adjacent keyframes are spliced ​​together in chronological order to obtain the basic video sequence. According to a preset timing logic, the multiple handheld display video clips are inserted into the corresponding positions of the basic video sequence; and Transition frames are generated at the junctions of the base video sequence and the plurality of handheld display video clips to form the target video.

9. A video generation apparatus, comprising: The acquisition module is configured to acquire user input information, which includes video duration information of the video to be generated, description information of the target object, and reference information, wherein the reference information is used to provide a target image of the target object; The parsing module is configured to parse the input information to generate a video script, wherein the video script includes a video theme field and a display method field for the target object; The first generation module is configured to generate a preview image of the first frame of the video based on the video script and the target image; The determination module is configured to, in response to the user's confirmation instruction on the preview image of the first frame of the video, determine the total number of video segments contained in the video to be generated based on the video script and the video duration information; The second generation module is configured to use a raw image model to generate multiple keyframes on the timeline based on the first frame preview image of the video and the total number of video segments. The third generation module is configured to generate sequential video segments based on each pair of keyframes that are adjacent on the time axis, using a video generation model; and The stitching module is configured to stitch together the video segments corresponding to each pair of adjacent keyframes in chronological order to generate the target video.

10. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.

12. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Method and terminal for displaying image and video and responding network request

    CN107105338A

  • Special effect video generation method and device, electronic equipment and storage medium

    CN119583736A