Video generation method, device, and medium
By generating a global guiding image and combining it with the motion information of the storyboard to constrain the video generation model, the problem of scene edge breakage and inconsistent environmental connection that occurs in multi-storyboard scenarios in AI video generation tools is solved, thereby improving the integrity and coherence of game videos.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YIDIAN LINGXI INFORMATION TECHNOLOGY (GUANGZHOU) CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-06-02
AI Technical Summary
Existing AI video generation tools are prone to glitches and inconsistencies in scene edge transitions and environmental continuity when generating game videos with multi-scene sequences, which reduces video quality.
By acquiring scene footage and first video generation information, a global guiding image is generated. Combined with the storyboard motion information, the video generation model is constrained to generate the target video, ensuring consistency between camera movement logic and scene.
It eliminates the problems of broken scene edges and disjointed environmental transitions, improves the integrity and continuity of game videos, and enhances video quality.
Smart Images

Figure CN122138016A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to a video generation method, apparatus, and medium. Background Technology
[0002] With the rapid development of AI video technology, the gaming industry can now leverage relevant AI video generation tools. For example, by directly inputting scene materials into an AI video generation tool, a complete generated video can be output, significantly improving the production efficiency of various game videos such as game promotional videos, game cutscenes, and game story short films, and effectively reducing the time and labor costs of traditional manual production.
[0003] However, existing AI video generation tools are prone to glitches and inconsistencies in scene edge transitions and environmental continuity when generating game videos with multi-scene sequences, which reduces the quality of the generated game videos. Summary of the Invention
[0004] One objective of this disclosure is to provide a new technical solution for video generation.
[0005] According to a first aspect of this disclosure, a video generation method is provided, comprising: Acquire scene materials and first video generation information; wherein, the first video generation information is used to describe the constituent elements of the video frame, and the first video generation information includes the motion information of the virtual camera. The scene material is processed based on the first video generation information to generate a global guide image that conforms to the first video generation information; wherein, the scene area covered by the global guide image includes multiple scene areas corresponding to the scene motion information; The first video generation model is invoked, and the video generation task is performed using at least the global guiding image and the storyboard motion information as generation constraints to obtain a target video that meets the generation constraints.
[0006] Optionally, the spatial range of the scene material is smaller than the scene area covered by the global guide image.
[0007] Optionally, the scene materials include scene original artwork and character portraits of the characters appearing on screen, and the first video generation information also includes storyboard composition information, which is used to determine the appearance information of the characters appearing on screen in the multiple storyboard areas; The step of processing the scene material based on the first video generation information to generate a global guide image that conforms to the first video generation information includes: The scene original drawing is expanded based on the storyboard motion information to generate an expanded scene image; wherein the expanded scene image covers the multiple storyboard areas; The character portrait is implanted into the target position of the extended scene image based on the on-screen information to generate the global guidance image; wherein the target position is determined based on the on-screen information.
[0008] Optionally, the first video generation information may also include text information. The step of processing the scene material based on the first video generation information to generate a global guide image that conforms to the first video generation information further includes: Generate text image elements based on the text information; The text image elements are embedded into the extended scene image to generate the global guide image.
[0009] Optionally, the generation constraints further include second video generation information; wherein the second video generation information is used to describe the expression form of the video frame constituent elements.
[0010] Optionally, the method further includes: Obtain third-party video generation information used to describe the audio information; If the third video generation information indicates that audio is generated in the first language, the first language model is invoked to generate candidate audio content information based on the third video generation information; If the third video generation information indicates that audio is generated in a second language, the second language model is invoked to generate candidate audio content information based on the third video generation information; The generation constraints also include the audio content information, which is used to guide the first video generation model to generate an audio file that matches the video frame when generating the video frame.
[0011] Optionally, the step of calling the first video generation model, using at least the global guiding image and the storyboard motion information as generation constraints, to perform a video generation task and obtain a target video that meets the generation constraints, includes: The first video generation model is invoked, and the video generation task is performed using at least the global guidance image and the storyboard motion information as generation constraints to obtain candidate videos that meet the generation constraints. The video quality of the candidate videos is evaluated to obtain the video quality evaluation results of the candidate videos; The target video is generated based on the video quality assessment results.
[0012] Optionally, generating the target video based on the video quality assessment result includes: If the video quality assessment result indicates that the candidate video is qualified, the candidate video will be used as the target video. If the video quality assessment result indicates that the candidate video is unqualified, the unqualified video segments in the candidate video are obtained; based on the video segments, at least one keyframe is determined from the candidate video; a second video generation model is invoked to reconstruct the video segments based on the at least one keyframe to obtain reconstructed video segments; and the reconstructed video segments are used to replace the video segments in the candidate video to obtain the target video.
[0013] According to a second aspect of this disclosure, an electronic device is provided, including a memory and a processor, the memory being configured to store a computer program, and the processor being configured to execute the method described according to the first aspect of this disclosure under the control of the computer program.
[0014] One beneficial effect of this disclosure is that, through the embodiments of this disclosure, a global guiding image covering the multi-scene area can be generated based on the scene material of the first video generation information processing, and the target video can be generated by combining the scene motion information to constrain the first video generation model. This method can accurately control the global scene information and standardize the camera motion logic, eliminating the problems of scene edge breakage and incoherent environmental connection in the dynamic multi-scene game video generation process, significantly improving the integrity and coherence of the game video, and improving the video quality of the generated game video.
[0015] The features and advantages of the embodiments of this specification will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments of this specification and, together with their description, serve to explain the principles of these embodiments.
[0017] Figure 1 A schematic diagram is shown that can be used to implement hardware configurations according to embodiments of the present disclosure; Figure 2 A flowchart illustrating a video generation method according to some embodiments is shown; Figure 3 A flowchart illustrating a video generation method according to some embodiments is shown; Figure 4 A schematic diagram of a scene concept art according to some embodiments is shown; Figure 5 It shows according to Figure 4 A schematic diagram of the global guide image generated from the original scene artwork shown; Figure 6A flowchart illustrating a video generation method according to some embodiments is shown; Figure 7 A schematic diagram of the composition structure of a video generation apparatus according to some embodiments is shown; Figure 8 A block diagram of an electronic device according to some embodiments is shown. Detailed Implementation
[0018] Various exemplary embodiments of this specification will now be described in detail with reference to the accompanying drawings.
[0019] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the embodiments of this specification or their application or use.
[0020] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0021] It should be noted that all actions involving the acquisition of signals, information, or data in this embodiment are carried out in compliance with the relevant data protection laws and regulations of the country where the location is situated, and with authorization from the owner of the relevant equipment.
[0022] This disclosure provides a video generation method. For ease of understanding, the application scenarios of the video generation method provided in this disclosure are illustrated below. Figure 1 As shown, Figure 1 This is a schematic diagram of the hardware structure of an electronic device for a video generation method provided in an embodiment of this disclosure.
[0023] Electronic device 1000 is a device capable of running computer programs. These computer programs can be local applications installed on the electronic device, or web applications, lightweight applications, or mini-programs, etc., without limitation. The electronic device 1000 can be a mobile phone, tablet computer, PC, etc., without limitation.
[0024] like Figure 1 As shown, the electronic device 1000 may include a processor 1101, a memory 1102, an interface device 1103, a communication device 1104, an output device 1105, an input device 1106, etc. Figure 1 The hardware configuration shown is illustrative only and is not intended to limit this disclosure, its application, or its use.
[0025] The processor 1101 executes computer programs, which can be written using instruction sets of architectures such as x86, Arm, RISC, MIPS, and SSE. The memory 1102 includes, for example, ROM (Read-Only Memory), RAM (Random Access Memory), and non-volatile memory such as a hard disk. The interface device 1103 includes, for example, a USB interface, a network cable interface, and a headphone jack. The communication device 1104 is capable of wired or wireless communication. The communication device 1104 may include at least one short-range communication module, such as any module for short-range wireless communication based on short-range wireless communication protocols such as Hilink, WiFi (IEEE 802.11), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, and LiFi. The communication device 1104 may also include a long-range communication module, such as any module for WLAN, GPRS, or 2G / 3G / 4G / 5G long-range communication. The output device 1105 may include, for example, an LCD screen or touch screen, and a speaker. Input device 1106 may include, for example, a touch screen, a keyboard, a microphone, various sensors, etc.
[0026] In this embodiment, the memory 1102 of the electronic device 1000 is used to store a computer program that controls the processor 1101 to operate in order to execute a video generation method according to any embodiment of the present disclosure.
[0027] The following is combined with Figure 1 The electronic device shown illustrates the video generation method provided in the embodiments of this disclosure. It should be noted that... Figure 1 This is merely one application scenario of the video generation method provided in this disclosure embodiment, and does not mean that the video generation method can only be applied to... Figure 1 The application scenarios shown.
[0028] It should be noted that the target video mentioned in the following embodiments usually refers to the game video used to set up the game project, generated by the video generation method based on the embodiments of this disclosure, such as game cutscenes, game promotional videos, or game story short films. The video generation method of the embodiments of this disclosure can at least solve the problem of scene errors during game video production.
[0029] <First Embodiment> Figure 2 The diagram illustrates a flow chart of a video generation method according to some embodiments, which is implemented by an electronic device, for example, the electronic device may be as follows: Figure 1 The electronic device 1000. The video generation method may include the following steps S210 to S230: Step S210: Obtain scene materials and first video generation information; wherein, the first video generation information is used to describe the constituent elements of the video frame, and the first video generation information includes the motion information of the virtual camera.
[0030] Scene materials are digital assets that are pre-prepared and fixed to form the basic visual elements of the video screen for the target video to be generated. Scene materials may include original scene artwork and character portraits of the characters appearing in the video.
[0031] The first video generation information is typically used to describe the constituent elements of the video frame. This first video generation information may include the motion information of the virtual camera's shots. This motion information may include pre-configured virtual camera motion parameters for each of the multiple shot regions divided along the complete timeline of the target video to be generated. These virtual camera motion parameters may include motion type, motion direction, motion trajectory, and motion speed.
[0032] Step S220: Process the scene material based on the first video generation information to generate a global guide image that conforms to the first video generation information; wherein, the scene area covered by the global guide image includes multiple storyboard areas corresponding to the storyboard motion information.
[0033] Specifically, the electronic device can process scene material based on the motion information of the virtual camera in the first video generation information to generate a global guide image containing multiple shot regions corresponding to the motion information of the shot. The global guide image can be a single guide image.
[0034] In this embodiment, by generating scene material based on the first video generation information containing segmentation motion information, a global guide image covering the entire segmentation area is generated. This can eliminate glitches such as virtual camera movement, scene edge breaks, and inconsistent environmental transitions that are prone to occur during segmentation switching, significantly improving the integrity and continuity of the game video.
[0035] Step S230: Invoke the first video generation model, and use at least the global guidance image and the storyboard motion information as generation constraints to perform the video generation task and obtain a target video that meets the generation constraints.
[0036] The first video generation model can be the SORA2 video generation model. Performing video generation tasks based on the SORA2 video generation model can maximize the visual consistency of the generated target video.
[0037] In this embodiment, the input to the first video generation model includes a global guide image and scene motion information. When performing the video generation task, the first video generation model uses the global guide image as a visual reference to anchor the spatial layout and environmental details of the scene, while using the scene motion information as a camera logic constraint to strictly regulate the motion type, direction, trajectory, and speed of the virtual camera, thereby controlling the rhythm of scene switching and the continuity of camera movement. Through the synergistic effect of the above dual constraints, the first video generation module can output a target video with smooth camera movement trajectory, highly consistent scene visual features in each scene area, and strong scene stability, effectively avoiding scene continuity errors under dynamic shots.
[0038] Through the embodiments of this disclosure, a global guiding image covering a multi-scene area can be generated based on the scene material of the first video generation information processing, and the target video can be generated by combining the scene motion information to constrain the first video generation model. This method can accurately control the global scene information and standardize the camera motion logic, eliminating the glitches such as scene edge breakage and inconsistent environmental connection in the process of dynamic multi-scene game video generation, significantly improving the integrity and continuity of the game video, and improving the video quality of the generated game video.
[0039] <Second Embodiment> In this embodiment, the spatial range of the scene footage is typically smaller than the scene area covered by the global guide image. In other words, it can expand a global guide image larger than the spatial range of the scene footage based on a small area of scene footage. Specifically, the scene area covered by the global guide image not only includes multiple scene regions corresponding to the scene motion information, but also the scene area covered by the global guide image is larger than the spatial range of the scene footage.
[0040] Based on this, on the one hand, it can completely avoid the scene edge breakage and background loss caused by virtual camera panning and zooming exceeding the material boundaries at the base level. On the other hand, it eliminates the need to consume a lot of resources to collect and produce large-scale complete scene materials; it can support the generation of multi-shot videos with only small-scale scene materials, significantly reducing the cost of material acquisition and preprocessing. At the same time, the extended area of the global guide image strictly follows the visual characteristics of the original scene materials, ensuring the style consistency of the entire scene. The elimination of the aforementioned continuity errors, the control of material costs, and the consistency of scene style can effectively improve the integrity and consistency of the target video, thereby improving the video quality of the generated target video.
[0041] <Third Embodiment> In related technologies, AI video generation tools often encounter issues such as character consistency defects and failure of multiple characters appearing in the same shot when generating game videos for a given game project. Character consistency defects manifest as characters easily distorting and shifting in position within the frame, with low similarity in the same character's appearance across different shots, making cross-shot image reuse difficult. Failure of multiple characters appearing in the same shot manifests as the inability to stably present two or more characters in the same frame, easily leading to character distortion and spatial misalignment. In this embodiment, the image, clothing, and spatial position of each character can be pre-embedded in the extended scene image, fundamentally ensuring cross-shot consistency of character appearance. Furthermore, when there are multiple characters, their layout and embedding in the extended scene image can be completed according to the storyboard composition information, achieving stable multi-character simultaneous appearance. This effective guarantee of character image stability and multi-character simultaneous appearance effectively avoids problems such as character drift, distortion, or disappearance during camera movement, significantly improving the narrative coherence and visual expressiveness of the video, thereby enhancing the video quality of the generated target video.
[0042] In these embodiments, compared to the first embodiment described above, the scene materials may specifically include original scene artwork and character portraits of the characters appearing in the scene.
[0043] Scene concept art can be background environment images drawn in advance by artists for the game's storyline. Scene concept art can usually include a complete background structure, prop layout (such as tables, chairs, and buildings), and environmental details (such as lighting, vegetation, and decorations), without obvious cropping or missing edges.
[0044] It should be noted that if the artists have not pre-drawn scene concept art that matches the game's storyline, a stylistic transfer method based on existing materials can be used to obtain scene concept art for the game project. Specifically, an initial scene image similar to the game's storyline in terms of spatial layout, main content, or atmosphere can be obtained first. Subsequently, the initial scene image undergoes image style transfer processing, image editing, or image redrawing to adjust its visual style to match the established art style of the game's storyline, thereby generating scene concept art adapted to the game project.
[0045] The character portrait of the on-screen character can include either a character description or an image of the character. The character description can include descriptions of the character's physical features and clothing / equipment. The character image can include three-view drawings or dynamic reference images of the character. The three-view drawings can be a front view, side view, and back view of the character, clearly showing the character's basic shape and clothing details. The dynamic reference images can be a 360° panoramic video showcasing the character, providing information on how the character's appearance changes from different perspectives.
[0046] In these embodiments, relative to the first embodiment described above, the first video generation information may further include storyboard composition information, which is used to determine the appearance information of the on-screen character in multiple storyboard areas. The storyboard composition information includes, for example, the character's position information and the character's posture information within the storyboard area.
[0047] Figure 3 The diagram illustrates a video generation method according to some embodiments. In these embodiments, relative to the first embodiment described above, the above-described generation of a global guide image conforming to the first video generation information by processing scene material may include the following steps S310 to S320: Step S310: Expand the scene original drawing based on the storyboard motion information to generate an expanded scene image.
[0048] In this context, the extended scene image covers multiple storyboard areas, meaning that the scene area of the extended scene image can usually completely cover all storyboard areas corresponding to the storyboard motion information.
[0049] In this embodiment, the electronic device can invoke a first image editing tool to perform image expansion processing on the original scene artwork based on the storyboard motion information, generating an expanded scene image covering multiple storyboard areas. Specifically, the electronic device can first analyze the motion type, motion direction, motion trajectory, and motion speed in the storyboard motion information to determine the entire viewport range that the virtual camera needs to cover during dynamic movements such as zooming out, panning, and switching; then, using the original scene artwork as the core base, it extends the scene boundary outward according to the virtual camera panning direction, expands the scene spatial depth according to the virtual camera zooming out trajectory, and connects the scene segments of the corresponding storyboards according to the virtual camera switching sequence to generate the expanded scene image.
[0050] For example, Figure 4 The scene concept art shown is the basic material for video generation, containing only visual elements of the core storyboard areas such as the castle, three rows of trees, and grass. Expanding scene concept art 41 yields... Figure 5The expanded scene image 51 shown may include the following expansions: adding more trees to the left of the existing trees to form a continuous forest, adding a winding river to the right of the image, and adding a continuous mountain range at the top of the image as a distant view. The expansion process strictly preserves the core visual features of the original scene artwork 41 (such as the castle layout). The final expanded scene image 51 can completely cover the entire viewport range required for dynamic movements such as virtual camera panning and zooming, providing a unified global visual benchmark for subsequent video generation and thus avoiding scene continuity errors during camera movement.
[0051] Step S320: Based on the on-screen information, the character portrait is implanted into the target position of the extended scene image to generate a global guide image.
[0052] The target location can be determined based on the aforementioned on-screen information.
[0053] In this embodiment, the electronic device further uses the on-screen information to call a second image editing tool to accurately embed the character portraits of each on-screen character into the target position in the extended scene image. Specifically, the electronic device can first parse the position information in the on-screen information to determine the coordinate parameters of the on-screen character in the extended scene image, and parse the posture information in the on-screen information to determine the action description of the on-screen character in the extended scene image; then, using the extended scene image as a spatial base, the character portraits are embedded into the target position in the extended scene image according to the coordinate parameters in the position information, and the limb shape of the character portraits is adapted according to the action description in the posture information to generate a global guidance image.
[0054] It should be noted that if the character image is character description information, the character description information needs to be converted into a character image first, and then the character image needs to be implanted into the target position of the extended scene image.
[0055] The first image editing tool and the second image editing tool can be the same tool or different tools.
[0056] <Fourth Embodiment> In related technologies, AI video generation tools suffer from instability issues with Chinese elements in game scenes when generating game videos for specific game projects. Specifically, this manifests as garbled characters and distorted forms in Chinese signage, prop text, and environmental slogans during camera movement, affecting the video's information delivery and visual experience. In this embodiment, the Chinese text in the scene can be pre-defined as inherent image elements within the extended scene image, achieving stable presentation of Chinese elements. When the camera changes, the Chinese elements in the scene will change synchronously with the overall image, eliminating the need for the video generation model to generate or redraw text. This fundamentally avoids the garbled characters and distortions caused by independent rendering of Chinese elements during video generation. This precise control over the presentation of Chinese elements effectively eliminates visual defects, ensuring the clarity and integrity of Chinese information in the video, while maintaining the coordination between Chinese elements and the image during camera movement. This significantly optimizes the overall visual appeal of the video, thereby achieving the desired video quality.
[0057] In these embodiments, compared to the third embodiment described above, the first video generation information may also include text information.
[0058] The text information can be the target text in the target video to be generated, as well as the visual attribute information of that target text. The target text can be a text element in the scene, such as a Chinese character element, which can be a Chinese sign, prop text, environmental slogan, etc. The visual attribute information can include the position, font, font size, color, and transparency of the text element.
[0059] Unlike the third embodiment described above, in these embodiments, relative to the third embodiment described above, step S220, which processes scene material based on the first video generation information to generate a global guide image that conforms to the first video generation information, may further include: generating text image elements based on text information; and embedding the text image elements into an extended scene image to generate a global guide image.
[0060] In this embodiment, the electronic device further uses a third image editing tool to generate text image elements based on the text information, and then precisely embeds these text image elements into the extended scene image to generate a global guidance image. Specifically, the electronic device can first parse the target text in the text information to determine the text content to be embedded, and parse the visual attribute information in the text information to determine the font style, font size, color parameters, layout format, and presentation effect of the target text; then, using vector rendering or bitmap rendering, the target text is converted into text image elements with independent layer attributes according to the parameter requirements of the visual attribute information; and using the extended scene image as the spatial base, the text image elements are embedded into the extended scene image according to the position parameters in the visual attribute information to generate a global guidance image.
[0061] The first image editing tool, the second image editing tool, and the third image editing tool can be the same tool or different tools.
[0062] <Fifth Embodiment> In related technologies, AI video generation tools typically lack a unified art style constraint when generating game videos for a given game project. This can easily lead to stylistic dissonance issues such as deviations in scene texture and inconsistencies between character and scene art styles, thus compromising visual consistency and immersion. In this embodiment, art style information can be used as another core constraint for video generation. This ensures that scenes conform to a predetermined art style, preventing stylistic gaps during camera movements and guaranteeing visual consistency and immersion. This global and consistent constraint on the video's art style effectively avoids stylistic dissonance issues that can easily occur during camera movements or scene transitions, significantly enhancing the video's visual consistency and viewing immersion, optimizing the overall visual presentation, and ultimately achieving the desired video quality.
[0063] In these embodiments, compared to the first embodiment described above, the generation constraints may further include second video generation information. The second video generation information can be used to describe the expression of the elements that make up the video frame. This expression of the elements that make up the video frame can be understood as art style information. Using art style information as a generation constraint for generating the target video can ensure that all elements that make up the target video, such as scenes and characters, fit the preset art style. Furthermore, it can prevent style discontinuities during camera movement and shot switching, ultimately ensuring the visual unity and immersive viewing experience of the target video.
[0064] <Sixth Embodiment> In related technologies, AI video generation tools typically operate on an independent basis, generating video footage and audio separately when creating game videos for a given game project. Specifically, video footage is generated first, followed by separate audio production and synthesis. Audio generated under this model, especially in Chinese, often suffers from stiff intonation, illogical phrasing, and poor naturalness. In this embodiment, video footage generation and audio generation are performed simultaneously, achieving synchronized audio-visual generation. The generated audio-visual content precisely matches the video footage, ensuring that the tone and rhythm of the voice align with the character actions and camera transitions within the video. This synchronized audio-visual generation effectively solves the problems of audio-visual asynchrony and rhythmic disconnect that can occur with traditional post-production dubbing, greatly enhancing the narrative expressiveness and emotional delivery of the video, significantly improving the overall audiovisual effect, and thus improving the video quality of the generated target video.
[0065] In these embodiments, compared to the first embodiment described above, the video generation method of this disclosure further includes: obtaining third video generation information for describing audio information; when the third video generation information indicates that audio is generated in a first language, calling a first language model to generate candidate audio content information based on the third video generation information; and when the third video generation information indicates that audio is generated in a second language, calling a second language model to generate candidate audio content information based on the third video generation information.
[0066] In these embodiments, compared to the first embodiment described above, the generation constraints may further include audio content information, which is used to guide the first video generation model to synchronously generate an audio file that matches the video frame when generating the video frame.
[0067] Specifically, in addition to performing the core steps of scene material acquisition, global guide image generation, and video generation in the first embodiment, third video generation information is also acquired. This third video generation information is used to define the constituent elements and generation specifications of the audio accompanying the target video to be generated. For example, it may include the audio language type (first language and second language), emotional tone, and duration. The first language is set to English, and a third-party language model (such as ElevenLabs) independent of the first video generation model is called; the second language is set to Chinese, and a native language model integrated into the first video generation model is called. This differentiated model adaptation ensures the quality of audio generation.
[0068] Based on the language type indicated by the third-party video generation information, the corresponding model is invoked to generate candidate audio content information. English audio relies on a third-party model to achieve natural intonation and emotional expression, while Chinese audio utilizes a native model to ensure standard pronunciation and seamlessness, controlling audio quality from the source. Audio content information is incorporated into the generation constraints, working together with global guidance images and scene motion information to constrain the first video generation model. This achieves synchronized generation of video and audio, ensuring precise matching of audio rhythm and emotion with character actions and shot transitions, avoiding audio-visual disconnect, strengthening the unity of video audiovisuals, and thus significantly improving the video quality of the generated target video.
[0069] <Seventh Embodiment> In this embodiment, the first video generation model can generate candidate videos based on static guide images, focusing on overall coherence and shot control. The second video generation model can accurately repair local defects in the candidate videos through keyframe localization, excelling in high-fidelity rendering of short clips and inter-frame interpolation. The two models work collaboratively and complement each other's strengths, controlling overall computational costs while specifically improving video quality. Through this hybrid workflow, the overall video quality and generation controllability achieved are difficult to attain with a single model, ultimately improving the video quality of the generated target video.
[0070] Figure 6 The diagram illustrates a video generation method according to some embodiments. In these embodiments, relative to the first embodiment described above, the invocation of the first video generation model, using at least the global guiding image and storyboard motion information as generation constraints, to perform the video generation task and obtain a target video that meets the generation constraints may include the following steps S610 to S630: Step S610: Call the first video generation model, using at least the global guide image and storyboard motion information as generation constraints, to perform the video generation task and obtain candidate videos that meet the generation constraints.
[0071] The first video generation model is used to quickly output complete candidate videos that conform to the global guiding image, and usually does not need to pursue excessive detail accuracy.
[0072] Step S620: Evaluate the video quality of the candidate videos to obtain the video quality evaluation results of the candidate videos.
[0073] In this embodiment, the visual performance of the candidate video can be comprehensively evaluated by manual assessment combined with preset video quality assessment dimensions to obtain the video quality assessment result of the candidate video.
[0074] The aforementioned preset video quality evaluation dimensions may include scene coherence, character consistency, text stability, and art style uniformity.
[0075] Step S630: Generate the target video based on the video quality assessment results.
[0076] In one example, step S630, which generates the target video based on the video quality assessment result, can be implemented in the following way: if the video quality assessment result indicates that the candidate video is of acceptable quality, the candidate video is directly used as the target video.
[0077] In another example, step S630, which generates the target video based on the video quality assessment result, can be implemented as follows: if the video quality assessment result indicates that the candidate video is of unqualified quality, obtain the unqualified video segments from the candidate video; determine at least one keyframe from the candidate video based on the video segments; input the at least one keyframe into the second video generation model to generate a reconstructed video segment; replace the video segments in the candidate video with the reconstructed video segment to obtain the target video.
[0078] The second video generation model focuses on local quality optimization and does not participate in the construction of the overall video framework; it only repairs local defects in the candidate video. The second video generation model can be either Wanxiang or Keling.
[0079] Keyframes can be one or more frames of relatively high visual quality extracted from candidate videos. Keyframes typically accurately represent the basic visual features and spatiotemporal boundaries of the corresponding video segment, providing a reliable benchmark for local repair by the second video generation model. That is, keyframes can be located within unqualified video segments or within other qualified segments of the candidate video.
[0080] In this example, the specific repair mode can be flexibly selected based on the number of keyframes extracted. For instance, if both a qualified start and end keyframe can be extracted simultaneously, a start-end frame-driven generation model can be used. This model relies on the two frames to accurately define the spatiotemporal boundaries of the defective segment, enabling inter-frame interpolation and quality optimization within the segment. Alternatively, if only a suitable start keyframe can be extracted, a single-frame guided optimization mode can be used. This mode uses the visual features of the keyframe as anchor points to complete the local repair of the defective segment.
[0081] <Eighth Embodiment> This embodiment provides a video generation device. Figure 7 A schematic diagram of the composition of a video generation apparatus according to an embodiment of the present disclosure is shown. Figure 7 As shown, the video generation device 700 may include an acquisition module 710, a first generation module 720, and a second generation module 730.
[0082] The acquisition module 710 is used to acquire scene materials and first video generation information; wherein, the first video generation information is used to describe the constituent elements of the video frame, and the first video generation information includes the motion information of the virtual camera. The first generation module 720 is used to process the scene material based on the first video generation information to generate a global guide image that conforms to the first video generation information; wherein, the scene area covered by the global guide image includes multiple segmentation areas corresponding to the segmentation motion information. The second generation module 730 is used to call the first video generation model, and at least use the global guide image and the storyboard motion information as generation constraints to perform a video generation task to obtain a target video that meets the generation constraints.
[0083] In some embodiments, the spatial extent of the scene material is smaller than the scene area covered by the global guide image.
[0084] In some embodiments, the scene materials include scene original artwork and character portraits of the characters appearing on screen, and the first video generation information further includes storyboard composition information, which is used to determine the appearance information of the characters appearing on screen in the multiple storyboard areas; The second generation module 720 is specifically used for: expanding the scene original drawing according to the storyboard motion information to generate an expanded scene image; wherein the expanded scene image covers the multiple storyboard areas; and embedding the character portrait into the target position of the expanded scene image according to the appearance information to generate the global guidance image; wherein the target position is determined based on the appearance information.
[0085] In some embodiments, the first video generation information further includes text information, and the second generation module 720 is specifically used to: generate text image elements based on the text information; and embed the text image elements into the extended scene image to generate the global guide image. In some embodiments, the generation constraints further include second video generation information; wherein the second video generation information is used to describe the expression form of the video frame constituent elements. In some embodiments, the acquisition module 710 is further configured to acquire third video generation information for describing audio information; The third generation module 730 is further configured to, when the third video generation information indicates that audio is generated in a first language, call a first language model to generate candidate audio content information based on the third video generation information; and when the third video generation information indicates that audio is generated in a second language, call a second language model to generate candidate audio content information based on the third video generation information. The generation constraints also include the audio content information, which is used to guide the first video generation model to generate an audio file that matches the video frame when generating the video frame. In some embodiments, the third generation module 730 is specifically configured to: invoke a first video generation model, using at least the global guiding image and the storyboard motion information as generation constraints, perform a video generation task, and obtain candidate videos that meet the generation constraints; evaluate the video quality of the candidate videos to obtain video quality evaluation results of the candidate videos; and generate the target video based on the video quality evaluation results. In some embodiments, the third generation module 730 is specifically configured to: if the video quality assessment result indicates that the candidate video is qualified, use the candidate video as the target video; if the video quality assessment result indicates that the candidate video is unqualified, obtain unqualified video segments from the candidate video; determine at least one keyframe from the candidate video based on the video segments; call a second video generation model to reconstruct the video segments based on the at least one keyframe to obtain reconstructed video segments; and replace the video segments in the candidate video with the reconstructed video segments to obtain the target video. <Electronic Device Examples> This embodiment provides an electronic device for implementing a video generation method according to any embodiment of this disclosure. Figure 8 The basic hardware components of this electronic device are shown. For example... Figure 8 As shown, the electronic device 800 includes a processor 810 and a memory 820. The memory 820 stores a computer program that controls the processor 810 to operate in order to control the electronic device 800 to execute a video generation method according to any embodiment of the present disclosure.
[0086] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a video generation method according to any embodiment of this disclosure.
[0087] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0088] This disclosure may be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement any of the methods in the foregoing embodiments of this disclosure.
[0089] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media may include, for example, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), compact disc-read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any combination thereof. The computer-readable storage medium used herein is not to be interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0090] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include one or more of copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to computer-readable storage media in the respective computing / processing device.
[0091] The computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source or object programs written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network (e.g., a local area network or a wide area network), or it may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays, or programmable logic arrays, can execute computer-readable program instructions to implement various aspects of the embodiments of this disclosure by utilizing state information from the computer-readable program instructions.
[0092] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0093] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0094] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It should be noted that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are all equivalent.
[0096] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of this disclosure is defined by the appended claims.
Claims
1. A video generation method, wherein, include: Acquire scene materials and first video generation information; wherein, the first video generation information is used to describe the constituent elements of the video frame, and the first video generation information includes the motion information of the virtual camera. The scene material is processed based on the first video generation information to generate a global guide image that conforms to the first video generation information; wherein, the scene area covered by the global guide image includes multiple scene areas corresponding to the scene motion information; The first video generation model is invoked, and the video generation task is performed using at least the global guiding image and the storyboard motion information as generation constraints to obtain a target video that meets the generation constraints.
2. The method according to claim 1, wherein, The spatial range of the scene material is smaller than the scene area covered by the global guide image.
3. The method according to claim 1, wherein, The scene materials include scene concept art and character portraits of the characters appearing on screen. The first video generation information also includes storyboard composition information, which is used to determine the appearance information of the characters appearing on screen in the multiple storyboard areas. The step of processing the scene material based on the first video generation information to generate a global guide image that conforms to the first video generation information includes: The scene original drawing is expanded based on the storyboard motion information to generate an expanded scene image; wherein the expanded scene image covers the multiple storyboard areas; The character portrait is implanted into the target position of the extended scene image based on the on-screen information to generate the global guidance image; wherein the target position is determined based on the on-screen information.
4. The method according to claim 3, wherein, The first video generation information also includes text information. The step of processing the scene material based on the first video generation information to generate a global guide image that conforms to the first video generation information further includes: Generate text image elements based on the text information; The text image elements are embedded into the extended scene image to generate the global guide image.
5. The method according to claim 1, wherein, The generation constraints also include second video generation information; wherein the second video generation information is used to describe the expression form of the video frame constituent elements.
6. The method according to claim 1, wherein, The method further includes: Obtain third-party video generation information used to describe the audio information; If the third video generation information indicates that audio is generated in the first language, the first language model is invoked to generate candidate audio content information based on the third video generation information; If the third video generation information indicates that audio is generated in a second language, the second language model is invoked to generate candidate audio content information based on the third video generation information; The generation constraints also include the audio content information, which is used to guide the first video generation model to generate an audio file that matches the video frame when generating the video frame.
7. The method according to claim 1, wherein, The step of calling the first video generation model, using at least the global guidance image and the storyboard motion information as generation constraints, to perform a video generation task and obtain a target video that meets the generation constraints, includes: The first video generation model is invoked, and the video generation task is performed using at least the global guidance image and the storyboard motion information as generation constraints to obtain candidate videos that meet the generation constraints. The video quality of the candidate videos is evaluated to obtain the video quality evaluation results of the candidate videos; The target video is generated based on the video quality assessment results.
8. The method according to claim 6, wherein, The step of generating the target video based on the video quality assessment result includes: If the video quality assessment result indicates that the candidate video is qualified, the candidate video will be used as the target video. If the video quality assessment result indicates that the candidate video is unqualified, the unqualified video segments in the candidate video are obtained; based on the video segments, at least one keyframe is determined from the candidate video; a second video generation model is invoked to reconstruct the video segments based on the at least one keyframe to obtain reconstructed video segments; and the reconstructed video segments are used to replace the video segments in the candidate video to obtain the target video.
9. An electronic device, wherein, It includes a memory and a processor, the memory being used to store a computer program, and the processor being used, under the control of the computer program, to execute the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, wherein, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the method according to any one of claims 1 to 8.