Video generation method and device, equipment, storage medium and product
By generating video information from multiple different perspectives and angles using a large model, the problem of single-viewpoint generation in existing technologies is solved, and a video generation method that quickly selects the best viewpoint is realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING QIHOOD TECHNOLOGY CO LTD
- Filing Date
- 2026-02-05
- Publication Date
- 2026-04-14
AI Technical Summary
Existing video generation technologies can only generate videos from a single perspective or shot, requiring users to re-enter prompts to obtain videos from different perspectives, resulting in a long creative verification cycle.
The system generates scene information corresponding to scene description prompts using a large model, and generates multiple video information with different shot sizes and/or angles based on this information, displaying these videos for users to choose the best perspective.
It enables the generation of multi-view videos, allowing users to quickly compare the visual effects of different shots and angles, shortening the creative verification cycle and improving generation efficiency.
Smart Images

Figure CN121865064A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of AI video generation technology, and in particular to video generation methods, apparatus, devices, storage media and products. Background Technology
[0002] Currently, with the development of generative artificial intelligence, text-to-video technology can directly generate video content from natural language descriptions (e.g., video generation methods based on diffusion models). A common approach involves the user providing a scene description cue, and the model directly generating a single video clip based on that cue. However, the generated video content typically corresponds to only one fixed camera angle or shot (e.g., a default medium frontal shot). If the user wants to see a video from a different angle, they need to draw a new card. Summary of the Invention
[0003] The main objective of this application is to provide a video generation method, apparatus, device, storage medium, and product, which aims to solve the technical problem that existing video generation methods can only generate a single perspective or scene.
[0004] To achieve the above objectives, this application proposes a video generation method, which includes: In response to the input scene description prompts, scene information corresponding to the scene description prompts is generated through a large model; Based on the scene information, the large model generates video information corresponding to multiple different scene types and / or angles; The generated video information will be displayed.
[0005] Optionally, the step of generating multiple video information corresponding to different scene sizes and / or angles based on the scene information using the large model includes: Multiple target scene types and / or angle combinations are generated based on the scene information; Video information is generated using the large model based on the scene information and the target shot size and / or angle combination.
[0006] Optionally, generating multiple target scene types and / or angle combinations based on the scene information includes: Multiple initial shot types and / or angle combinations are generated based on the scene information; The initial shot size and / or angle combination is randomly perturbed to generate the target shot size and / or angle combination.
[0007] Optionally, the initial shot size and / or angle combination includes video shot size and video angle; The step of randomly perturbing the initial shot size and / or angle combination to generate the target shot size and / or angle combination includes: The target scene is obtained by randomly perturbing the video scene based on a preset scene neighborhood range. And / or, The video angle is randomly perturbed to obtain the target angle, and the target shot and / or angle combination is determined based on the initial shot and / or angle combination, the target shot or the target angle.
[0008] Optionally, before randomly perturbing the video scene based on a preset scene neighborhood range to obtain the target scene, the method further includes: Determine the candidate shot sizes in the preset shot candidate library; The perturbable shot type is determined based on the shot type characteristics of the candidate shot types; A preset scene neighborhood range is generated based on the perturbable scene size.
[0009] Optionally, the step of randomly perturbing the video angle to obtain the target angle includes: Based on the scene information, scene features are determined, and the scene features include at least one of the following: scene environment type, entity action information, spatial scale, and scene emotion information; The target angle is obtained by randomly perturbing the video angle based on the scene characteristics.
[0010] Optionally, the step of randomly perturbing the video angle based on the scene features to obtain the target angle includes: Determine the video angle preference range corresponding to the scene information based on the scene features; The target angle is obtained by randomly perturbing the video angle based on the video angle preference range.
[0011] Optionally, after displaying the generated multiple video information, the process further includes: Upon receiving a video regeneration instruction, the target video information and lens adjustment parameters corresponding to the video regeneration instruction are obtained; Determine the video shot and video angle corresponding to the target video information; The target lens parameters are determined based on the video shot type, the video angle, and the lens adjustment parameters. Based on the scene information, the large model generates video information corresponding to the target lens parameters.
[0012] Optionally, in response to the input scene description prompt, the process of generating scene information corresponding to the scene description prompt through a large model includes: In response to the input scene description prompts, the large model determines the video generation elements based on the scene description prompts; Based on the video generation elements, the three-dimensional spatial structure of the video scene is laid out using the large model to obtain scene information corresponding to the scene description prompts.
[0013] Optionally, the scene information is three-dimensional spatial scene information, which includes at least one of the following: the spatial position of the main body in the scene, the geometric layout data of the scene, the spatial relationship constraints between scene entities, and the lighting direction of the scene.
[0014] Optionally, displaying the generated multiple video information includes: The number of videos is determined based on the video information; The video display layout information is determined based on the number of videos; The video information is displayed according to the video display layout information.
[0015] Optionally, displaying the video information according to the video display layout information includes: Based on the video display layout information, the video information is synthesized according to spatial location to obtain a target video that simultaneously displays multiple video information; The target video is time-aligned by aligning multiple videos within it, and the time-aligned target video is then displayed.
[0016] Furthermore, to achieve the above objectives, this application also proposes a video generation apparatus, which includes: The response module is used to respond to the input scene description prompts and generate scene information corresponding to the scene description prompts through the large model; The video generation module is used to generate multiple video information corresponding to different scene types and / or angles based on the scene information and the large model; The display module is used to showcase multiple generated video information items.
[0017] Optionally, the video generation module is further configured to generate multiple target shot types and / or angle combinations based on the scene information; Video information is generated using the large model based on the scene information and the target shot size and / or angle combination.
[0018] Optionally, the video generation module is further configured to generate multiple initial shot sizes and / or angle combinations based on the scene information; The initial shot size and / or angle combination is randomly perturbed to generate the target shot size and / or angle combination.
[0019] Optionally, the initial shot size and / or angle combination includes a video shot size and a video angle; the video generation module is further configured to randomly perturb the video shot size based on a preset shot size neighborhood range to obtain a target shot size; And / or, The video angle is randomly perturbed to obtain the target angle, and the target shot and / or angle combination is determined based on the initial shot and / or angle combination, the target shot or the target angle.
[0020] Optionally, the video generation module is further configured to determine candidate shot sizes in a preset shot candidate library; The perturbable shot type is determined based on the shot type characteristics of the candidate shot types; A preset scene neighborhood range is generated based on the perturbable scene size.
[0021] Optionally, the video generation module is further configured to determine scene features based on the scene information, wherein the scene features include at least one of the following: scene environment type, entity action information, spatial scale, and scene emotion information; The target angle is obtained by randomly perturbing the video angle based on the scene characteristics.
[0022] Optionally, the video generation module is further configured to determine the video angle preference range corresponding to the scene information based on the scene features; The target angle is obtained by randomly perturbing the video angle based on the video angle preference range.
[0023] Optionally, the display module is further configured to, upon receiving a video regeneration instruction, obtain the target video information and lens adjustment parameters corresponding to the video regeneration instruction; Determine the video shot and video angle corresponding to the target video information; The target lens parameters are determined based on the video shot type, the video angle, and the lens adjustment parameters. Based on the scene information, the large model generates video information corresponding to the target lens parameters.
[0024] Optionally, the response module is further configured to, in response to the input scene description prompt, determine video generation elements based on the scene description prompt using a large model; Based on the video generation elements, the three-dimensional spatial structure of the video scene is laid out using the large model to obtain scene information corresponding to the scene description prompts.
[0025] Optionally, the scene information is three-dimensional spatial scene information, which includes at least one of the following: the spatial position of the main body in the scene, the geometric layout data of the scene, the spatial relationship constraints between scene entities, and the lighting direction of the scene.
[0026] Optionally, the display module is further configured to determine the number of videos based on the video information; The video display layout information is determined based on the number of videos; The video information is displayed according to the video display layout information.
[0027] Optionally, the display module is further configured to synthesize the video information according to the spatial position based on the video display layout information to obtain a target video that displays multiple video information simultaneously; The target video is time-aligned by aligning multiple videos within it, and the time-aligned target video is then displayed.
[0028] In addition, to achieve the above objectives, this application also proposes a video generation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the video generation method as described above.
[0029] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the video generation method described above.
[0030] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the video generation method described above.
[0031] In response to input scene description prompts, this application generates scene information corresponding to the scene description prompts through a large model; based on the scene information, it generates multiple video information corresponding to different shot sizes and / or angles through the large model; and displays the generated multiple video information. Compared with existing video content that only targets a fixed camera angle or shot size, this application can generate multiple video information with different shot sizes and / or angles, allowing users to compare the visual effects of different shot sizes and / or angles, quickly select the best shot size and / or angle, and shorten the creative verification cycle. Attached Figure Description
[0032] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0033] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 This is a flowchart illustrating an embodiment of the video generation method of this application. Figure 2 This is a schematic diagram illustrating the nine-grid video information display effect provided in Embodiment 1 of the video generation method of this application; Figure 3 This is a schematic diagram illustrating the display effect of a four-grid video information provided in Embodiment 1 of the video generation method of this application; Figure 4 This is a flowchart illustrating Embodiment 2 of the video generation method of this application; Figure 5 This is a flowchart illustrating Embodiment 3 of the video generation method of this application; Figure 6 This is a schematic diagram of the module structure of the video generation device according to an embodiment of this application; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the video generation method in this application embodiment.
[0035] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0036] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0037] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0038] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or video generation device capable of performing the above functions. The following description uses a video generation device as an example to illustrate this embodiment and the subsequent embodiments.
[0039] Based on this, embodiments of this application provide a video generation method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the video generation method of this application, provided in Embodiment 1.
[0040] In this embodiment, the video generation method includes the following steps: Step S10: In response to the input scene description prompt, generate scene information corresponding to the scene description prompt through the large model; It should be noted that the scene description prompts can be user-inputted descriptions of the desired video content, and can be in text or audio format. Scene description prompts typically include information such as characters, actions, environment, and emotions. For example, in the prompt "Generate a video of a cat sweeping the floor at home," the scene information can include semantic information, spatial layout, and visual attributes, providing a unified reference for subsequent multi-view generation. The semantic information can include subject category, role, action behavior, environment type, time and weather, and emotional atmosphere; the spatial layout can include the size of the scene; and the visual attributes can include the direction of light sources, ambient light color, and weather effect parameters in the video. The large model can be a deep learning model with video generation capabilities.
[0041] Step S20: Based on the scene information, generate multiple video information corresponding to different scene sizes and / or angles using the large model; It should be noted that the shot size can be categorized by the size of the camera's field of view, commonly including long shot, full shot, medium shot, close-up, and extreme close-up. The shot size determines the proportion of the subject in the frame and the amount of environmental information. Angle can refer to the camera's orientation and height relative to the subject, including: camera height: used to determine the camera's vertical position relative to the ground; horizontal azimuth (horizontal orientation): front, side, back, oblique, etc.; vertical elevation (vertical height): eye-level, upward, downward, bird's-eye view, etc. In this embodiment, the angle can also be described in specific degrees. For example, for the horizontal azimuth, which is the angle of the camera relative to the subject on a horizontal plane, it is usually measured clockwise or counterclockwise with the front being 0°. For example: 0° represents the front; 90° represents the right side; 180° or -180° represents the back; 45° or 135° represents the oblique side. The horizontal azimuth of the camera relative to the subject can be adjusted by adjusting the degree. For the vertical elevation angle, it can be the camera's height angle relative to the subject on a vertical plane, which can be 0° for eye-level, positive for upward, and negative for downward. For example, a level view corresponds to a vertical tilt angle of 0°, while 30° is a 30° upward tilt, meaning looking up from below at a 30° angle. Video information corresponding to different shot sizes and / or angles can be a video segment generated from different combinations of shot sizes and / or angles. For example, video information can include: a video clip with a long shot and a frontal angle, a video clip with a close-up shot and an oblique angle, and a video clip with a close-up and a frontal angle. This also includes the camera height, adjusted based on the subject's height or the visual effect during the specific shooting.
[0042] Step S30: Display the generated video information.
[0043] It should be noted that displaying the generated multiple video information can mean presenting the video information to the user in a visual form, such as arranging and playing it in a grid layout (such as a nine-square grid) on the same interface, or providing a function to switch and preview multiple video information.
[0044] Furthermore, in order to enable users to see video information corresponding to different shot sizes and / or angles at the same time, and to facilitate users to select the most satisfactory viewpoint, step S30 may include: determining the number of videos based on the video information; The video display layout information is determined based on the number of videos; The video information is displayed according to the video display layout information.
[0045] It should be noted that the number of videos can be the number of video information corresponding to different shot sizes and / or angles. In this embodiment, the video information can be displayed simultaneously using a grid layout. Determining the video display layout information based on the number of videos can be based on determining the screen division method in the grid layout according to the number of videos. For example, if there are 9 videos, they can be displayed in a nine-grid format; if there are 2 videos, they can be displayed in one screen by arranging them vertically or horizontally.
[0046] Furthermore, displaying the video information according to the video display layout information includes: Based on the video display layout information, the video information is synthesized according to spatial location to obtain a target video that simultaneously displays multiple video information; The target video is time-aligned by aligning multiple videos within it, and the time-aligned target video is then displayed.
[0047] It should be noted that the process of synthesizing the video information according to its spatial position based on the video display layout information can be achieved by displaying each video in a corresponding grid based on the video arrangement in the video display layout information. For example, if there are nine videos, they can be displayed in a 3x3 grid. (See reference...) Figure 2 , Figure 2 This is a schematic diagram illustrating the nine-grid video information display effect provided in Embodiment 1 of the video generation method of this application; Figure 2 Each small square in the diagram corresponds to video information under a specific shot and / or angle combination. If there are four videos, refer to... Figure 3 , Figure 3 This is a schematic diagram illustrating the display effect of a four-grid video information provided in Embodiment 1 of the video generation method of this application; Figure 3 Each small cell in the image corresponds to video information under a specific shot and / or angle combination. For example, the first cell in the first row and first column displays video information of a long shot and a frontal view; the first cell in the first row and second column displays video information of a close-up and a side view; the second cell in the second row and first column displays video information of a close-up and a frontal view; and the second cell in the second row and second column displays video information of a close-up and a back view. The step of time-aligning multiple videos in the target video and displaying the time-aligned target video can be achieved by mapping each frame of each video in the target video to a timeline, thus resolving the issues of asynchronous fast-forwarding, slow-motion, and start / end times.
[0048] In this embodiment, in response to the input scene description prompt, scene information corresponding to the scene description prompt is generated through a large model; based on the scene information, multiple video information corresponding to different shot sizes and / or angles are generated through the large model; the generated multiple video information are displayed. Compared with existing video content that only targets a fixed camera angle or shot size, this embodiment can generate multiple video information with different shot sizes and / or angles, allowing users to compare the visual effects of different shot sizes and / or angles, quickly select the best shot size and / or angle, and shorten the creative verification cycle.
[0049] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 , Figure 4 This is a flowchart illustrating a second embodiment of the video generation method of this application. Step S20 further includes the following steps: Step S201: Generate multiple target scene types and / or angle combinations based on the scene information; It should be noted that generating multiple target shot types and / or angle combinations based on the scene information can be determined by identifying suitable target shot types and / or angle combinations based on the semantic information and spatial layout of the scene information. For example, if the scene information is a city panorama and the protagonist is walking in the distance, the target shot types and / or angle combinations can be a combination of long shot and bird's-eye view, a combination of long shot and side view, a combination of close-up and front view, a combination of close-up and side view, etc. In addition to determining target shot types and / or angle combinations based on scene information, for the overall effect of the video, this embodiment can also determine target shot types and / or angle combinations based on the storyline of the scene. For example, at the beginning of the video (used to set the environment), a combination of long shot + bird's-eye view / overhead view target shot types and / or angle combinations can be used more often; in the middle of the video (emphasizing character interaction), a combination of medium shot + side / oblique side target shot types and / or angle combinations can be used more often; at the climax of the story (emotional outburst), a combination of close-up / close-up + varying angles can be used more often.
[0050] Furthermore, in order to increase the randomness of the target shot size and / or angle combination, to discover creative camera positions that are difficult to find under the mindset of setting up a fixed camera perspective, to generate video information from a better perspective than the user expects, to break the user's perspective layout limitations, and to inspire the user's perspective layout ideas, step S201 may include: generating multiple initial shot size and / or angle combinations based on the scene information. The initial shot size and / or angle combination is randomly perturbed to generate the target shot size and / or angle combination.
[0051] It should be noted that generating multiple initial shot sizes and / or angle combinations based on the scene information can be described as determining multiple shot sizes and / or angle combinations suitable for the scene information based on the scene information. Randomly perturbing the initial shot size and / or angle combination to generate a target shot size and / or angle combination can be achieved by randomly changing the shot size and / or angle in the initial shot size and / or angle combination to obtain creative camera positions. For example, if the initial shot size and / or angle combination is a long shot and a frontal eye-level angle, the target shot size and / or angle combination obtained after random perturbation can be a long shot and a side-view angle; if the initial shot size and / or angle combination is a close-up and a frontal eye-level angle, the target shot size and / or angle combination obtained after random perturbation can be a long shot and a rear-view angle; if the initial shot size and / or angle combination is a close-up and a frontal eye-level angle, the target shot size and / or angle combination obtained after random perturbation can be a close-up and a horizontal 30° angle with a vertical tilt angle of 45°; if the initial shot size and / or angle combination is a close-up and a frontal eye-level angle, the target shot size and / or angle combination obtained after random perturbation can be a medium shot and a horizontal 60° angle with a vertical tilt angle of -20°.
[0052] Furthermore, to ensure that the randomly perturbed viewpoint conforms to the viewpoint constraints of a specific scene, the initial shot size and / or angle combination includes a video shot size and a video angle; the step of randomly perturbing the initial shot size and / or angle combination to generate a target shot size and / or angle combination includes: The target scene is obtained by randomly perturbing the video scene based on a preset scene neighborhood range. And / or, The video angle is randomly perturbed to obtain the target angle, and the target shot and / or angle combination is determined based on the initial shot and / or angle combination, the target shot or the target angle.
[0053] It should be noted that the preset shot size neighborhood range limits the shot sizes that can be perturbed by a certain shot size. For example, for a grand environment display, even if shot size perturbation is required, it cannot be performed from a long shot to a close-up, but it can be performed from a long shot to a full shot or medium shot. For character emotion portrayal, even if shot size perturbation is required, it cannot be performed from a close-up to a long shot, but it can be performed from a close-up to a near shot. Randomly perturbing the video shot size based on the preset shot size neighborhood range can be done by randomly perturbing the shot sizes that can be perturbed by each shot size within the preset shot size neighborhood range to obtain the target shot size.
[0054] Randomly perturbing the video angle to obtain the target angle can be achieved by randomly changing the vertical elevation angle and horizontal azimuth angle of the video angle. For example, perturbing the front view to the side or back view, or perturbing the eye-level view to the upward view, downward view, or bird's-eye view. For video angles described in degrees, the degrees of the video angle can be directly perturbed randomly. For example, perturbing the horizontal angle 0° to 10° or 30°, or perturbing the vertical elevation angle 0° to 30°, -20°, or 60°.
[0055] Furthermore, before randomly perturbing the video scene based on a preset scene neighborhood range to obtain the target scene, the method further includes: determining candidate scene types in a preset shot candidate library; The perturbable shot type is determined based on the shot type characteristics of the candidate shot types; A preset scene neighborhood range is generated based on the perturbable scene size.
[0056] It should be noted that the preset shot candidate library includes multiple selectable candidate shot types, such as: long shot, full shot, medium shot, close-up, extreme close-up, etc. Determining the perturbable shot type based on the shot type characteristics of the candidate shot type can be done by determining the perturbable shot type based on the commonly used shooting scene types corresponding to the candidate shot type. For example, a long shot can be perturbed to a full shot and a medium shot, a close-up can be perturbed to an extreme close-up and a medium shot, a medium shot can be perturbed to a full shot and a close-up, and in some suspenseful scenes, an extreme close-up can also be perturbed to a long shot and a medium shot. Generating a preset shot type neighborhood range based on the perturbable shot type can be done by determining the perturbable shot type corresponding to each candidate shot type based on the perturbable shot type and the candidate shot types, thus obtaining the preset shot type neighborhood range.
[0057] In practice, the user-input scene description prompts could be: "A detective gazes into the distance at a rainy night dock, searching for clues." The scene information after the large model is parsed could be: Subject: Detective; Environment: Rainy night dock; Atmosphere: Suspenseful, cold; Plot point: Gazing into the distance (emotion + environment); The generated multiple different shot sizes and / or angles could be: Long shot + bird's-eye view (showing the dock's night view and rain); Full shot + back view (the detective stands on the dock, his back blending into the environment); Medium shot + side view (walking in the rain, holding an umbrella); Close-up + upward view (the detective looks up at the dark sea); Extreme close-up + eye-level view (sharp eyes, raindrops on glasses); Full shot + downward view (viewing the detective and the entire dock layout from a high place); Medium shot + oblique side view (talking to other characters); Long shot + upward view (lighthouse beam piercing through the rain); Extreme close-up + downward view (wet hands holding an old photograph).
[0058] Furthermore, the step of randomly perturbing the video angle to obtain the target angle includes: Based on the scene information, scene features are determined, and the scene features include at least one of the following: scene environment type, entity action information, spatial scale, and scene emotion information; The target angle is obtained by randomly perturbing the video angle based on the scene characteristics.
[0059] It should be noted that scene environment types can include macroscopic natural landscapes, enclosed indoor spaces, etc. Entity action information can include looking up, large-scale fighting actions, and subtle movements. Spatial scale can include small and large spatial scales. Scene emotional information can include suspense, sadness, joy, etc. Randomly perturbing the video angle based on the scene characteristics to obtain the target angle can be done by determining the video angle preference range corresponding to the scene characteristics, and randomly perturbing the video angle under the constraints of the video angle preference range. For example: in a magnificent natural landscape environment, the sense of scale can be emphasized, and the angle can be a tilt angle, prioritizing bird's-eye view or overhead view, while the azimuth angle can be frontal or surrounding. In an enclosed indoor space, it needs to conform to the physical space, and the tilt angle can be limited to eye level or a slight overhead view. When the main action is looking up, the angle needs to be consistent with the action, leaning towards looking up or eye level, avoiding overhead views. In a suspenseful mood, the azimuth angle often uses the back, side, or oblique side, and the tilt angle often uses an overhead view or a low angle.
[0060] Step S202: Generate video information using the large model based on the scene information and the target shot size and / or angle combination.
[0061] It should be noted that the step of generating video information through the large model based on the scene information and the target shot size and / or angle combination can be achieved by converting each target shot size and / or angle combination into a corresponding virtual camera pose through the large model, and then rendering and generating the video information corresponding to different target shot sizes and / or angle combinations based on the scene information.
[0062] This embodiment generates multiple target scene types and / or angle combinations based on the scene information; Based on the scene information and the target shot size and / or angle combination, video information is generated through the large model. This embodiment automatically generates video information under different target shot sizes and / or angle combinations based on scene information. Compared with the existing method where users have to repeatedly try different prompts to obtain multi-view videos, this reduces the trial and error cost. Moreover, all videos share the same set of scene information, which avoids the randomness of the scene when drawing cards, and realizes the integrated, high-quality and controllable generation of multi-view videos.
[0063] Based on the above embodiments of this application, in the third embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 5 , Figure 5 This is a flowchart illustrating the video generation method of this application in Embodiment 3. Step S30 further includes the following steps: Step S40: Upon receiving a video regeneration instruction, obtain the target video information and lens adjustment parameters corresponding to the video regeneration instruction; It should be noted that the video regeneration command can be a user-triggered command to regenerate a video based on multiple displayed video information. The video regeneration command can include the user-selected lens adjustment parameters and video information, i.e., the target video information. Specifically, the user can further optimize the perspective in the displayed video information based on the multiple video information. The perspective can include the video framing and video angle. For example, for a video with a long-range framing, the user can select the video and input lens adjustment parameters such as: zoom in a little more, rotate 5 degrees clockwise horizontally, lower the lens height, etc., to adjust the perspective.
[0064] Step S50: Determine the video shot and video angle corresponding to the target video information; It should be noted that the video shot and angle corresponding to the target video information can be the video shot and angle corresponding to the video for which the user selects the video whose perspective needs to be adjusted. Users can select videos by clicking on the displayed video information, and the system will automatically obtain the video shot and angle of the selected video.
[0065] Step S60: Determine the target lens parameters based on the video shot size, the video angle, and the lens adjustment parameters; It should be noted that determining the target lens parameters based on the video shot size, the video angle, and the lens adjustment parameters can be achieved by adjusting the video shot size and the video angle according to the user's lens adjustment parameters, thereby obtaining the adjusted video shot size and video angle, i.e., the target lens parameters.
[0066] Step S70: Generate video information corresponding to the target lens parameters using the large model based on the scene information.
[0067] It should be noted that the process of generating video information corresponding to the target lens parameters based on the scene information using the large model can be the generation of video information under the target lens parameters based on the scene information using the large model.
[0068] Furthermore, in order to efficiently generate video information corresponding to different shot sizes and / or angles based on scene information, step S20 includes: responding to the input scene description prompts, determining video generation elements based on the scene description prompts using a large model; Based on the video generation elements, the three-dimensional spatial structure of the video scene is laid out using the large model to obtain scene information corresponding to the scene description prompts.
[0069] It should be noted that determining video generation elements based on scene description prompts using a large model can be achieved by parsing the semantics of the scene description prompts using the large model's natural language understanding capabilities. These video generation elements can be key information units used to guide subsequent video generation, including subject category, character, action, environment type, time, weather, emotional atmosphere, light source, entity relationships, and environmental object relationships. The 3D spatial structure layout of the video scene based on these video generation elements using the large model can be based on these elements, where the large model determines the position, size, orientation, and spatial relationships of the main objects in the scene within a 3D coordinate system, resulting in a 3D scene model. The scene information can be a structured data set output after the 3D spatial structure layout.
[0070] Furthermore, in order to ensure cross-view consistency and perspective rationality, the scene information is three-dimensional spatial scene information, which includes at least one of the following: the spatial position of the main body in the scene, the geometric layout data of the scene, the spatial relationship constraints between scene entities, and the lighting direction of the scene.
[0071] It should be noted that the spatial position of the main body in the scene can be the positional relationship between various main bodies in the scene, for example, chairs around a table; the geometric layout data of the scene can be the three-dimensional shape and distribution data of various entities (main bodies, environmental objects, terrain, buildings, etc.) in the scene, for example: using a cuboid to represent the outer boundary of an object, and using a large number of discrete points to represent the surface shape; the spatial relationship constraints between scene entities can be rules or restrictions describing the relative position, distance, orientation, etc. of different entities in space, and can be the minimum or maximum distance allowed between entities, for example: for a scene where "a person stands on the edge of a dock, facing the sea", the spatial relationship constraints between scene entities can be: the positional relationship between the person and the edge of the dock is 1 meter apart, and the person's orientation is facing the sea, as well as positional relationships and spacing constraints such as "streetlights are located on both sides of the road, 10 meters apart"; the lighting direction of the scene can be information such as the divergence direction of important light sources (sunlight, lamplight) in the scene.
[0072] In this embodiment, upon receiving a video regeneration command, the system acquires the target video information and lens adjustment parameters corresponding to the command; determines the video framing and angle corresponding to the target video information; determines the target lens parameters based on the framing, angle, and adjustment parameters; and generates the video information corresponding to the target lens parameters using the large model based on the scene information. In this embodiment, the user can further optimize the perspective based on multiple displayed video framing and / or angles, and generate the video based on pre-determined scene information. This allows for further optimization of the perspective while ensuring the integration of multi-view videos, improving user experience and reducing trial-and-error costs.
[0073] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the video generation method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0074] This application also provides a video generation apparatus, please refer to... Figure 6 The video generation device includes: The response module 10 is used to respond to the input scene description prompt words and generate scene information corresponding to the scene description prompt words through the large model; The video generation module 20 is used to generate multiple video information corresponding to different scene types and / or angles based on the scene information and the large model; The display module 30 is used to display multiple generated video information.
[0075] In this embodiment, in response to the input scene description prompt, scene information corresponding to the scene description prompt is generated through a large model; based on the scene information, multiple video information corresponding to different shot sizes and / or angles are generated through the large model; the generated multiple video information are displayed. Compared with existing video content that only targets a fixed camera angle or shot size, this embodiment can generate multiple video information with different shot sizes and / or angles, allowing users to compare the visual effects of different shot sizes and / or angles, quickly select the best shot size and / or angle, and shorten the creative verification cycle.
[0076] The video generation apparatus provided in this application, employing the video generation method described in the above embodiments, can solve the technical problem that existing video generation methods can only generate single-viewpoint or single-scene formats. Compared with the prior art, the beneficial effects of the video generation apparatus provided in this application are the same as those of the video generation method described in the above embodiments, and other technical features in the video generation apparatus are the same as those disclosed in the methods described in the above embodiments, and will not be repeated here.
[0077] This application provides a video generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the video generation method in Embodiment 1 above.
[0078] The following is for reference. Figure 7 The diagram illustrates a structural schematic of a video generation device suitable for implementing embodiments of this application. The video generation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The video generation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0079] like Figure 7As shown, the video generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the video generation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the video generating device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows video generating devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.
[0080] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0081] The video generation device provided in this application, employing the video generation method described in the above embodiments, can solve the technical problem that existing video generation methods can only generate single-viewpoint or single-scene formats. Compared with the prior art, the beneficial effects of the video generation device provided in this application are the same as those of the video generation method described in the above embodiments, and other technical features of this video generation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0082] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0083] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0084] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the video generation method described in the above embodiments.
[0085] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0086] The aforementioned computer-readable storage medium may be included in the video generating device; or it may exist independently and not assembled into the video generating device.
[0087] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Python, Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0089] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0090] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described video generation method, thereby solving the technical problem that existing video generation methods can only generate a single viewpoint or scene. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the video generation method provided in the above embodiments, and will not be repeated here.
[0091] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the video generation method described above.
[0092] The computer program product provided in this application can solve the technical problem that existing video generation methods can only generate a single viewpoint or scene. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the video generation method provided in the above embodiments, and will not be repeated here.
[0093] All user-related data (e.g., user input data) involved in this application were obtained with the user's permission or consent. In other words, when this application is applied to specific products or technologies, user permission is required to acquire and process the relevant data, and the processing of the data must comply with the relevant laws, regulations, and regulatory standards of the relevant countries and regions. For example, when it is necessary to obtain a user's current geographical location, a location acquisition prompt can be displayed on the user's terminal. After receiving confirmation from the user regarding the location acquisition prompt, the terminal can obtain the user's current geographical location.
[0094] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.
[0095] This invention discloses A1. A video generation method, the video generation method comprising the following steps: In response to the input scene description prompts, scene information corresponding to the scene description prompts is generated through a large model; Based on the scene information, the large model generates video information corresponding to multiple different scene types and / or angles; The generated video information will be displayed.
[0096] A2. The video generation method as described in A1, wherein generating multiple video information corresponding to different scene sizes and / or angles based on the scene information using the large model includes: Multiple target scene types and / or angle combinations are generated based on the scene information; Video information is generated using the large model based on the scene information and the target shot size and / or angle combination.
[0097] A3. The video generation method as described in A2, wherein generating multiple target scene types and / or angle combinations based on the scene information includes: Multiple initial shot types and / or angle combinations are generated based on the scene information; The initial shot size and / or angle combination is randomly perturbed to generate the target shot size and / or angle combination.
[0098] A4. The video generation method as described in A3, wherein the initial shot and / or angle combination includes video shot and video angle; The step of randomly perturbing the initial shot size and / or angle combination to generate the target shot size and / or angle combination includes: The target scene is obtained by randomly perturbing the video scene based on a preset scene neighborhood range. And / or, The video angle is randomly perturbed to obtain the target angle, and the target shot and / or angle combination is determined based on the initial shot and / or angle combination, the target shot or the target angle.
[0099] A5. The video generation method as described in A4, before randomly perturbing the video scene based on a preset scene neighborhood range to obtain the target scene, further includes: Determine the candidate shot sizes in the preset shot candidate library; The perturbable shot type is determined based on the shot type characteristics of the candidate shot types; A preset scene neighborhood range is generated based on the perturbable scene size.
[0100] A6. The video generation method as described in A4, wherein randomly perturbing the video angle to obtain the target angle includes: Based on the scene information, scene features are determined, and the scene features include at least one of the following: scene environment type, entity action information, spatial scale, and scene emotion information; The target angle is obtained by randomly perturbing the video angle based on the scene characteristics.
[0101] A7. The video generation method as described in A6, wherein the step of randomly perturbing the video angle according to the scene features to obtain the target angle includes: Determine the video angle preference range corresponding to the scene information based on the scene features; The target angle is obtained by randomly perturbing the video angle based on the video angle preference range.
[0102] A8. The video generation method as described in any one of A1-A7, further comprising, after displaying the generated multiple video information: Upon receiving a video regeneration instruction, the target video information and lens adjustment parameters corresponding to the video regeneration instruction are obtained; Determine the video shot and video angle corresponding to the target video information; The target lens parameters are determined based on the video shot type, the video angle, and the lens adjustment parameters. Based on the scene information, the large model generates video information corresponding to the target lens parameters.
[0103] A9. The video generation method as described in any one of A1-A7, wherein the step of generating scene information corresponding to the scene description prompt word through a large model in response to the input scene description prompt word includes: In response to the input scene description prompts, the large model determines the video generation elements based on the scene description prompts; Based on the video generation elements, the three-dimensional spatial structure of the video scene is laid out using the large model to obtain scene information corresponding to the scene description prompts.
[0104] A10. The video generation method as described in A9, wherein the scene information is three-dimensional spatial scene information, and the three-dimensional spatial scene information includes at least one of the following: the spatial position of the main body in the scene, the geometric layout data of the scene, the spatial relationship constraints between scene entities, and the lighting direction of the scene.
[0105] A11. The video generation method as described in any one of A1-A7, wherein displaying the generated multiple video information includes: The number of videos is determined based on the video information; The video display layout information is determined based on the number of videos; The video information is displayed according to the video display layout information.
[0106] A12. The video generation method as described in A11, wherein displaying the video information according to the video display layout information includes: Based on the video display layout information, the video information is synthesized according to spatial location to obtain a target video that simultaneously displays multiple video information; The target video is time-aligned by aligning multiple videos within it, and the time-aligned target video is then displayed.
[0107] This invention discloses B13. A video generation apparatus, the video generation apparatus comprising: The response module is used to respond to the input scene description prompts and generate scene information corresponding to the scene description prompts through the large model; The video generation module is used to generate multiple video information corresponding to different scene types and / or angles based on the scene information and the large model; The display module is used to showcase multiple generated video information items.
[0108] B14. The video generation apparatus as described in B13, wherein the video generation module is further configured to generate multiple target scene types and / or angle combinations based on the scene information; Video information is generated using the large model based on the scene information and the target shot size and / or angle combination.
[0109] B15. The video generation apparatus as described in B14, wherein the video generation module is further configured to generate multiple initial shot sizes and / or angle combinations based on the scene information; The initial shot size and / or angle combination is randomly perturbed to generate the target shot size and / or angle combination.
[0110] B16. The video generation apparatus as described in B15, wherein the initial shot and / or angle combination includes a video shot and a video angle; The video generation module is also used to randomly perturb the video scene based on a preset scene neighborhood range to obtain the target scene. And / or, The video angle is randomly perturbed to obtain the target angle, and the target shot and / or angle combination is determined based on the initial shot and / or angle combination, the target shot or the target angle.
[0111] B17. The video generation apparatus as described in B16, wherein the video generation module is further configured to determine candidate shot types in a preset shot candidate library; The perturbable shot type is determined based on the shot type characteristics of the candidate shot types; A preset scene neighborhood range is generated based on the perturbable scene size.
[0112] This invention discloses C18. A video generation apparatus, the apparatus comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the video generation method as described in any one of A1 to A12.
[0113] The present invention discloses D19. A storage medium, which is a computer-readable storage medium, wherein a computer program is stored on the storage medium, and the computer program, when executed by a processor, implements the steps of the video generation method as described in any one of A1 to A12.
[0114] This invention discloses E20. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the video generation method as described in any one of A1 to A12.
Claims
1. A video generation method, characterized in that, The video generation method includes the following steps: In response to the input scene description prompts, scene information corresponding to the scene description prompts is generated through a large model; Based on the scene information, the large model generates video information corresponding to multiple different scene types and / or angles; The generated video information will be displayed.
2. The video generation method as described in claim 1, characterized in that, The generation of multiple video information corresponding to different scene sizes and / or angles based on the scene information using the large model includes: Multiple target scene types and / or angle combinations are generated based on the scene information; Video information is generated using the large model based on the scene information and the target shot size and / or angle combination.
3. The video generation method as described in claim 2, characterized in that, The generation of multiple target scene types and / or angle combinations based on the scene information includes: Multiple initial shot types and / or angle combinations are generated based on the scene information; The initial shot size and / or angle combination is randomly perturbed to generate the target shot size and / or angle combination.
4. The video generation method as described in claim 3, characterized in that, The initial shot size and / or angle combination includes video shot size and video angle; The step of randomly perturbing the initial shot size and / or angle combination to generate the target shot size and / or angle combination includes: The target scene is obtained by randomly perturbing the video scene based on a preset scene neighborhood range. And / or, The video angle is randomly perturbed to obtain the target angle, and the target shot and / or angle combination is determined based on the initial shot and / or angle combination, the target shot or the target angle.
5. The video generation method as described in claim 4, characterized in that, Before obtaining the target scene by randomly perturbing the video scene based on a preset scene neighborhood range, the method further includes: Determine the candidate shot sizes in the preset shot candidate library; The perturbable shot type is determined based on the shot type characteristics of the candidate shot types; A preset scene neighborhood range is generated based on the perturbable scene size.
6. The video generation method as described in claim 4, characterized in that, The step of randomly perturbing the video angle to obtain the target angle includes: Based on the scene information, scene features are determined, and the scene features include at least one of the following: scene environment type, entity action information, spatial scale, and scene emotion information; The target angle is obtained by randomly perturbing the video angle based on the scene characteristics.
7. A video generation apparatus, characterized in that, The video generation device includes: The response module is used to respond to the input scene description prompts and generate scene information corresponding to the scene description prompts through the large model; The video generation module is used to generate multiple video information corresponding to different scene types and / or angles based on the scene information and the large model; The display module is used to showcase multiple generated video information items.
8. A video generation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the video generation method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the video generation method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the video generation method as described in any one of claims 1 to 6.