Video generation method, video generation model training method, device and equipment
By using voxel features in three-dimensional space during the video generation process, the target video at a set viewing angle is generated, and the problem of poor quality in multi-view videos is solved, and higher quality multi-view video generation is achieved.
Patent Information
- Application Number
- CN202510346380.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
AI Technical Summary
The video quality generated in the prior art is poor, especially in the multi-view video generation, there are problems of object deformation, depth distortion and edge blur.
By obtaining the original video frame from the observation perspective, using the voxel characteristics in the three-dimensional space to generate the target video from the set perspective, retaining the object's three-dimensional information to improve the consistency and authenticity of multi-view videos.
It improves the three-dimensional consistency between multi-view videos and the quality of generated videos, and is suitable for multi-view autonomous driving video generation scenarios.
Smart Images

Figure CN120302022A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of autonomous driving and video generation, and particularly to a video generation method, a training method of a video generation model, an apparatus, a device, a vehicle, a chip, and a storage medium. Background Art
[0002] In the related art, most rely on two-dimensional information (such as two-dimensional detection boxes) to generate single-view or multi-view videos. For example, in a driving scenario, a diffusion model is used to generate a multi-view driving video using the two-dimensional detection boxes of obstacles and road images. The generated video has problems such as object deformation, depth distortion, and edge blurring, that is, the quality of the generated video is poor. Summary of the Invention
[0003] The present disclosure provides a video generation method, a training method of a video generation model, an apparatus, a device, a vehicle, a chip, and a computer-readable storage medium to at least solve the problem of poor quality of the generated video in the related art. The technical solution of the present disclosure is as follows:
[0004] According to the first aspect of the embodiments of the present disclosure, a video generation method is provided, including: obtaining at least one frame in an original video from an observation perspective; determining the feature of the voxel of the corresponding frame in a three-dimensional space based on the feature of the image of the at least one frame in the original video in a two-dimensional space; and generating a target video from a set perspective based on the features of the voxels of each frame in the three-dimensional space.
[0005] According to the second aspect of the embodiments of the present disclosure, a training method of a video generation model is provided. The video generation model includes a first generation model and a second generation model. The video generation model is used to generate a target video from a set perspective. The method includes: generating a second sample video from an observation perspective based on a plurality of first sample videos from a set perspective, and obtaining at least one frame in the second sample video; determining the sample feature of the voxel of the corresponding frame in a three-dimensional space based on the sample feature of the image of the at least one frame in the second sample video in a two-dimensional space; training the first generation model based on the second sample video and the sample features of the voxels of each frame in the three-dimensional space; and / or training the second generation model based on at least one of the first sample videos from the set perspective and the sample features of the voxels of each frame in the three-dimensional space.
[0006] According to a third aspect of the embodiments of the present disclosure, there is provided a video generation device, including: an acquisition module configured to acquire at least one frame in an original video from an observation perspective; a determination module configured to determine the characteristics of voxels of a corresponding frame in a three-dimensional space based on the characteristics of an image of at least one frame in the original video in a two-dimensional space; and a generation module configured to generate a target video from a set perspective based on the characteristics of voxels of each frame in the three-dimensional space.
[0007] According to a fourth aspect of the embodiments of the present disclosure, there is provided a training device for a video generation model. The video generation model includes a first generation model and a second generation model, and is used to generate a target video from a set perspective. The device includes: an acquisition module configured to generate a second sample video from an observation perspective based on a first sample video from a plurality of set perspectives, and acquire at least one frame in the second sample video; a determination module configured to determine the sample characteristics of voxels of a corresponding frame in a three-dimensional space based on the sample characteristics of an image of at least one frame in the second sample video in a two-dimensional space; a first training module configured to train the first generation model based on the second sample video and the sample characteristics of voxels of each frame in the three-dimensional space; and / or, a second training module configured to train the second generation model based on the first sample video from at least one of the set perspectives and the sample characteristics of voxels of each frame in the three-dimensional space.
[0008] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the video generation method according to the first aspect of the embodiments of the present disclosure are implemented, and / or, the steps of the training method of the video generation model according to the second aspect of the embodiments of the present disclosure are implemented.
[0009] According to a sixth aspect of the embodiments of the present disclosure, there is provided a vehicle, including a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of the video generation method according to the first aspect of the embodiments of the present disclosure, and / or, the steps of the training method of the video generation model according to the second aspect of the embodiments of the present disclosure.
[0010] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the video generation method according to the first aspect of the embodiments of the present disclosure are implemented, and / or, the steps of the training method of the video generation model according to the second aspect of the embodiments of the present disclosure are implemented.
[0011] According to an eighth aspect of the embodiments of the present disclosure, a chip is provided. The chip includes an interface circuit and a processing circuit that are coupled to each other. The interface circuit is configured to input or output signals, and the processing circuit is configured to implement the steps of the video generation method described in the first aspect of the embodiments of the present disclosure, and / or implement the steps of the training method of the video generation model described in the second aspect of the embodiments of the present disclosure.
[0012] According to a ninth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements the steps of the video generation method described in the first aspect of the embodiments of the present disclosure, and / or implements the steps of the training method of the video generation model described in the second aspect of the embodiments of the present disclosure.
[0013] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: obtaining at least one frame in the original video from the observation perspective, determining the voxel features of the corresponding frame in the three-dimensional space based on the features of the image of at least one frame in the original video in the two-dimensional space, and generating a target video from the set perspective based on the voxel features of each frame in the three-dimensional space. Thus, the voxel features retain the three-dimensional information of the object, enabling the generated target video from the set perspective to maintain the original three-dimensional features of the object, that is, maintaining visual consistency and authenticity. In the multi-view video generation scenario, it helps to improve the three-dimensional consistency between multi-view videos, improves the quality of the generated videos, and is applicable to the generation scenario of multi-view autonomous driving videos.
[0014] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0016] Figure 1 is a flowchart showing a video generation method according to an exemplary embodiment.
[0017] Figure 2 is a flowchart showing a video generation method according to another exemplary embodiment.
[0018] Figure 3 is a flowchart showing a video generation method according to another exemplary embodiment.
[0019] Figure 4 is a flowchart showing a video generation method according to another exemplary embodiment.
[0020] Figure 5 It is a schematic flowchart of a method for training a video generation model shown according to an exemplary embodiment.
[0021] Figure 6 It is a schematic diagram of a video generation model shown according to an exemplary embodiment.
[0022] Figure 7 It is a schematic structural diagram of a video generation device shown according to an exemplary embodiment.
[0023] Figure 8 It is a schematic structural diagram of a training device for a video generation model shown according to an exemplary embodiment.
[0024] Figure 9 It is a schematic structural diagram of a vehicle shown according to an exemplary embodiment.
[0025] Figure 10 It is a schematic structural diagram of a chip shown according to an exemplary embodiment. Detailed implementation manners
[0026] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0028] Next, the video generation method, the method for training a video generation model, the device, the electronic device, the vehicle, the chip, and the storage medium in the embodiments of the present disclosure will be described with reference to the drawings.
[0029] Figure 1 It is a schematic flowchart of a video generation method shown according to an exemplary embodiment. As Figure 1 shown, the video generation method in the embodiments of the present disclosure includes the following steps.
[0030] S101, obtain at least one frame in the original video from the observation perspective.
[0031] It should be noted that the execution subject of the video generation method in the embodiments of the present disclosure is an electronic device, such as an in-vehicle terminal, an in-vehicle controller, a server, an ISP (Image Signal Processor, image processing chip), etc. The video generation method in the embodiments of the present disclosure can be executed by the video generation device in the embodiments of the present disclosure. The video generation device in the embodiments of the present disclosure can be configured in any electronic device to execute the video generation method in the embodiments of the present disclosure.
[0032] There are no excessive limitations on the observation perspective. For example, it may include the perspective of looking down at the ground from a high altitude, that is, the "bird's-eye" perspective. When the observation perspective is the bird's-eye perspective, the original video includes multiple frames of bird's-eye images. There are no excessive limitations on the frame images in any video in the present disclosure. For example, it may include RGB, HSV, HSL, YCbCr, Lab, YUV, etc. There are no excessive limitations on any video in the present disclosure. For example, it may include a driving video (such as an autonomous driving video), an XR (Extended Reality, extended reality) video, a navigation video, a film and television video, a game video, a surveillance video, etc.
[0033] Optionally, obtaining the original video from the observation perspective includes capturing the original video from the observation perspective through the camera corresponding to the observation perspective. For example, taking a driving scenario (such as an autonomous driving scenario) as an example, the drone can be controlled to move to the upper area of the vehicle, and the original video from the bird's-eye perspective can be captured through the camera set on the drone.
[0034] Optionally, obtaining the original video from the observation perspective includes obtaining candidate videos from multiple set perspectives and generating the original video from the observation perspective based on the candidate videos from multiple set perspectives.
[0035] There are no excessive limitations on the set perspectives. For example, it may include front view, left front, left rear, right front, right rear, rear view, etc.
[0036] For example, taking a driving scenario as an example, multiple cameras can be set on the vehicle. The cameras correspond to the set perspectives one by one, and the candidate videos from their respective corresponding set perspectives can be captured through the multiple cameras.
[0037] Among them, generating the original video from the observation perspective based on the candidate videos from multiple set perspectives can be implemented by any video generation method in related technologies, and there are no excessive limitations here.
[0038] Optionally, obtaining the original video from the observation perspective includes generating at least one frame of image from the observation perspective through a mapping tool and generating the original video from the observation perspective based on the images from each frame in the observation perspective.
[0039] S102. Determine the voxel features of the corresponding frame in three-dimensional space based on the features of at least one frame of the image in two-dimensional space in the original video.
[0040] It should be noted that the features of the image of any frame in two-dimensional space refer to the features of any frame image in the original video, that is, the features of the image of any frame in the two-dimensional space corresponding to the observation perspective. A voxel (Volume Pixel) is short for volume element, which is a basic unit for discretely representing information on a three-dimensional scale and represents the information of a volume unit in three-dimensional space.
[0041] The features of the image are the features of the object shown in the original video in two-dimensional space, and the features of the voxel are the features of the object shown in the original video in three-dimensional space. Compared with the features of the image, the features of the voxel retain the three-dimensional information of the object, such as the three-dimensional geometric details of the object, so that the generated target video under the set perspective can maintain the original three-dimensional features of the object, that is, maintain visual consistency and authenticity.
[0042] In the scenario of multi-view video generation, it can make the generated multi-view video maintain the original three-dimensional features of the object, that is, maintain visual consistency and authenticity, improve the three-dimensional consistency between multi-view videos, and provide a consistent and realistic visual experience whether the multi-view video is a single view or across multiple views.
[0043] Avoid problems such as object deformation, depth distortion, and edge blurring in the generated single-view and multi-view videos, and improve the quality of the generated single-view and multi-view videos.
[0044] For example, taking the driving scenario as an example, the features of the voxel can include three-dimensional information such as the trajectory of dynamic obstacles, road topology, the inclination of the facade of buildings, and road curvature. For example, based on the trajectory of dynamic obstacles and road topology in a traffic accident scenario, a multi-view driving video in the traffic accident scenario can be generated as a simulated driving video to replace high-risk real vehicle tests.
[0045] For example, the features of the image and the voxel can both include color features, texture features, shape features, spatial relationship features, object category, object action, object pose, object emotion, etc.
[0046] For example, based on the features of the image of the first frame in the original video in two-dimensional space, that is, the features of the first frame image of the original video, determine the voxel features of the first frame in three-dimensional space.
[0047] Based on the features of the image of the second frame in the original video in two-dimensional space, that is, the features of the second frame image of the original video, determine the voxel features of the second frame in three-dimensional space.
[0048] Determine the voxel features of the third frame in three-dimensional space based on the features of the image of the third frame in the original video in two-dimensional space, i.e., the features of the image of the third frame of the original video.
[0049] It should be noted that to determine the voxel features of the corresponding frame in three-dimensional space based on the features of the image of at least one frame in two-dimensional space, any three-dimensional reconstruction method in related technologies can be used to achieve this, and no excessive limitations are imposed here.
[0050] For example, to determine the voxel features of the corresponding frame in three-dimensional space based on the features of the image of at least one frame in the original video in two-dimensional space, it includes constructing a three-dimensional model of the corresponding frame based on the features of the image of at least one frame in the original video in two-dimensional space, and performing voxelization processing on the three-dimensional model of at least one frame to obtain the voxel features of the corresponding frame in three-dimensional space.
[0051] S103. Generate a target video from the voxel features of each frame in three-dimensional space at a set viewing angle.
[0052] For example, based on the voxel features of the 1st to 20th frames in the original video from a bird's-eye view in three-dimensional space, generate target video 1 from a front view, target video 2 from a left-front view, target video 3 from a left-rear view, target video 4 from a right-front view, target video 5 from a right-rear view, and target video 6 from a rear view. Among them, target videos 1 to 6 each include 20 frames of images.
[0053] Optionally, to generate a target video from the voxel features of each frame in three-dimensional space at a set viewing angle, it includes generating an image of the corresponding frame at the set viewing angle based on the voxel features of at least one frame, and generating a target video at the set viewing angle based on the images of each frame at the set viewing angle.
[0054] For example, generate an image of the first frame from a front view based on the voxel features of the first frame in the original video from a bird's-eye view in three-dimensional space.
[0055] Generate an image of the second frame from a front view based on the voxel features of the second frame in the original video from a bird's-eye view in three-dimensional space.
[0056] Generate an image of the third frame from a front view based on the voxel features of the third frame in the original video from a bird's-eye view in three-dimensional space.
[0057] Optionally, based on the features of voxels of each frame in three-dimensional space, a target video from a set perspective is generated, including obtaining a description text of the generation requirements for the target video from the set perspective, and generating the target video from the set perspective based on the features of voxels of each frame in three-dimensional space and the description text. Thus, the generation requirements for a video from a certain perspective and the features of voxels of each frame in three-dimensional space can be comprehensively considered to generate the target video from this perspective, making the generated target video from this perspective meet the corresponding video generation requirements, with strong controllability for multi-perspective video generation and enabling personalization of multi-perspective video generation. In addition, the features of voxels of each frame in three-dimensional space can be kept unchanged, and the description text can be adjusted multiple times to batch-generate diverse multi-perspective videos.
[0058] It should be noted that the generation requirements for target videos from different set perspectives may be different or the same, and there are no excessive restrictions on the generation requirements for target videos from the set perspective. For example, they may include weather, time, etc. For example, the description text of the generation requirements for the target video from the set perspective may include "Please generate a video segment in heavy snow weather at night".
[0059] For example, taking the driving scenario as an example, video generation requirements such as weather and time can be set. For example, the weather can be set to inclement weather such as rain, snow, and fog, and the time can be set to night, and multi-perspective driving videos in low visibility environments such as rain, snow, fog, and night can be generated as simulated driving videos to simulate the interaction between vehicles and pedestrians in low visibility environments and replace high-risk real vehicle tests.
[0060] For example, the method further includes adding the simulated driving video to the driving video library to expand the driving video library, so that the driving video library can cover a variety of driving scenarios, such as driving scenarios under various weather conditions, road conditions, different perspectives, and different time periods, improving the richness and diversity of the driving video library.
[0061] In addition, when the driving video library is used for training and / or testing an autonomous driving model, it helps to improve the adaptability of the autonomous driving model to a variety of driving scenarios and enhances the robustness and reliability of the autonomous driving model.
[0062] Optionally, based on the features of voxels of each frame in three-dimensional space, a target video from a set perspective is generated, including adjusting the features of voxels of at least one frame in three-dimensional space according to the generation requirements of the target video from the set perspective, and generating the target video from the set perspective based on the final features of voxels of each frame in three-dimensional space. Thus, considering the video generation requirements from a certain perspective, the features of voxels of at least one frame in three-dimensional space can be adjusted, and the target video from this perspective can be generated based on the final features of voxels of each frame in three-dimensional space. The controllability of multi-perspective video generation is relatively strong, and the personalization of multi-perspective video generation can be realized. In addition, the generation requirements can be adjusted multiple times, and then the features of voxels of at least one frame in three-dimensional space can be adjusted multiple times to batch generate diverse multi-perspective videos.
[0063] For example, taking a driving scenario as an example, the object position, road curvature, etc. in the features of voxels can be adjusted to batch generate diverse multi-perspective driving videos as simulated driving videos to simulate driving scenarios such as occlusion and small targets at a distance, and the autonomous driving model can be trained and / or tested based on the above simulated driving videos, which helps to improve the generalization ability of the autonomous driving model in driving scenarios such as occlusion and small targets at a distance.
[0064] Optionally, as Figure 6 shown, the video generation model includes a first generation model and a second generation model. It should be noted that there are no excessive limitations on the video generation model. For example, it can include a diffusion model, GAN (Generative Adversarial Networks), VAEs (Variational Autoencoders), etc. It should be noted that VAEs include an encoder and a decoder.
[0065] The original video from the observed perspective is input into the first generation model, and the first generation model outputs the features of voxels of each frame in three-dimensional space. The features of voxels of each frame in three-dimensional space are input into the second generation model, and the second generation model outputs the target video from the set perspective.
[0066] The video generation method provided by the embodiments of the present disclosure obtains at least one frame in the original video from the observed perspective, determines the features of voxels of the corresponding frame based on the features of the image of at least one frame in the original video in two-dimensional space, and generates the target video from the set perspective based on the features of voxels of each frame in three-dimensional space. Thus, the features of voxels retain the three-dimensional information of the object, enabling the generated target video from the set perspective to maintain the original three-dimensional features of the object, that is, maintaining visual consistency and authenticity. In the multi-perspective video generation scenario, it helps to improve the three-dimensional consistency between multi-perspective videos, improves the quality of the generated videos, and is applicable to the generation scenario of multi-perspective autonomous driving videos.
[0067] The video generation method provided by the embodiments of the present disclosure is applicable to video generation scenarios such as driving scenarios, XR scenarios, navigation scenarios, film and television scenarios, game scenarios, monitoring scenarios, traffic planning, etc.
[0068] In the first case, taking the driving scenario as an example, before obtaining at least one frame of the original video from the observation perspective, it further includes obtaining a first instruction, where the first instruction is triggered based on an operation on the target vehicle.
[0069] Obtaining the original video includes obtaining the real driving video of the target vehicle from the observation perspective as the original video.
[0070] Optionally, the first instruction is triggered based on a touch operation and / or a voice input operation on the target vehicle. For example, the first instruction is triggered based on a touch operation on the interaction interface of the target vehicle, and / or the first instruction is triggered based on a voice input operation on the voice collection device of the target vehicle. For example, the user can say "I want to watch a multi-perspective driving video" or "I want to watch a driving video from a set perspective".
[0071] Optionally, the target video from the set perspective is used for visual display of the target vehicle.
[0072] Thus, after obtaining the first instruction, the real driving video of the target vehicle from the observation perspective can be obtained to generate the target video of the target vehicle from the set perspective, and the visual display of the multi-perspective driving video can be realized.
[0073] In addition, the generated target video from the set perspective can maintain the original three-dimensional features of the object. In the multi-perspective driving video generation scenario, it helps to improve the three-dimensional consistency between multi-perspective driving videos and avoid object deformation and depth distortion in multi-perspective driving videos.
[0074] In the second case, taking the driving scenario as an example, obtaining the original video from the observation perspective includes obtaining real driving videos from multiple set perspectives and generating the original video from the observation perspective based on the real driving videos from multiple set perspectives.
[0075] After generating the target video from the set perspective, it further includes using the target video from the set perspective as the simulated driving video from the set perspective and training and / or testing the autonomous driving model based on the simulated driving video from the set perspective.
[0076] It should be noted that the autonomous driving model is not overly limited. For example, it can include a perception model, map construction, route planning, behavior decision-making model, etc.
[0077] Thus, the generated multi-view driving videos can be used to train and / or test the autonomous driving model, which helps to improve the accuracy of the autonomous driving model, reduce the acquisition cost of real driving information, and enhance the reliability of the autonomous driving model in extreme scenarios.
[0078] As another possible implementation, taking the traffic planning scenario as an example, the method further includes performing traffic planning based on the simulated driving videos from at least one set perspective to obtain traffic planning results. The traffic planning results are not limited too much. For example, they can include traffic flow prediction results, signal light optimization results, etc.
[0079] In the third case, taking the XR scenario as an example, the original video from the observation perspective is obtained, including obtaining the original extended reality videos from multiple set perspectives, and generating the original video from the observation perspective based on the original extended reality videos from multiple set perspectives.
[0080] After generating the target video from the set perspective, it further includes using the target video from the set perspective as the new extended reality video from the set perspective, where the new extended reality video from the set perspective is used for visual display by the extended reality device.
[0081] Thus, the generated target video from the set perspective can maintain the original three-dimensional features of the object. In the multi-view XR video generation scenario, it helps to improve the three-dimensional consistency among multi-view XR videos, avoid object deformation and depth distortion in multi-view XR videos, enable the construction of dynamic scenes in XR, and enable the visual display of multi-view XR videos.
[0082] In the fourth case, taking the robot navigation scenario as an example, the original video from the observation perspective is obtained, including obtaining the real robot navigation videos from multiple set perspectives, and generating the original video from the observation perspective based on the real robot navigation videos from multiple set perspectives.
[0083] After generating the target video from the set perspective, it further includes using the target video from the set perspective as the simulated robot navigation video from the set perspective, and training and / or testing the robot navigation model based on the simulated robot navigation video from the set perspective.
[0084] It should be noted that the robot navigation model is not limited too much. For example, it can include a perception model, map construction, route planning, behavior decision-making model, etc.
[0085] Therefore, the generated target video from the set perspective can maintain the original three-dimensional features of the object. In the scenario of multi-perspective robot navigation video generation, it helps to improve the three-dimensional consistency among multi-perspective robot navigation videos, avoid object deformation and depth distortion in multi-perspective robot navigation videos, and use the generated multi-perspective robot navigation videos to train and / or test the robot navigation model, which helps to improve the accuracy of the robot navigation model, improve the acquisition efficiency of robot navigation videos, and reduce the acquisition cost of real robot navigation information.
[0086] In the fifth case, taking a film and television scenario as an example, the original video from the observation perspective is obtained, including obtaining real-scene videos from multiple set perspectives, and based on the real-scene videos from multiple set perspectives, the original video from the observation perspective is generated.
[0087] After generating the target video from the set perspective, it further includes using the target video from the set perspective as the virtual-scene video from the set perspective, and based on the virtual-scene video from the set perspective, generating a film and television video, which is used for visual display by a film and television video playback device.
[0088] It should be noted that the real-scene video refers to the video of the real environment captured by a camera, and the virtual-scene video refers to the video of the virtual environment generated by a computer.
[0089] Therefore, the generated target video from the set perspective can maintain the original three-dimensional features of the object. In the scenario of multi-perspective virtual-scene video generation, it helps to improve the three-dimensional consistency among multi-perspective virtual-scene videos, avoid object deformation and depth distortion in multi-perspective virtual-scene videos, and use the generated multi-perspective virtual-scene videos for film and television production and realize the visual display of multi-perspective virtual-scene videos.
[0090] Figure 2 It is a flowchart showing a video generation method according to another exemplary embodiment. As Figure 2 shown, the video generation method of the embodiments of the present disclosure includes the following steps.
[0091] S201, obtain at least one frame from the original video from the observation perspective.
[0092] S202, based on the features of the image of the at least one frame from the original video in the two-dimensional space, determine the features of the voxels of the corresponding frame in the three-dimensional space.
[0093] The relevant content of steps S201 - S202 can be referred to the above embodiments and will not be elaborated here.
[0094] S203. Project the features of at least one frame of voxels in three-dimensional space onto the two-dimensional space corresponding to the set viewing angle to obtain the features of the image of the corresponding frame in the two-dimensional space corresponding to the set viewing angle.
[0095] For example, taking the set viewing angle as the front view as an example, the features of the voxels of the first frame in three-dimensional space can be projected onto the two-dimensional space corresponding to the front view to obtain the features of the image of the first frame in the two-dimensional space corresponding to the front view, and the features of the voxels of the second frame in three-dimensional space can be projected onto the two-dimensional space corresponding to the front view to obtain the features of the image of the second frame in the two-dimensional space corresponding to the front view.
[0096] It should be noted that there are no excessive restrictions on the projection method. For example, it may include ray projection, ray casting, etc.
[0097] There are no excessive restrictions on the image in the two-dimensional space corresponding to the set viewing angle. For example, it may include a semantic image, a depth image, a coordinate image, an MPI (Multi-Plane Images), etc. The semantic image carries the semantic information of the voxels corresponding to each pixel point. The depth image carries the distance information between the voxels corresponding to each pixel point and the acquisition device in the world coordinate system. The distance information between the voxels and the acquisition device in the world coordinate system is the depth information of the voxels. There are no excessive restrictions on the acquisition device. For example, it may include a vehicle-mounted camera, a camera, a lidar, etc. The coordinate image carries the position information of the voxels corresponding to each pixel point in the world coordinate system, and the multi-plane image carries the semantic information of the voxels corresponding to each pixel point under different depth information.
[0098] There are no excessive restrictions on the semantic information of the voxels. For example, it may include object category, action, pose, emotion, etc.
[0099] Projecting the features of at least one frame of voxels in three-dimensional space onto the two-dimensional space corresponding to the set viewing angle to obtain the features of the image of the corresponding frame in the two-dimensional space corresponding to the set viewing angle includes the following possible real-time methods:
[0100] Method 1: Project the semantic information of at least one frame of voxels in three-dimensional space onto the two-dimensional space corresponding to the set viewing angle to obtain the features of the semantic image of the corresponding frame in the two-dimensional space corresponding to the set viewing angle.
[0101] It should be noted that the features of the semantic image refer to the semantic information of the voxels corresponding to each pixel point.
[0102] Method 2: Project the depth information of at least one frame of voxels in three-dimensional space onto the two-dimensional space corresponding to the set viewing angle to obtain the features of the depth image of the corresponding frame in the two-dimensional space corresponding to the set viewing angle.
[0103] It should be noted that the features of the depth image refer to the depth information of the voxels corresponding to each pixel point.
[0104] Method 3: Project the coordinate information of the voxels of at least one frame in three-dimensional space onto the two-dimensional space corresponding to the set viewing angle to obtain the characteristics of the coordinate image of the corresponding frame in the two-dimensional space corresponding to the set viewing angle.
[0105] It should be noted that the characteristics of the coordinate image refer to the position information of the voxels corresponding to each pixel point in the world coordinate system.
[0106] Method 4: Project the semantic information of the voxels of at least one frame in three-dimensional space at the j-th depth information onto the j-th two-dimensional space corresponding to the set viewing angle to obtain the characteristics of the image of the corresponding frame in the j-th two-dimensional space corresponding to the set viewing angle, and the two-dimensional space corresponding to the set viewing angle corresponds one-to-one with the depth information.
[0107] It should be noted that the characteristics of the image in the j-th two-dimensional space refer to the semantic information of the voxels corresponding to each pixel point at the j-th depth information.
[0108] For example, 3 depth information can be set. Taking the first frame as an example, project the semantic information of the voxels of the first frame in three-dimensional space at the first depth information onto the first two-dimensional space corresponding to the set viewing angle to obtain the characteristics of the image of the first frame in the first two-dimensional space corresponding to the set viewing angle.
[0109] Project the semantic information of the voxels of the first frame in three-dimensional space at the second depth information onto the second two-dimensional space corresponding to the set viewing angle to obtain the characteristics of the image of the first frame in the second two-dimensional space corresponding to the set viewing angle.
[0110] Project the semantic information of the voxels of the first frame in three-dimensional space at the third depth information onto the third two-dimensional space corresponding to the set viewing angle to obtain the characteristics of the image of the first frame in the third two-dimensional space corresponding to the set viewing angle.
[0111] Optionally, before projecting the characteristics of the voxels of at least one frame in three-dimensional space onto the two-dimensional space corresponding to the set viewing angle, it further includes adjusting the characteristics of the voxels of at least one frame in three-dimensional space based on the generation requirements of the target video at the set viewing angle.
[0112] Projecting the characteristics of the voxels of at least one frame in three-dimensional space onto the two-dimensional space corresponding to the set viewing angle to obtain the characteristics of the image of the corresponding frame in the two-dimensional space corresponding to the set viewing angle includes projecting the final characteristics of the voxels of at least one frame in three-dimensional space onto the two-dimensional space corresponding to the set viewing angle to obtain the characteristics of the image of the corresponding frame in the two-dimensional space corresponding to the set viewing angle.
[0113] S204: Generate the target video at the set viewing angle based on the characteristics of the images of each frame in the two-dimensional space corresponding to the set viewing angle.
[0114] For example, taking the set viewing angle as the front view as an example, the target video in the front view can be generated based on the features of the images of the first to twentieth frames in the original video from the bird's-eye view in the two-dimensional space corresponding to the front view.
[0115] It should be noted that for the relevant content of step S204, reference can be made to the relevant content of step S103, which will not be elaborated here.
[0116] Optionally, generating the target video in the set viewing angle based on the features of the images of each frame in the two-dimensional space corresponding to the set viewing angle includes obtaining the description text of the generation requirements of the target video in the set viewing angle, and generating the target video in the set viewing angle based on the features of the images of each frame in the two-dimensional space corresponding to the set viewing angle and the description text.
[0117] For example, input the features of the images of each frame in the two-dimensional space corresponding to the set viewing angle and the description text into the second generation model, and the second generation model outputs the target video in the set viewing angle.
[0118] The video generation method provided by the embodiments of the present disclosure projects the features of the voxels of at least one frame in the three-dimensional space onto the two-dimensional space corresponding to the set viewing angle, obtains the features of the images of the corresponding frames in the two-dimensional space corresponding to the set viewing angle, and generates the target video in the set viewing angle based on the features of the images of each frame in the two-dimensional space corresponding to the set viewing angle. Thus, the features of the voxels can be projected according to the set viewing angle to obtain the features of the images in the set viewing angle, so as to generate the target video in the set viewing angle.
[0119] Based on any of the above embodiments, the image in the two-dimensional space corresponding to the set viewing angle includes a multi-modal image.
[0120] Optionally, the image in the two-dimensional space corresponding to the set viewing angle includes at least two modal images among semantic images, depth images, coordinate images, and multi-plane images.
[0121] Figure 3 It is a flowchart showing a video generation method according to another exemplary embodiment. As Figure 3 shown, the video generation method of the embodiments of the present disclosure includes the following steps.
[0122] S301, obtain at least one frame in the original video from the observation viewing angle.
[0123] S302, determine the features of the voxels of the corresponding frames in the three-dimensional space based on the features of the images of at least one frame in the original video in the two-dimensional space.
[0124] S303, project the features of the voxels of at least one frame in the three-dimensional space onto the two-dimensional space corresponding to the set viewing angle, and obtain the features of the images of the corresponding frames in the two-dimensional space corresponding to the set viewing angle.
[0125] For the relevant content of steps S301 - S303, reference can be made to the above - mentioned embodiments, which will not be elaborated here.
[0126] S304, extract the features of the i - th modal image in the two - dimensional space corresponding to the set perspective for at least one frame, to obtain the i - th modal feature of the corresponding frame in the set perspective.
[0127] For example, if the set perspective is the front view, and the images in the two - dimensional space corresponding to the set perspective include semantic images, depth images, coordinate images, and multi - plane images. Taking the first frame as an example, extract the features of the semantic image in the two - dimensional space corresponding to the front view of the first frame, to obtain the first modal feature of the first frame in the front view.
[0128] Extract the features of the depth image in the two - dimensional space corresponding to the front view of the first frame, to obtain the second modal feature of the first frame in the front view.
[0129] Extract the features of the coordinate image in the two - dimensional space corresponding to the front view of the first frame, to obtain the third modal feature of the first frame in the front view.
[0130] Extract the features of the multi - plane image in the two - dimensional space corresponding to the front view of the first frame, to obtain the fourth modal feature of the first frame in the front view.
[0131] Optionally, extracting the features of the i - th modal image in the two - dimensional space corresponding to the set perspective for at least one frame to obtain the i - th modal feature of the corresponding frame in the set perspective includes extracting the features of the i - th modal image in the two - dimensional space corresponding to the set perspective for at least one frame according to a unified feature extraction method, to obtain the i - th modal feature of the corresponding frame in the set perspective, so that the multi - modal features of each frame in the set perspective are spatially aligned. Thus, the features of the multi - modal images in the two - dimensional space corresponding to the set perspective for at least one frame can be extracted according to a unified feature extraction method, making the multi - modal features of each frame in the set perspective spatially aligned, which helps to improve the visual consistency and authenticity of the generated target video.
[0132] It should be noted that the feature extraction method is not overly limited. For example, it can include the down - sampling factor.
[0133] Optionally, extracting the features of the i - th modal image in the two - dimensional space corresponding to the set perspective for at least one frame to obtain the i - th modal feature of the corresponding frame in the set perspective includes, in response to the i - th modal image being any one of a semantic image, a depth image, and a coordinate image, inputting the features of the i - th modal image in the two - dimensional space corresponding to the set perspective for at least one frame into the first encoder, and outputting the i - th modal feature of the corresponding frame in the set perspective by the first encoder.
[0134] Alternatively, in response to the i-th modal image being a multi-plane image, interpolate the features of at least one frame of the i-th modal image in the two-dimensional space corresponding to the set perspective to obtain the interpolated features of the corresponding frame in the set perspective, and input the interpolated features of at least one frame in the set perspective into the second encoder, and the second encoder outputs the i-th modal feature of the corresponding frame in the set perspective.
[0135] It should be noted that no excessive limitations are imposed on the first encoder and the second encoder. For example, the first encoder is the encoder of VAEs, and the second encoder is the MPI encoder. For example, the interpolated features of each frame in the set perspective are spatially aligned with the multi-modal features of each frame in the set perspective.
[0136] S305, perform a fusion process on the multi-modal features of any frame in the set perspective to obtain the fused feature of any frame in the set perspective.
[0137] It should be noted that the fused feature is generated based on the multi-modal features and carries rich and diverse information, which helps to improve the quality of the generated target video and enables the generated target video to maintain the original multi-modal features of the object, that is, to maintain visual consistency and authenticity.
[0138] For example, when the set perspective is the front view, the images in the two-dimensional space corresponding to the set perspective include semantic images, depth images, coordinate images, and multi-plane images. Taking the first frame as an example, perform a fusion process on the first to fourth modal features of the first frame in the front view to obtain the fused feature of the first frame in the front view.
[0139] It should be noted that no excessive limitations are imposed on the fusion processing method.
[0140] Optionally, perform a fusion process on the multi-modal features of any frame in the set perspective to obtain the fused feature of any frame in the set perspective, including inputting the multi-modal features of any frame in the set perspective into a convolutional layer, and the convolutional layer outputs the fused feature of any frame in the set perspective. It should be noted that no excessive limitations are imposed on the size of the convolutional kernel of the convolutional layer. For example, it can be 1*1.
[0141] S306, generate a target video in the set perspective based on the fused features of each frame in the set perspective.
[0142] For example, taking the set perspective as the front view as an example, a target video in the front view can be generated based on the fused features of the first to 20th frames in the front view of the original video from the bird's-eye view.
[0143] It should be noted that for the relevant content of step S306, reference can be made to the relevant content of step S103, which will not be elaborated here.
[0144] Optionally, based on the fusion features of each frame from a set perspective, a target video from the set perspective is generated, including obtaining a description text of the generation requirements of the target video from the set perspective, and generating the target video from the set perspective based on the fusion features of each frame from the set perspective and the description text.
[0145] For example, the fusion features of each frame from the set perspective and the description text are input into a second generation model, and the second generation model outputs the target video from the set perspective.
[0146] The video generation method provided by the embodiments of the present disclosure extracts the features of the i-th modality image of at least one frame in the two-dimensional space corresponding to the set perspective to obtain the i-th modality feature of the corresponding frame from the set perspective, performs a fusion process on the multi-modal features of any frame from the set perspective to obtain the fusion feature of any frame from the set perspective, and generates a target video from the set perspective based on the fusion features of each frame from the set perspective. Thus, the fusion feature is generated based on the multi-modal features and carries rich and diverse information, enabling the generated target video to maintain the original multi-modal features of the object, that is, to maintain visual consistency and authenticity, which helps to improve the quality of the generated target video.
[0147] Figure 4 is a schematic flowchart of a video generation method shown according to another exemplary embodiment, as Figure 4 shown, the video generation method of the embodiments of the present disclosure includes the following steps.
[0148] S401, obtain at least one frame in the original video from the observation perspective.
[0149] S402, based on the features of the image of at least one frame in the original video in the two-dimensional space, determine the features of the voxels of the corresponding frame in the three-dimensional space.
[0150] S403, project the features of the voxels of at least one frame in the three-dimensional space onto the two-dimensional space corresponding to the set perspective to obtain the features of the image of the corresponding frame in the two-dimensional space corresponding to the set perspective.
[0151] S404, extract the features of the i-th modality image of at least one frame in the two-dimensional space corresponding to the set perspective to obtain the i-th modality feature of the corresponding frame from the set perspective.
[0152] S405, perform a fusion process on the multi-modal features of any frame from the set perspective to obtain the fusion feature of any frame from the set perspective.
[0153] For the relevant content of steps S401 - S405, reference can be made to the above embodiments and will not be elaborated here.
[0154] S406. Optimize at least one frame of the fused features at a set perspective to obtain the optimized features of the corresponding frame at the set perspective.
[0155] It should be noted that there are no excessive restrictions on the optimization process, such as including jitter elimination, smoothing processing, etc.
[0156] In the related art, when fusing multi-modal features such as semantics and depth, there may be conflicts in the fused features of adjacent frames, resulting in inter-frame jitter problems in the generated video. For example, taking a driving scenario as an example, there are sudden changes in the vehicle positions in adjacent frames of the generated driving video.
[0157] In this embodiment, the fused features can be optimized, such as jitter elimination, smoothing processing, etc., to avoid conflicts in the optimized features of adjacent frames, thereby avoiding inter-frame jitter problems in the generated target video.
[0158] Optionally, optimizing at least one frame of the fused features at a set perspective to obtain the optimized features of the corresponding frame at the set perspective includes inputting at least one frame of the fused features at the set perspective into an adapter. Among them, the adapter includes a spatial convolutional layer, a temporal convolutional layer, and a temporal self-attention layer. The spatial convolutional layer extracts spatial features from at least one frame of the fused features at the set perspective to obtain the spatial features of the corresponding frame at the set perspective. The temporal convolutional layer extracts spatio-temporal features from the spatial features of at least one frame at the set perspective to obtain the spatio-temporal features of the corresponding frame at the set perspective. The temporal self-attention layer enhances the attention of the spatio-temporal features of at least one frame at the set perspective to obtain the optimized features of the corresponding frame at the set perspective.
[0159] In this embodiment, an adapter composed of a spatial convolutional layer, a temporal convolutional layer, and a temporal self-attention layer can be used to optimize the fused features, dynamically balance the weights of multi-modal features, achieve adaptive fusion of multi-modal features, ensure the long-range consistency of the generated video and the smoothness of inter-frame transitions, and improve the temporal stability of the generated long video.
[0160] Optionally, the spatial convolutional layer includes units such as upsampling, two-dimensional convolution, and group normalization. The temporal convolutional layer includes units such as three-dimensional convolution and group normalization. The temporal self-attention layer includes a feed-forward network and a temporal self-attention unit.
[0161] S407. Generate a target video at the set perspective based on the optimized features of each frame at the set perspective.
[0162] Optionally, based on the optimized features of each frame from a set perspective, a target video from the set perspective is generated, including inputting a pure noise image and the optimized features of each frame from the set perspective into a diffusion model. The diffusion model denoises the pure noise image based on the optimized features of at least one frame from the set perspective to obtain the predicted features of the corresponding frame from the set perspective. Then, the predicted features of each frame from the set perspective are input into a decoder, and the decoder outputs the target video from the set perspective. Thus, the optimized features of at least one frame from the set perspective can be used to guide the diffusion model to denoise the pure noise image, obtain the predicted features of the corresponding frame from the set perspective, and further obtain the target video from the set perspective.
[0163] For example, denoising can be performed through a diffusion model based on a spatially-flattened attention mechanism, improving the consistency of multi-perspective videos in multi-perspective video generation scenarios.
[0164] It should be noted that there are no excessive restrictions on the decoder. For example, it can include a decoder of VAEs.
[0165] For example, the images in the two-dimensional space corresponding to the set perspective include semantic images, depth images, coordinate images, and multi-plane images. As Figure 6 shown, the second generation model includes a feature processing model, a diffusion model, an adapter, and a decoder. Among them, the feature processing model includes a projection network, a first encoder, an interpolation network, a second encoder, and a convolutional layer. The diffusion model includes a Control Block, a third encoder, and a Base Block.
[0166] Input the original video from the observation perspective into the first generation model. The first generation model outputs the features of each frame in the voxels of the three-dimensional space. Input the features of at least one frame in the voxels of the three-dimensional space into the projection network, and the projection network outputs the features of the corresponding frame in the image in the two-dimensional space corresponding to the set perspective.
[0167] Input the features of at least one frame in the semantic image in the two-dimensional space corresponding to the set perspective into the first encoder, and the first encoder outputs the first modal feature of the corresponding frame from the set perspective.
[0168] Input the features of at least one frame in the depth image in the two-dimensional space corresponding to the set perspective into the first encoder, and the first encoder outputs the second modal feature of the corresponding frame from the set perspective.
[0169] Input the features of at least one frame in the coordinate image in the two-dimensional space corresponding to the set perspective into the first encoder, and the first encoder outputs the third modal feature of the corresponding frame from the set perspective.
[0170] Input the features of at least one multi-plane image in the two-dimensional space corresponding to the set perspective into the interpolation network, and the interpolation network outputs the interpolation features of the corresponding frame at the set perspective. Input the interpolation features of at least one frame at the set perspective into the second encoder, and the second encoder outputs the fourth modal features of the corresponding frame at the set perspective.
[0171] Input the first and second modal features of any frame at the set perspective into the convolutional layer, and the convolutional layer outputs the fused features of any frame at the set perspective. Alternatively, input the third and fourth modal features of any frame at the set perspective into the convolutional layer, and the convolutional layer outputs the fused features of any frame at the set perspective.
[0172] Input the fused features of at least one frame at the set perspective into the control block, and the control block outputs the processed features of the corresponding frame at the set perspective. Input the processed features of at least one frame at the set perspective into the adapter, and the adapter outputs the optimized features of the corresponding frame at the set perspective.
[0173] Input the description text of the generation requirement of the target video at the set perspective into the third encoder, and the third encoder outputs the text features of the description text.
[0174] Input the pure noise image, the optimized features of at least one frame at the set perspective, and the text features into the control block, and the control block outputs the predicted features of the corresponding frame at the set perspective.
[0175] Input the predicted features of each frame at the set perspective into the decoder, and the decoder outputs the target video at the set perspective.
[0176] The video generation method provided by the embodiments of the present disclosure performs fusion processing on the multi-modal features of any frame at the set perspective to obtain the fused features of any frame at the set perspective, and generates the target video at the set perspective based on the optimized features of each frame at the set perspective. Thus, the fused features can be optimized, such as jitter elimination, smoothing processing, etc., to avoid conflicts between the optimized features of adjacent frames, and further avoid the problem of frame jitter in the generated target video.
[0177] Figure 5 It is a schematic flowchart of a training method of a video generation model shown according to an exemplary embodiment, as Figure 5 shown, the training method of the video generation model of the embodiments of the present disclosure includes the following steps.
[0178] S501, based on the first sample videos under multiple set perspectives, generate the second sample video under the observation perspective, and obtain at least one frame in the second sample video.
[0179] S502. Determine the sample features of the voxels corresponding to at least one frame in the three-dimensional space based on the sample features of the images of the at least one frame in the second sample video in the two-dimensional space.
[0180] For the relevant content of steps S501 - S502, reference can be made to the above embodiments and will not be elaborated here.
[0181] S503. Train the first generation model based on the second sample video and the sample features of the voxels of each frame in the three-dimensional space.
[0182] Optionally, training the first generation model based on the second sample video and the sample features of the voxels of each frame in the three-dimensional space includes inputting the second sample video into the first generation model, outputting the predicted features of the voxels of each frame in the three-dimensional space by the first generation model, and training the first generation model based on the sample features of the voxels of each frame in the three-dimensional space and the predicted features of the voxels of each frame in the three-dimensional space.
[0183] S504. Train the second generation model based on the first sample video under at least one set perspective and the sample features of the voxels of each frame in the three-dimensional space.
[0184] Optionally, training the second generation model based on the first sample video under at least one set perspective and the sample features of the voxels of each frame in the three-dimensional space includes inputting the sample features of the voxels of each frame in the three-dimensional space into the second generation model, outputting the predicted features of the predicted video under at least one set perspective by the second generation model, and training the second generation model based on the sample features of the first sample video under the set perspective and the predicted features of the predicted video under the set perspective.
[0185] Optionally, training the second generation model based on the sample features of the first sample video under the set perspective and the predicted features of the predicted video under the set perspective includes obtaining the total difference between the sample features of the first sample video under the set perspective and the predicted features of the predicted video under the set perspective, determining the target difference from the total difference, where the target difference is the difference between the sample features corresponding to the foreground region in the first sample video under the set perspective and the predicted features corresponding to the foreground region, amplifying the target difference in the total difference to update the total difference, and training the second generation model based on the updated total difference. Thus, the target difference corresponding to the foreground region can be amplified, so that the second generation model can focus on strengthening the generation quality of the foreground region during the training process, which helps to improve the generation quality of the foreground region generated by the second generation model.
[0186] For example, taking the driving scenario as an example, the foreground region may include the image regions where vehicles and pedestrians are located.
[0187] Optionally, a target difference is determined from the total difference, including obtaining masks of foreground regions of each frame, and determining the target difference from the total difference based on the masks of foreground regions of each frame.
[0188] Optionally, the target difference in the total difference is amplified to update the total difference, including obtaining the product of the target difference and a set weight, obtaining the sum value of the total difference and the product, and updating the total difference to the sum value.
[0189] Optionally, continuing with Figure 6 as an example, the method further includes inputting a first sample video under a set perspective into a first encoder, and outputting sample features of the first sample video under the set perspective by the first encoder, so that the multi-class features of each obtained video are spatially aligned.
[0190] It should be noted that step S503 and / or step S504 can be executed, and the execution order of steps S503 and S504 is not overly limited. For example, they can be executed serially or in parallel.
[0191] The training method of the video generation model provided by the embodiments of the present disclosure is based on first sample videos under multiple set perspectives, generates a second sample video under an observation perspective, and obtains at least one frame in the second sample video. Based on the sample features of the image of at least one frame in the second sample video in a two-dimensional space, the sample features of the voxels of the corresponding frame in a three-dimensional space are determined. Based on the second sample video and the sample features of the voxels of each frame in the three-dimensional space, the first generation model is trained, and / or based on the first sample videos under at least one set perspective and the sample features of the voxels of each frame in the three-dimensional space, the second generation model is trained. The video generation model is used to generate a target video under a set perspective. Thus, during the training process, the first generation model can learn the correlation between the second sample video and the sample features of the voxels of each frame in the three-dimensional space, so that the trained first generation model can obtain the features of the voxels of the corresponding frame in the three-dimensional space based on at least one frame in the original video. In addition, during the training process, the second generation model can learn the correlation between the first sample videos under the set perspective and the sample features of the voxels of each frame in the three-dimensional space, so that the trained second generation model can generate a target video under the set perspective based on the features of the voxels of each frame in the three-dimensional space. Finally, the trained video generation model can generate a target video under the set perspective based on the original video under the observation perspective.
[0192] Figure 7 is a schematic structural diagram of a video generation device shown according to an exemplary embodiment. Referring to Figure 7 this, the video generation device 100 of the embodiments of the present disclosure includes: an acquisition module 110, a determination module 120, and a generation module 130.
[0193] An acquisition module 110, configured to acquire at least one frame in an original video from an observation perspective;
[0194] A determination module 120, configured to determine the characteristics of voxels of a corresponding frame in three-dimensional space based on the characteristics of an image of at least one frame in the original video in two-dimensional space;
[0195] A generation module 130, configured to generate a target video from a set perspective based on the characteristics of voxels of each frame in the three-dimensional space.
[0196] In some possible implementation manners, before acquiring at least one frame in the original video from the observation perspective, the acquisition module 110 is further configured to: acquire a first instruction, where the first instruction is triggered based on an operation on a target vehicle;
[0197] The acquisition module 110 is further configured to: acquire a real driving video of the target vehicle from the observation perspective as the original video.
[0198] In some possible implementation manners, the first instruction is triggered based on a touch operation and / or a voice input operation on the target vehicle.
[0199] In some possible implementation manners, the target video from the set perspective is used for visual display of the target vehicle.
[0200] In some possible implementation manners, the generation module 130 is further configured to: project the characteristics of voxels of at least one frame in three-dimensional space onto a two-dimensional space corresponding to the set perspective to obtain the characteristics of an image of the corresponding frame in the two-dimensional space corresponding to the set perspective; generate the target video from the set perspective based on the characteristics of the images of each frame in the two-dimensional space corresponding to the set perspective.
[0201] In some possible implementation manners, the image in the two-dimensional space corresponding to the set perspective includes a multi-modal image.
[0202] In some possible implementation manners, the image in the two-dimensional space corresponding to the set perspective includes at least two modal images among a semantic image, a depth image, a coordinate image, and a multi-plane image.
[0203] In some possible implementation manners, the generation module 130 is further configured to: perform feature extraction on the characteristics of the i-th modal image of at least one frame in the two-dimensional space corresponding to the set perspective to obtain the i-th modal characteristics of the corresponding frame from the set perspective; perform fusion processing on the multi-modal characteristics of any frame from the set perspective to obtain the fusion characteristics of the any frame from the set perspective; generate the target video from the set perspective based on the fusion characteristics of each frame from the set perspective.
[0204] In some possible embodiments, the generation module 130 is further configured to: extract features of at least one frame of the i-th modal image in the two-dimensional space corresponding to the set perspective according to a unified feature extraction method, so as to obtain the i-th modal feature of the corresponding frame in the set perspective, so that the multi-modal features of each frame in the set perspective are spatially aligned.
[0205] In some possible embodiments, the generation module 130 is further configured to: in response to the i-th modal image being any one of a semantic image, a depth image, and a coordinate image, input the features of at least one frame of the i-th modal image in the two-dimensional space corresponding to the set perspective into the first encoder, and output the i-th modal feature of the corresponding frame in the set perspective by the first encoder; or,
[0206] In response to the i-th modal image being a multi-plane image, perform interpolation processing on the features of at least one frame of the i-th modal image in the two-dimensional space corresponding to the set perspective to obtain the interpolation feature of the corresponding frame in the set perspective; input the interpolation features of at least one frame in the set perspective into the second encoder, and output the i-th modal feature of the corresponding frame in the set perspective by the second encoder.
[0207] In some possible embodiments, the generation module 130 is further configured to: input the multi-modal features of any frame in the set perspective into a convolutional layer, and output the fused feature of any frame in the set perspective by the convolutional layer.
[0208] In some possible embodiments, the generation module 130 is further configured to: optimize the fused features of at least one frame in the set perspective to obtain the optimized features of the corresponding frame in the set perspective; generate the target video in the set perspective based on the optimized features of each frame in the set perspective.
[0209] In some possible embodiments, the generation module 130 is further configured to: input the fused features of at least one frame in the set perspective into an adapter, where the adapter includes a spatial convolutional layer, a temporal convolutional layer, and a temporal self-attention layer; perform spatial feature extraction on the fused features of at least one frame in the set perspective through the spatial convolutional layer to obtain the spatial features of the corresponding frame in the set perspective; perform spatio-temporal feature extraction on the spatial features of at least one frame in the set perspective through the temporal convolutional layer to obtain the spatio-temporal features of the corresponding frame in the set perspective; perform attention enhancement on the spatio-temporal features of at least one frame in the set perspective through the temporal self-attention layer to obtain the optimized features of the corresponding frame in the set perspective.
[0210] In some possible embodiments, the generating module 130 is further configured to: input the pure noise image and the optimized features of each frame at the set viewing angle into a diffusion model, and denoise the pure noise image by the diffusion model based on the optimized features of at least one frame at the set viewing angle to obtain the predicted features of the corresponding frame at the set viewing angle; input the predicted features of each frame at the set viewing angle into a decoder, and output the target video at the set viewing angle by the decoder.
[0211] In some possible embodiments, the generating module 130 is further configured to: obtain a description text of the generation requirement of the target video at the set viewing angle; generate the target video at the set viewing angle based on the features of the voxels of each frame in the three-dimensional space and the description text.
[0212] In some possible embodiments, the generating module 130 is further configured to: adjust the features of the voxels of at least one frame in the three-dimensional space based on the generation requirement of the target video at the set viewing angle; generate the target video at the set viewing angle based on the final features of the voxels of each frame in the three-dimensional space.
[0213] In some possible embodiments, the obtaining module 110 is further configured to: obtain real driving videos at multiple set viewing angles, and obtain the original video at the observation viewing angle based on the real driving videos at the multiple set viewing angles;
[0214] After generating the target video at the set viewing angle, the generating module 130 is further configured to: use the target video at the set viewing angle as the simulated driving video at the set viewing angle; train and / or test an autonomous driving model based on the simulated driving video at the set viewing angle.
[0215] In some possible embodiments, the obtaining module 110 is further configured to: obtain original extended reality videos at multiple set viewing angles, and obtain the original video at the observation viewing angle based on the original extended reality videos at the multiple set viewing angles;
[0216] After generating the target video at the set viewing angle, the generating module 130 is further configured to: use the target video at the set viewing angle as the new extended reality video at the set viewing angle, and the new extended reality video at the set viewing angle is used for visual display by an extended reality device.
[0217] In some possible embodiments, the obtaining module 110 is further configured to: obtain real robot navigation videos at multiple set viewing angles, and obtain the original video at the observation viewing angle based on the real robot navigation videos at the multiple set viewing angles;
[0218] After generating the target video from the set perspective, the generating module 130 is further configured to: use the target video from the set perspective as the simulated robot navigation video from the set perspective; and train and / or test the robot navigation model based on the simulated robot navigation video from the set perspective.
[0219] In some possible implementation manners, the obtaining module 110 is further configured to: obtain real-scene videos from multiple set perspectives, and obtain the original video from the observation perspective based on the real-scene videos from the multiple set perspectives;
[0220] After generating the target video from the set perspective, the generating module 130 is further configured to: use the target video from the set perspective as the virtual-scene video from the set perspective; and generate a film and television video based on the virtual-scene video from the set perspective, where the film and television video is used for visual display by a film and television video playing device.
[0221] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0222] The video generation device provided by the embodiments of the present disclosure obtains at least one frame from the original video from the observation perspective, determines the feature of the voxel in the three-dimensional space corresponding to the corresponding frame based on the feature of the image of the at least one frame from the original video in the two-dimensional space, and generates the target video from the set perspective based on the feature of the voxel in the three-dimensional space of each frame. Thus, the feature of the voxel retains the three-dimensional information of the object, so that the generated target video from the set perspective can maintain the original three-dimensional feature of the object, that is, maintain visual consistency and authenticity. In the multi-perspective video generation scenario, it helps to improve the three-dimensional consistency between multi-perspective videos and improve the quality of the generated videos, and is applicable to the generation scenario of multi-perspective autonomous driving videos.
[0223] Figure 8 is a schematic structural diagram of a training device for a video generation model shown according to an exemplary embodiment. Refer to Figure 8 , the training device 200 for the video generation model according to the embodiments of the present disclosure includes: an obtaining module 210, a determining module 220, a first training module 230, and / or a second training module 240.
[0224] The obtaining module 210 is configured to generate a second sample video from the observation perspective based on the first sample videos from multiple set perspectives, and obtain at least one frame of the second sample video;
[0225] A determination module 220, configured to determine sample features of voxels of a corresponding frame in three-dimensional space based on sample features of an image of at least one frame in the second sample video in two-dimensional space;
[0226] A first training module 230, configured to train a first generation model based on the second sample video and sample features of voxels of each frame in three-dimensional space; and / or,
[0227] A second training module 240, configured to train a second generation model based on the first sample video under at least one of the set perspectives and sample features of voxels of each frame in three-dimensional space.
[0228] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0229] It should be noted that Figure 8 The training device of the video generation model shown includes both a first training module and a second training module, which is only an example of the training device of the video generation model in the embodiments of the present disclosure, and does not limit the training device of the video generation model in the embodiments of the present disclosure.
[0230] The training device of the video generation model provided in the embodiments of the present disclosure generates a second sample video from the first sample videos under multiple set perspectives, obtains at least one frame in the second sample video, determines sample features of voxels of the corresponding frame in three-dimensional space based on the sample features of the image of at least one frame in the second sample video in two-dimensional space, trains a first generation model based on the second sample video and sample features of voxels of each frame in three-dimensional space, and / or trains a second generation model based on the first sample video under at least one set perspective and sample features of voxels of each frame in three-dimensional space. The video generation model is used to generate a target video under a set perspective. Thus, during the training process, the first generation model can learn the correlation between the second sample video and the sample features of voxels of each frame in three-dimensional space, so that the trained first generation model can obtain the features of voxels of the corresponding frame in three-dimensional space based on at least one frame in the original video. In addition, during the training process, the second generation model can learn the correlation between the first sample video under the set perspective and the sample features of voxels of each frame in three-dimensional space, so that the trained second generation model can generate a target video under the set perspective based on the features of voxels of each frame in three-dimensional space. Finally, the trained video generation model can generate a target video under the set perspective based on the original video under the observation perspective.
[0231] To implement the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the video generation method provided by the present disclosure are implemented, and / or the steps of the training method of the video generation model provided by the present disclosure are implemented.
[0232] To implement the above embodiments, the present disclosure also provides a vehicle, including a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the steps of the video generation method provided by the present disclosure, and / or the steps of the training method of the video generation model provided by the present disclosure.
[0233] Figure 9 FIG. 7 is a schematic structural diagram of a vehicle shown according to an exemplary embodiment. For example, vehicle 300 may be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. Vehicle 300 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0234] Refer to Figure 9 , vehicle 300 may include various subsystems. For example, the infotainment system 310, the perception system 320, the decision control system 330, the drive system 340, and the computing platform 350. Among them, vehicle 300 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of vehicle 300 may be interconnected by wired or wireless means.
[0235] In some embodiments, the infotainment system 310 may include a communication system, an entertainment system, and a navigation system, etc.
[0236] The perception system 320 may include several sensors for sensing information about the environment around vehicle 300. For example, the perception system 320 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), a lidar, a millimeter wave radar, an ultrasonic radar, and a camera device.
[0237] The decision control system 330 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.
[0238] The drive system 340 may include components that provide motive power for the vehicle 300. In one embodiment, the drive system 340 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.
[0239] Some or all functions of the vehicle 300 are controlled by the computing platform 350. The computing platform 350 may include at least one processor 351 and a memory 352, and the processor 351 may execute instructions 353 stored in the memory 352.
[0240] The processor 351 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0241] The memory 352 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0242] In addition to the instructions 353, the memory 352 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 352 can be used by the computing platform 350.
[0243] In an embodiment of the present disclosure, the processor 351 may execute the instructions 353 to implement all or part of the steps of the video generation method provided by the present disclosure, and / or implement all or part of the steps of the training method of the video generation model provided by the present disclosure.
[0244] The vehicle according to the embodiments of the present disclosure acquires at least one frame in the original video from the observation perspective, determines the characteristics of the voxels of the corresponding frame in the three-dimensional space based on the characteristics of the images of the at least one frame in the original video in the two-dimensional space, and generates the target video from the set perspective based on the characteristics of the voxels of each frame in the three-dimensional space. Thus, the characteristics of the voxels retain the three-dimensional information of the object, enabling the generated target video from the set perspective to maintain the original three-dimensional characteristics of the object, that is, to maintain visual consistency and authenticity. In the multi-perspective video generation scenario, it helps to improve the three-dimensional consistency among multi-perspective videos and improve the quality of the generated videos, and is applicable to the generation scenario of multi-perspective autonomous driving videos.
[0245] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the video generation method provided by the present disclosure are implemented, and / or the steps of the training method of the video generation model provided by the present disclosure are implemented.
[0246] Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0247] To implement the above embodiments, the present disclosure also proposes a chip, which includes an interface circuit and a processing circuit that are coupled to each other. The interface circuit is used to input or output signals, and the processing circuit is configured to implement the steps of the video generation method provided by the present disclosure, and / or the steps of the training method of the video generation model provided by the present disclosure.
[0248] Figure 10 It is a schematic structural diagram of a chip shown according to an exemplary embodiment. Reference may be made to Figure 10 the schematic structural diagram of the chip 400 shown, but not limited thereto.
[0249] The chip 400 includes a processing circuit 410, and the processing circuit 410 is configured to execute any of the above video generation methods and / or any of the training methods of the video generation model.
[0250] In some embodiments, the chip 400 further includes one or more interface circuits 420. Optionally, the interface circuit 420 is connected to the memory 430. The interface circuit 420 may be used to receive signals from the memory 430 or other devices, and the interface circuit 420 may be used to send signals to the memory 430 or other devices. For example, the interface circuit 420 may read the instructions stored in the memory 430 and send the instructions to the processing circuit 410.
[0251] In some embodiments, the interface circuit 420 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 410 performs other steps.
[0252] In some embodiments, terms such as interface circuit, interface, transceiver pin, transceiver, etc. may be used interchangeably.
[0253] In some embodiments, the chip 400 further includes one or more memories 430 for storing instructions. Optionally, all or part of the memory 430 may be outside the chip 400.
[0254] To implement the above embodiments, the present disclosure also proposes a computer program product, including a computer program, which when executed by a processor, implements the steps of the video generation method provided by the present disclosure, and / or implements the steps of the training method of the video generation model provided by the present disclosure.
[0255] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0256] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A video generation method, characterized in that, Including: Obtaining at least one frame in the original video from the observation perspective; Determining the feature of the voxel of the corresponding frame in the three-dimensional space based on the feature of the image of at least one frame in the original video in the two-dimensional space; Generating a target video from the set perspective based on the feature of the voxel of each frame in the three-dimensional space.
2. The method according to claim 1, wherein Before obtaining at least one frame in the original video from the observation perspective, it further includes: Obtaining a first instruction, where the first instruction is triggered based on the operation on the target vehicle; Obtaining the original video includes: Obtaining the real driving video of the target vehicle from the observation perspective as the original video.
3. The method according to claim 2, wherein The first instruction is triggered based on the touch operation and / or voice input operation on the target vehicle.
4. The method according to claim 1, wherein The target video from the set perspective is used for visual display of the target vehicle.
5. The method according to claim 1, characterized in that, The generating a target video from the set perspective based on the feature of the voxel of each frame in the three-dimensional space includes: Projecting the feature of the voxel of at least one frame in the three-dimensional space onto the two-dimensional space corresponding to the set perspective to obtain the feature of the image of the corresponding frame in the two-dimensional space corresponding to the set perspective; Generating the target video from the set perspective based on the feature of the image of each frame in the two-dimensional space corresponding to the set perspective.
6. The method according to claim 5, characterized in that, The image in the two-dimensional space corresponding to the set perspective includes a multi-modal image.
7. The method according to claim 6, characterized in that, The image in the two-dimensional space corresponding to the set perspective includes at least two-modal images among semantic image, depth image, coordinate image, and multi-plane image.
8. The method according to claim 6, wherein The generating the target video from the set perspective based on the feature of the image of each frame in the two-dimensional space corresponding to the set perspective includes: Performing feature extraction on the feature of the i-th modal image of at least one frame in the two-dimensional space corresponding to the set perspective to obtain the i-th modal feature of the corresponding frame in the set perspective; Performing fusion processing on the multi-modal features of any frame in the set perspective to obtain the fusion feature of the any frame in the set perspective; Generating the target video from the set perspective based on the fusion feature of each frame in the set perspective.
9. The method according to claim 8, wherein The performing feature extraction on the feature of the i-th modal image of at least one frame in the two-dimensional space corresponding to the set perspective to obtain the i-th modal feature of the corresponding frame in the set perspective includes: Performing feature extraction on the feature of the i-th modal image of at least one frame in the two-dimensional space corresponding to the set perspective in accordance with a unified feature extraction method to obtain the i-th modal feature of the corresponding frame in the set perspective, so as to achieve spatial alignment of the multi-modal features of each frame in the set perspective.
10. The method according to claim 8, characterized in that, The performing feature extraction on the feature of the i-th modal image of at least one frame in the two-dimensional space corresponding to the set perspective to obtain the i-th modal feature of the corresponding frame in the set perspective includes: In response to the i-th modal image being any one of a semantic image, a depth image, and a coordinate image, inputting the feature of the i-th modal image of at least one frame in the two-dimensional space corresponding to the set perspective into a first encoder, and outputting the i-th modal feature of the corresponding frame in the set perspective by the first encoder; or, In response to the i-th modal image being a multi-planar image, interpolate the features of at least one frame of the i-th modal image in the two-dimensional space corresponding to the set viewing angle to obtain the interpolated features of the corresponding frame at the set viewing angle; Input the interpolated features of at least one frame at the set viewing angle into the second encoder, and output the i-th modal feature of the corresponding frame at the set viewing angle by the second encoder.
11. The method according to claim 8, wherein The fusing process of the multi-modal features of any frame at the set viewing angle to obtain the fused features of any frame at the set viewing angle includes: Input the multi-modal features of any frame at the set viewing angle into the convolutional layer, and output the fused features of any frame at the set viewing angle by the convolutional layer.
12. The method according to claim 8, wherein Generating the target video at the set viewing angle based on the fused features of each frame at the set viewing angle includes: Perform an optimization process on the fused features of at least one frame at the set viewing angle to obtain the optimized features of the corresponding frame at the set viewing angle; Generate the target video at the set viewing angle based on the optimized features of each frame at the set viewing angle.
13. The method according to claim 12, wherein The optimizing process of the fused features of at least one frame at the set viewing angle to obtain the optimized features of the corresponding frame at the set viewing angle includes: Input the fused features of at least one frame at the set viewing angle into the adapter, where the adapter includes a spatial convolutional layer, a temporal convolutional layer, and a temporal self-attention layer; Extract spatial features of the fused features of at least one frame at the set viewing angle through the spatial convolutional layer to obtain the spatial features of the corresponding frame at the set viewing angle; Extract spatio-temporal features of the spatial features of at least one frame at the set viewing angle through the temporal convolutional layer to obtain the spatio-temporal features of the corresponding frame at the set viewing angle; Enhance the attention of the spatio-temporal features of at least one frame at the set viewing angle through the temporal self-attention layer to obtain the optimized features of the corresponding frame at the set viewing angle.
14. The method according to claim 12, characterized in that Generating the target video at the set viewing angle based on the optimized features of each frame at the set viewing angle includes: Input the pure noise image and the optimized features of each frame at the set viewing angle into the diffusion model, and denoise the pure noise image based on the optimized features of at least one frame at the set viewing angle through the diffusion model to obtain the predicted features of the corresponding frame at the set viewing angle; Input the predicted features of each frame at the set viewing angle into the decoder, and output the target video at the set viewing angle by the decoder.
15. The method according to any one of claims 1 to 14, characterized in that, Generating the target video at the set viewing angle based on the features of voxels of each frame in the three-dimensional space includes: Obtain the description text of the generation requirements of the target video at the set viewing angle; Generate the target video at the set viewing angle based on the features of voxels of each frame in the three-dimensional space and the description text.
16. The method according to any one of claims 1-14, characterized in that, Generating the target video at the set viewing angle based on the features of voxels of each frame in the three-dimensional space includes: Adjust the features of at least one frame of voxels in the three-dimensional space based on the generation requirements of the target video at the set viewing angle; Generate the target video at the set viewing angle based on the final features of the voxels of each frame in the three-dimensional space.
17. The method according to any one of claims 1 to 14, characterized in that, Obtain the original video, including: Obtain real driving videos at multiple set viewing angles, and based on the real driving videos at the multiple set viewing angles, obtain the original video at the observation viewing angle; After generating the target video at the set viewing angle, further include: Use the target video at the set viewing angle as the simulated driving video at the set viewing angle; Based on the simulated driving video at the set viewing angle, train and / or test the autonomous driving model.
18. The method according to any one of claims 1 to 14, characterized in that, Obtain the original video, including: Obtain original extended reality videos at multiple set viewing angles, and based on the original extended reality videos at the multiple set viewing angles, obtain the original video at the observation viewing angle; After generating the target video at the set viewing angle, further include: Use the target video at the set viewing angle as the new extended reality video at the set viewing angle, and the new extended reality video at the set viewing angle is used for visual display by an extended reality device.
19. The method according to any one of claims 1 to 14, characterized in that, Obtain the original video, including: Obtain real robot navigation videos at multiple set viewing angles, and based on the real robot navigation videos at the multiple set viewing angles, obtain the original video at the observation viewing angle; After generating the target video at the set viewing angle, further include: Use the target video at the set viewing angle as the simulated robot navigation video at the set viewing angle; Based on the simulated robot navigation video at the set viewing angle, train and / or test the robot navigation model.
20. The method according to any one of claims 1 to 14, characterized in that Obtain the original video, including: Obtain real scene videos at multiple set viewing angles, and based on the real scene videos at the multiple set viewing angles, obtain the original video at the observation viewing angle; After generating the target video at the set viewing angle, further include: Use the target video at the set viewing angle as the virtual scene video at the set viewing angle; Based on the virtual scene video at the set viewing angle, generate a film and television video, and the film and television video is used for visual display by a film and television video playback device.
21. A training method for a video generation model, characterized in that, The video generation model includes a first generation model and a second generation model, and the video generation model is used to generate the target video at the set viewing angle; The method includes: Based on the first sample videos at multiple set viewing angles, generate the second sample video at the observation viewing angle, and obtain at least one frame in the second sample video; Based on the sample features of the images of at least one frame in the second sample video in the two-dimensional space, determine the sample features of the voxels of the corresponding frame in the three-dimensional space; Based on the second sample video and the sample features of the voxels of each frame in the three-dimensional space, train the first generation model; and / or, Based on at least one of the first sample videos at the set viewing angle, and the sample features of the voxels of each frame in the three-dimensional space, train the second generation model.
22. A video generation device, characterized in that, Include: An acquisition module, configured to acquire at least one frame in the original video at the observation viewing angle; A determination module, configured to determine the features of the voxels of the corresponding frame in the three-dimensional space based on the features of the images of at least one frame in the original video in the two-dimensional space; A generation module, configured to generate a target video at a set viewing angle based on the features of voxels of each frame in the three-dimensional space.
23. A training device for a video generation model, characterized in that, The video generation model includes a first generation model and a second generation model. The video generation model is used to generate a target video at a set viewing angle. The apparatus includes: An acquisition module, configured to generate a second sample video at an observation viewing angle based on first sample videos at multiple set viewing angles, and acquire at least one frame in the second sample video; A determination module, configured to determine the sample features of voxels of the corresponding frame in the three-dimensional space based on the sample features of the images of at least one frame in the second sample video in the two-dimensional space; A first training module, configured to train the first generation model based on the second sample video and the sample features of voxels of each frame in the three-dimensional space; and / or, A second training module, configured to train the second generation model based on the first sample videos at at least one of the set viewing angles and the sample features of voxels of each frame in the three-dimensional space.
24. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 20, and / or, implements the steps of the method according to claim 21.
25. A vehicle, characterized in that, It includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to: Implement the steps of the method according to any one of claims 1-20, and / or, implement the steps of the method according to claim 21.
26. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instruction is executed by the processor, it implements the steps of the method according to any one of claims 1-20, and / or, implements the steps of the method according to claim 21.
27. A chip, characterized in that, The chip includes an interface circuit and a processing circuit coupled to each other. The interface circuit is used to input or output signals. The processing circuit is configured to implement the steps of the method according to any one of claims 1 to 20, and / or, implement the steps of the method according to claim 21.
28. A computer program product, characterized in that, It includes a computer program. When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 20, and / or, implements the steps of the method according to claim 21.
Citation Information
Cited By
Training method and device of video generation model
CN121056659A