A spatial video image generation method, apparatus, device, medium and product
By acquiring 360-degree panoramic video images with 3D depth information, constructing a physical world model, and generating a three-dimensional spatial scene with six degrees of freedom, the problem of the lack of spatial movement freedom in existing panoramic videos is solved, and a 360-degree panoramic video experience with a realistic sense of scene is achieved.
Patent Information
- Application Number
- CN202610807362.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-25
AI Technical Summary
Existing 360-degree panoramic videos only support three degrees of freedom: yaw, pitch, and roll. They lack the spatial movement capabilities of three degrees of freedom: undulation, forward and backward movement, and lateral movement, resulting in an inadequate user experience.
By acquiring 360-degree panoramic video images with 3D depth information, extracting three-dimensional spatial structure information, constructing a physical world model, and generating a three-dimensional spatial scene that supports six degrees of freedom: yaw, pitch, roll, undulation, forward and backward movement, and lateral movement, it can be viewed in 3D using VR or MR headsets.
It achieves realistic spatial video with six degrees of freedom and 360-degree panoramic video images, enhancing the user's spatial movement experience and avoiding the feeling of AI generation and video rendering.
Smart Images

Figure CN122637293A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video image generation, and in particular to a method, apparatus, device, medium, and product for generating spatial video images. Background Technology
[0002] Currently available 360-degree panoramic cameras (such as those from InStone, DJI, and Kankan) consist of two panoramic cameras, one front and one rear. Each camera has a field of view of approximately 200-210 degrees, allowing for sufficient overlap to stitch the images together and output 360-degree panoramic video. While this type of panoramic video allows for viewing from different angles, it is still essentially 2D. This 360-degree panoramic video supports 360-degree panoramic viewing at any given moment, but only in three degrees of freedom: yaw, pitch, and roll. It does not support three degrees of freedom: undulation, forward / backward movement, and lateral shift. Summary of the Invention
[0003] The purpose of this application is to provide a method, apparatus, device, medium, and product for generating spatial video images, which can generate spatial videos and 360-degree panoramic video images that support six degrees of freedom: yaw, pitch, roll, undulation, forward and backward movement, and lateral movement, creating a realistic sense of scene.
[0004] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a method for generating spatial video images, including: Acquire 360-degree panoramic video images with 3D depth information; Extract three-dimensional spatial structure information from the 360-degree panoramic video images; A physical world model is constructed based on the 360-degree panoramic video images and the three-dimensional spatial structure information; Generate a three-dimensional spatial scene based on the physical world model; Based on the three-dimensional spatial scene, the 360-degree panoramic video image is moved in six degrees of freedom: yaw, pitch, roll, undulation, forward and backward movement, and lateral movement, and a new 360-degree panoramic video image is generated according to any coordinate.
[0005] In one embodiment, the spatial video image generation method further includes: The 360-degree panoramic video images in the three-dimensional spatial scene are subjected to binocular split-viewing and viewed using a 3D viewing device.
[0006] In one embodiment, constructing a physical world model based on the 360-degree panoramic video image and the three-dimensional spatial structure information specifically includes: Semantic segmentation is performed on the 360-degree panoramic video image, and the three-dimensional spatial structure information is hierarchically divided and depth-aligned according to the segmentation results to obtain an RGB-D video sequence; The RGB-D video sequence is reconstructed in three dimensions, and a navigation mesh and collision bounding box are built to obtain a three-dimensional interactive geometric scene; By configuring physical parameters and setting dynamic constraints for the three-dimensional interactive geometric scene, a three-dimensional simulation scene is obtained. The geometric structure of the three-dimensional simulation scene is corrected and the physical rules are unified to obtain a physical world model.
[0007] In one embodiment, generating a three-dimensional spatial scene based on the physical world model specifically includes: The physical world model is analyzed to obtain model data; the model data includes: three-dimensional spatial structure, semantic information, and physical rules. The extracted model data is then subjected to lightweighting, structuring, and hierarchical optimization to obtain the processed model data. Generative scene construction is performed based on the processed model data and the physical constraints in the physical world model to obtain the generative scene structure; To obtain the matched scene structure, a unified semantic label and physical dynamic parameters are matched to the generative scene structure. Based on feature alignment and spatial fusion algorithms, the matched scene structure is spliced with the physical world model to obtain the overall scene; The overall scene is rendered, optimized, and calibrated to generate a new three-dimensional spatial scene.
[0008] In one embodiment, acquiring a 360-degree panoramic video image with 3D depth information specifically includes: Acquire 360-degree panoramic video images with 3D depth information captured by a 360-degree panoramic camera; the 360-degree panoramic camera includes at least two front-facing panoramic cameras and at least two rear-facing panoramic cameras; stitch together the front-facing panoramic images captured by the front-facing panoramic cameras and the corresponding rear-facing panoramic images captured by the rear-facing panoramic cameras to obtain a stitched image; add 3D depth information to the stitched image to obtain a 360-degree panoramic video image with 3D depth information; the 3D depth information includes: front-facing 3D depth information calculated from at least two front-facing panoramic images and rear-facing 3D depth information calculated from at least two rear-facing panoramic images.
[0009] In one embodiment, the 3D viewing device is a VR headset or an MR headset.
[0010] Secondly, this application provides a spatial video image generation apparatus, comprising: A panoramic video image reading unit is used to acquire 360-degree panoramic video images with 3D depth information; A three-dimensional information extraction unit is used to extract three-dimensional spatial structure information based on the 360-degree panoramic video image; The physical world modeling unit is used to construct a physical world model based on the 360-degree panoramic video image and the three-dimensional spatial structure information. A 3D spatial scene generation unit is used to generate a 3D spatial scene based on the physical world model. The six-degree-of-freedom spatial display unit is used to move the 360-degree panoramic video image in six degrees of freedom—yaw, pitch, roll, undulation, forward and backward, and lateral—according to the three-dimensional spatial scene, and to generate a new 360-degree panoramic video image according to any coordinate.
[0011] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the computer program to implement the spatial video image generation method described in any one of the above.
[0012] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the spatial video image generation method described in any one of the above descriptions.
[0013] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the spatial video image generation method described above.
[0014] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a method, apparatus, device, medium, and product for generating spatial video images. Based on a 360-degree panoramic video image with 3D depth information, it extracts three-dimensional spatial structure information and constructs a physical world model. A three-dimensional spatial scene is generated according to the physical world model, enabling movement with six degrees of freedom: yaw, pitch, roll, undulation, forward / backward, and lateral. A new 360-degree panoramic video image is generated according to any coordinate system. This application can generate spatial videos and 360-degree panoramic video images that support six degrees of freedom: yaw, pitch, roll, undulation, forward / backward, and lateral, creating a realistic scene feel. The new 360-degree panoramic video image has the same realistic scene feel as the original 360-degree panoramic video image, without obvious AI-generated or rendered visual effects, facilitating subsequent use by the user. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A schematic flowchart illustrating a spatial video image generation method provided in an embodiment of this application; Figure 2 A schematic diagram of the functional modules of a spatial video image generation device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] While current panoramic videos allow for free viewing from different angles, they are still essentially 2D, lacking 3D depth information and failing to support the three degrees of freedom: undulation, forward / backward movement, and lateral shift. Even if a user moves a handheld panoramic camera in space, they still lack the two degrees of freedom in the vertical plane of the direction of movement.
[0020] This application proposes a spatial video image generation method, apparatus, device, medium, and product based on a 3D panoramic camera. This spatial video image generation method can generate spatial video based on 360-degree panoramic video images with 3D depth information captured by a 3D panoramic camera, thereby supplementing the two missing degrees of freedom in the vertical plane of the motion direction, facilitating subsequent use by the user.
[0021] In one exemplary embodiment, such as Figure 1 As shown, a method for generating spatial video images is provided, including: Step 101: Obtain a 360-degree panoramic video image with 3D depth information.
[0022] Step 102: Extract three-dimensional spatial structure information based on the 360-degree panoramic video image.
[0023] Step 103: Construct a physical world model based on the 360-degree panoramic video image and the three-dimensional spatial structure information.
[0024] The physical world model is used to characterize the three-dimensional physical world.
[0025] Step 104: Generate a three-dimensional spatial scene based on the physical world model.
[0026] Step 105: Move the 360-degree panoramic video image according to the three-dimensional spatial scene using six degrees of freedom: yaw, pitch, roll, undulation, forward and backward, and lateral movement, and generate a new 360-degree panoramic video image according to any coordinate.
[0027] The new 360-degree panoramic video image has the same sense of realism as the original 360-degree panoramic video image obtained in step 101, without any obvious AI-generated or video image rendering feel.
[0028] In another exemplary embodiment of this application, such as Figure 1 As shown, the spatial video image generation method further includes: step 106, performing binocular split-viewing on the 360-degree panoramic video image in the three-dimensional spatial scene, and viewing it using a 3D viewing device.
[0029] The 3D viewing device can be a VR headset or a MR headset.
[0030] In another exemplary embodiment of this application, step 101 specifically includes: acquiring a 360-degree panoramic video image with 3D depth information captured by a 360-degree panoramic camera; the 360-degree panoramic camera includes: at least two front panoramic cameras and at least two rear panoramic cameras; stitching together the front panoramic image captured by the front panoramic camera and the corresponding rear panoramic image captured by the rear panoramic camera to obtain a stitched image; adding 3D depth information to the stitched image to obtain a 360-degree panoramic video image with 3D depth information; the 3D depth information includes: front 3D depth information calculated from at least two front panoramic images and rear 3D depth information calculated from at least two rear panoramic images.
[0031] In another exemplary embodiment of this application, step 103 specifically includes: (1) Semantic segmentation is performed on the 360-degree panoramic video image, and the three-dimensional spatial structure information is divided into layers and depth aligned according to the segmentation results to obtain an RGB-D video sequence.
[0032] (2) Perform three-dimensional reconstruction on the RGB-D video sequence, and build a navigation mesh and collision bounding box to obtain a three-dimensional interactive geometric scene.
[0033] (3) Configure physical parameters and set dynamic constraints for the three-dimensional interactive geometric scene to obtain a three-dimensional simulation scene.
[0034] (4) The geometric structure of the three-dimensional simulation scene is corrected and the physical rules are unified to obtain the physical world model.
[0035] In another exemplary embodiment of this application, step 104 specifically includes: (1) The physical world model is analyzed to obtain model data; the model data includes: three-dimensional spatial structure, semantic information and physical rules.
[0036] (2) The extracted model data is lightweighted, structured and hierarchically optimized to obtain the processed model data.
[0037] (3) Generative scene construction is carried out based on the processed model data and the physical constraints in the physical world model to obtain the generative scene structure.
[0038] (4) Match the generative scene structure with unified semantic tags and physical dynamic parameters to obtain the matched scene structure.
[0039] (5) Based on the feature alignment and spatial fusion algorithm, the matched scene structure is spliced with the physical world model to obtain the overall scene.
[0040] (6) Render the overall scene, optimize and calibrate it, and generate a new three-dimensional space scene.
[0041] In another exemplary embodiment of this application, the 360-degree panoramic camera includes: at least two front panoramic cameras, at least two rear panoramic cameras, an image processing module, a panoramic stitching module, and a 3D information calculation module.
[0042] The front-facing panoramic camera is used to capture a front-facing panoramic image. The rear-facing panoramic camera is used to capture a rear-facing panoramic image. The image processing module is used to process the front-facing panoramic image and the rear-facing panoramic image to obtain a front-facing processed panoramic image and a rear-facing processed panoramic image. The panoramic stitching module is used to stitch together any front-facing processed panoramic image with the rear-facing processed panoramic image from the rear-facing panoramic camera corresponding to the position of the front-facing panoramic camera to obtain a 360-degree panoramic image. The 3D information calculation module is used to perform 3D information calculation on the front-facing panoramic images from at least two front-facing panoramic cameras to obtain front-facing 3D depth information, and to perform 3D information calculation on the rear-facing panoramic images from at least two rear-facing panoramic cameras to obtain rear-facing 3D depth information. Based on the front-facing 3D depth information, the rear-facing 3D depth information, and the 360-degree panoramic image, a 360-degree panoramic video image with 3D depth information is obtained.
[0043] The implementation process of this spatial video image generation method is as follows: First, read a 360-degree panoramic video image with 3D depth information and extract the three-dimensional spatial structure information; Second, construct a representation of the three-dimensional physical world based on the above three-dimensional spatial structure information to obtain a physical world model, so as to understand the position and size of objects in space, the relative relationships between objects, and physical laws (gravity, occlusion, light reflection, etc.); Third, generate a new three-dimensional spatial scene according to the physical world model, and the physical laws are self-consistent; Fourth, users can freely browse the generated three-dimensional spatial scene, supporting six degrees of freedom and realizing arbitrary spatial display; Fifth, perform binocular split viewing on the generated 360-degree panoramic video image to produce a 3D image effect, which can be viewed in VR, MR and other devices.
[0044] Based on the same inventive concept, this application also provides a spatial video image generation apparatus for implementing the spatial video image generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more spatial video image generation apparatus embodiments provided below can be found in the limitations of the spatial video image generation method described above, and will not be repeated here.
[0045] In one exemplary embodiment, such as Figure 2 As shown, a spatial video image generation device is provided, comprising: The panoramic video image reading unit 201 is used to acquire 360-degree panoramic video images with 3D depth information.
[0046] The three-dimensional information extraction unit 202 is used to extract three-dimensional spatial structure information based on the 360-degree panoramic video image.
[0047] The physical world modeling unit 203 is used to construct a physical world model based on the 360-degree panoramic video image and the three-dimensional spatial structure information.
[0048] The three-dimensional space scene generation unit 204 is used to generate a three-dimensional space scene based on the physical world model.
[0049] The six-degree-of-freedom spatial display unit 205 is used to move the 360-degree panoramic video image in six degrees of freedom (yaw, pitch, roll, undulation, forward and backward, and lateral) according to the three-dimensional spatial scene, and generate a new 360-degree panoramic video image according to any coordinate.
[0050] As an optional implementation, the spatial video image generation apparatus further includes: The binocular split-view image generation unit 206 is used to perform binocular split-view on the 360-degree panoramic video image in the three-dimensional spatial scene and to view it using a 3D viewing device.
[0051] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 3 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores 360-degree panoramic video images of a three-dimensional spatial scene. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a spatial video image generation method.
[0052] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment to which the present application is applied. Specific computer equipment may include, for example, [the following is a list of possible additional structures]. Figure 3The embodiments show more or fewer components, combinations of certain components, or different component arrangements. In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, which the processor executes to implement the steps in the above-described method embodiments.
[0053] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0054] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0055] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0056] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0057] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, etc., and are not limited to these.
[0058] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0059] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for generating spatial video images, characterized in that, include: Acquire 360-degree panoramic video images with 3D depth information; Extract three-dimensional spatial structure information from the 360-degree panoramic video images; A physical world model is constructed based on the 360-degree panoramic video images and the three-dimensional spatial structure information; Generate a three-dimensional spatial scene based on the physical world model; Based on the three-dimensional spatial scene, the 360-degree panoramic video image is moved in six degrees of freedom: yaw, pitch, roll, undulation, forward and backward, and lateral movement, and a new 360-degree panoramic video image is generated according to any coordinate.
2. The spatial video image generation method according to claim 1, characterized in that, Also includes: The 360-degree panoramic video images in the three-dimensional spatial scene are subjected to binocular split-viewing and viewed using a 3D viewing device.
3. The spatial video image generation method according to claim 1, characterized in that, Constructing a physical world model based on the 360-degree panoramic video images and the three-dimensional spatial structure information specifically includes: Semantic segmentation is performed on the 360-degree panoramic video image, and the three-dimensional spatial structure information is hierarchically divided and depth-aligned according to the segmentation results to obtain an RGB-D video sequence; The RGB-D video sequence is reconstructed in three dimensions, and a navigation mesh and collision bounding box are built to obtain a three-dimensional interactive geometric scene; By configuring physical parameters and setting dynamic constraints for the three-dimensional interactive geometric scene, a three-dimensional simulation scene is obtained. The geometric structure of the three-dimensional simulation scene is corrected and the physical rules are unified to obtain a physical world model.
4. The spatial video image generation method according to claim 1, characterized in that, Generating a three-dimensional spatial scene based on the physical world model specifically includes: The physical world model is analyzed to obtain model data; the model data includes: three-dimensional spatial structure, semantic information, and physical rules. The extracted model data is then subjected to lightweighting, structuring, and hierarchical optimization to obtain the processed model data. Generative scene construction is performed based on the processed model data and the physical constraints in the physical world model to obtain the generative scene structure; To obtain the matched scene structure, a unified semantic label and physical dynamic parameters are matched to the generative scene structure. Based on feature alignment and spatial fusion algorithms, the matched scene structure is spliced with the physical world model to obtain the overall scene; The overall scene is rendered, optimized, and calibrated to generate a new three-dimensional spatial scene.
5. The spatial video image generation method according to claim 1, characterized in that, Acquiring 360-degree panoramic video images with 3D depth information, specifically including: Acquire 360-degree panoramic video images with 3D depth information captured by a 360-degree panoramic camera; the 360-degree panoramic camera includes at least two front-facing panoramic cameras and at least two rear-facing panoramic cameras; stitch together the front-facing panoramic images captured by the front-facing panoramic cameras and the corresponding rear-facing panoramic images captured by the rear-facing panoramic cameras to obtain a stitched image; add 3D depth information to the stitched image to obtain a 360-degree panoramic video image with 3D depth information; the 3D depth information includes: front-facing 3D depth information calculated from at least two front-facing panoramic images and rear-facing 3D depth information calculated from at least two rear-facing panoramic images.
6. The spatial video image generation method according to claim 2, characterized in that, The 3D viewing device is a VR headset or an MR headset.
7. A spatial video image generation device, characterized in that, include: A panoramic video image reading unit is used to acquire 360-degree panoramic video images with 3D depth information; A three-dimensional information extraction unit is used to extract three-dimensional spatial structure information based on the 360-degree panoramic video image; The physical world modeling unit is used to construct a physical world model based on the 360-degree panoramic video image and the three-dimensional spatial structure information. A 3D spatial scene generation unit is used to generate a 3D spatial scene based on the physical world model. The six-degree-of-freedom spatial display unit is used to move the 360-degree panoramic video image in six degrees of freedom—yaw, pitch, roll, undulation, forward and backward, and lateral—according to the three-dimensional spatial scene, and to generate a new 360-degree panoramic video image according to any coordinate.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the spatial video image generation method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the spatial video image generation method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the spatial video image generation method according to any one of claims 1-6.