3D video generation method, 3D video viewing method, and electronic device

By employing a novel perspective synthesis algorithm based on deep learning and artificial intelligence, combined with data acquisition guided by depth sensors and inertial measurement units, high-quality 3D videos are generated. This solves the problems of limited accuracy and viewpoint in existing 3D video generation technologies, providing an immersive 6-DOF viewing experience.

WO2025208971A9PCT designated stage Publication Date: 2026-05-15HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-12-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively generate high-quality 3D videos, especially those based on monocular 2D video or small-baseline binocular video, which suffer from accuracy issues and limited viewpoints. Furthermore, ordinary users find it difficult to obtain depth information and 3D models.

Method used

A novel perspective synthesis algorithm based on deep learning and artificial intelligence is adopted. By analyzing the 3D video scene type, the corresponding novel perspective generation model is used to generate 3D video data. The data acquisition is guided by depth sensors and inertial measurement units to ensure the integrity of the acquisition. Finally, neural rendering technology is used to generate 3D videos that meet the 6 degrees of freedom requirement.

Benefits of technology

It enables the generation of high-quality, immersive 3D video experiences on ordinary devices. Users can freely change their viewing angle, providing rich 3D perception and stereoscopic effect, avoiding limited viewing angle and noise, and reducing device costs and data acquisition difficulties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024143731_15052026_PF_FP_ABST
    Figure CN2024143731_15052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a 3D video generation method and an electronic device. The method is applied to the electronic device. The method comprises: on the basis of captured data, acquiring 3D video material data, wherein the captured data comprises a 2D video and / or a 2D image; on the basis of the 3D video material data, analyzing a 3D video scene type; on the basis of the 3D video scene type, acquiring a new-view generation model corresponding to the 3D video scene type; and using the new-view generation model corresponding to the 3D video scene type to generate 3D video data on the basis of the 3D video material data. On the basis of the method of the embodiments of the present application, at a 3D video data generation stage, a corresponding new-view generation model is called on the basis of a 3D video scene type, and thus a more complete scene synthesis view can be acquired on the basis of a new-view synthesis algorithm that uses deep learning and artificial intelligence, such that when 3D video data is played, a user can immersively view 3D video content from different angles, thereby providing a stronger sense of 3D and richer 3D video content.
Need to check novelty before this filing date? Find Prior Art

Description

A method for generating 3D video, a method for viewing 3D video, and an electronic device.

[0001] This application claims priority to Chinese Patent Application No. 2024103860747, filed on March 31, 2024, entitled "A method for generating, viewing and an electronic device for 3D video", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, and in particular to a method for generating 3D video, a method for viewing 3D video, and an electronic device. Background Technology

[0003] With the rapid development of multimedia-related technologies and hardware, various 3D devices have been launched on the market. Among them, virtual reality (VR) glasses have gradually been commercialized and are increasingly accepted by ordinary consumers.

[0004] VR glasses can play 3D videos, providing users with a stereoscopic and immersive viewing experience. To obtain 3D videos, a method for generating 3D videos is needed. Summary of the Invention

[0005] Regarding the question of how to generate 3D videos, this application provides a method for generating 3D videos, a method for viewing 3D videos, and an electronic device. This application also provides a computer-readable storage medium.

[0006] The embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, this application provides a 3D video generation method, which is applied to an electronic device, and the method includes:

[0008] Based on the captured data, 3D video material data is obtained, including 2D video and / or 2D images;

[0009] Analyze 3D video scene types based on 3D video material data;

[0010] Based on the type of 3D video scene, obtain the corresponding new perspective generation model for the 3D video scene type;

[0011] Using the new perspective generation model corresponding to the 3D video scene type, 3D video data is generated based on the 3D video material data.

[0012] According to the first aspect of the method, in the 3D video data generation stage, the corresponding new perspective generation model is called according to the 3D video scene type to generate 3D video data. Thus, the new perspective synthesis algorithm based on deep learning and artificial intelligence can obtain a more complete scene synthesis perspective, so that when playing 3D video data, users can immerse themselves in watching different angles of 3D video content, providing a stronger 3D feeling and richer 3D video content.

[0013] In one implementation of the first aspect, the captured data includes one or more 2D videos and / or 2D images containing a target scene and / or a target object; the acquisition of 3D video material data based on the captured data includes:

[0014] Generate 3D video material data for the target scene and / or the target object based on one or more 2D videos and / or 2D images containing the target scene and / or the target object.

[0015] According to the above implementation method, the source of 3D video material data can be 2D video and / or 2D images captured using any image acquisition device (e.g., VR glasses or mobile devices), and the 2D video and / or 2D images can come from the same scene material captured by different users at different times.

[0016] The method described above reduces the difficulty of acquiring 3D video material data.

[0017] In one implementation of the first aspect, generating 3D video material data for the target scene and / or the target object based on one or more 2D videos and / or 2D images containing the target scene and / or the target object includes:

[0018] Identify video frames and / or images in one or more 2D videos and / or 2D images that are related to the target scene and / or the target object;

[0019] The 3D video material data is generated based on the video frames and / or images related to the target scene and / or the target object.

[0020] In one implementation of the first aspect, the method further includes:

[0021] Collecting the captured data includes guiding the collection of the captured data based on data collection integrity.

[0022] According to the above implementation method, in the data acquisition stage, the target scene is captured by a device including but not limited to a camera module, an inertial measurement unit, a depth sensor, etc. The data already collected by the sensor is analyzed in real time to determine the completeness of the target scene acquisition and guide the user to capture the area where the acquisition has not reached the completeness threshold, thereby obtaining complete capture data.

[0023] The method described above ensures the integrity of the generated 3D video, allowing users to immerse themselves in watching the 3D video content from different angles, providing a stronger 3D feel and richer 3D video content, thus improving the user's 3D video viewing experience.

[0024] In one implementation of the first aspect, guiding the acquisition of the captured data based on data acquisition integrity includes:

[0025] The acquisition of the captured data is guided by the completeness of the acquisition trajectory and / or acquisition point location of the captured data.

[0026] In one implementation of the first aspect, guiding the acquisition of the captured data based on the acquisition trajectory and / or acquisition point location of the captured data includes:

[0027] Obtain the ideal acquisition method that matches the target scene and / or target object, wherein the ideal acquisition method includes the ideal acquisition trajectory and / or the ideal acquisition point position;

[0028] By comparing the acquisition trajectory and / or acquisition point location of the captured data with the ideal acquisition trajectory and / or ideal acquisition point location, the user is guided to supplement the acquisition data at the missing acquisition trajectory and / or acquisition point location based on the comparison results.

[0029] In one implementation of the first aspect, guiding the acquisition of the captured data based on data acquisition integrity includes:

[0030] The acquisition of the captured data is guided by the completeness of the acquired data.

[0031] In one implementation of the first aspect, guiding the acquisition of the captured data based on the integrity of the acquired data includes:

[0032] Based on the captured data, calculate the captured location and determine the coverage of the captured data on the target scene and / or target object;

[0033] Based on the coverage of the target scene and / or target object by the already collected capture data, guide the user to supplement the collection of capture data for the corresponding missing coverage areas.

[0034] In one implementation of the first aspect, guiding the acquisition of the captured data based on the integrity of the acquired data includes:

[0035] Based on the captured data that has been collected, calculate the spatial geometric structure corresponding to the captured data that has been collected;

[0036] Based on the completeness of the spatial geometry corresponding to the captured data, the user is guided to supplement the captured data of the missing spatial geometry.

[0037] In one implementation of the first aspect, the method further includes:

[0038] Calculate the observable area of ​​the 3D video virtual space corresponding to the 3D video data.

[0039] Based on the above implementation method, during the 3D video playback stage, the user can be guided to watch 3D videos based on the observable area of ​​the 3D video virtual space, avoiding the user's gaze / viewpoint from moving to areas outside the 3D virtual video space, thereby improving the user's 3D video viewing experience.

[0040] In one implementation of the first aspect, the method further includes playing the 3D video data, wherein playing the 3D video data includes:

[0041] The playback environment is detected while the 3D video data is being played;

[0042] Match the 3D video virtual space corresponding to the 3D video data with the playback environment to obtain the matching result;

[0043] Based on the matching results, the playback of the 3D video data is guided.

[0044] In one implementation of the first aspect, the instruction to play the 3D video data includes:

[0045] Determine the spatial correspondence between the playback environment and the 3D video virtual space;

[0046] Based on the spatial location correspondence, an active area is set in the 3D video virtual space, and / or the location of the viewing point in the 3D video virtual space is determined.

[0047] Based on the method described above, the playback of 3D video data is guided by the matching results between the 3D video virtual space and the playback environment, and an active area is set within the 3D video virtual space. This effectively prevents users from colliding with actual objects in the playback environment when changing their viewing angle or position while watching 3D videos.

[0048] Based on the method described above, the matching results between the 3D video virtual space and the playback environment guide the playback of 3D video data, determining the viewing point's location within the 3D video virtual space. This allows for the integration of the 3D video virtual space with the playback environment, enhancing the user's 3D video viewing experience.

[0049] In a second aspect, this application provides an electronic device, which includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to perform the steps of the method described in the first aspect.

[0050] Thirdly, this application provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the method described in the first aspect. Attached Figure Description

[0051] Figure 1 shows a flowchart of a DIBR implementation according to an embodiment of this application;

[0052] Figure 2 shows a flowchart of neural rendering implementation according to an embodiment of this application;

[0053] Figure 3 shows a flowchart of a DIBR implementation according to an embodiment of this application;

[0054] Figure 4 shows a flowchart of a DIBR implementation according to an embodiment of this application;

[0055] Figure 5 shows a flowchart of an MBR implementation according to an embodiment of this application;

[0056] Figure 6 is a schematic diagram of a 3D video generation device according to an embodiment of this application;

[0057] Figure 7 is a flowchart of a 3D video generation method according to an embodiment of this application;

[0058] Figure 8 shows a 3D image left and right eye view according to an embodiment of this application;

[0059] Figure 9 shows a 3D image left and right eye view according to an embodiment of this application;

[0060] Figure 10 shows a flowchart of data acquisition according to an embodiment of this application;

[0061] Figure 11 is a schematic diagram of a recommended data collection method according to an embodiment of this application;

[0062] Figure 12 shows a flowchart of data acquisition according to an embodiment of this application;

[0063] Figure 13 is a schematic diagram of a scene point cloud structure according to an embodiment of this application;

[0064] Figure 14 is a schematic diagram of a scene point cloud structure according to an embodiment of this application;

[0065] Figure 15 shows a flowchart of 3D video material data generation according to an embodiment of this application;

[0066] Figure 16 is a schematic diagram of an electronic device structure according to an embodiment of this application. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0068] The terminology used in the implementation section of this application is for the purpose of explaining specific embodiments of this application only, and is not intended to limit this application.

[0069] 3D videos can provide users with a three-dimensional and immersive viewing experience. 3D videos are mainly supported by a rich variety of binocular 3D video resources. They utilize the parallax of images seen by the left and right eyes, which are then merged in the brain to obtain three-dimensional information, allowing users to perceive objects as having a sense of distance, like a real three-dimensional scene.

[0070] One technical solution for acquiring 3D video is active stereoscopic shooting, which uses a binocular camera to ensure consistency in lens, aperture, and chromaticity, and synchronization of the two signals. However, currently common mobile shooting devices, such as mobile phones and cameras, can only shoot monocular 2D video or small-baseline binocular video. Active stereoscopic shooting requires specialized shooting equipment, therefore, this stereoscopic shooting method is relatively expensive.

[0071] Another technical solution for acquiring 3D video is passive video generation, which converts monocular 2D video into 3D video. This is mainly achieved by estimating binocular video (two-channel video) using computer vision and computer graphics methods.

[0072] Generally, in the process of converting monocular 2D video into 3D video, the key technology lies in virtual viewpoint synthesis technology based on 2D video. Virtual viewpoint synthesis technology can be divided into depth image based rendering (DIBR), model based rendering (MBR), and neural rendering technology, depending on the implementation methods and auxiliary tools.

[0073] DIBR is a method for rendering virtual viewpoint images based on a reference viewpoint texture map and a corresponding depth map using a 3D mapping equation. Due to its high speed and low complexity, it is widely used in the field of new viewpoint compositing.

[0074] Figure 1 shows a flowchart of a DIBR implementation according to an embodiment of this application.

[0075] In one embodiment, the basic flowchart of DIBR is shown in Figure 1.

[0076] The DIBR process can be simply described as follows:

[0077] S110, Inversely project one or more reference viewpoint images (e.g., reference image 101, depth image 102) onto the world coordinate system;

[0078] S120, reprojected onto the virtual viewpoint plane;

[0079] S130, generate a virtual image of the target viewpoint (virtual viewpoint image) through image fusion;

[0080] S140, perform post-processing (e.g., S141, remove artifacts and overlaps; S142, fill holes);

[0081] S150, obtain the final image and output it.

[0082] The core technology of DIBR is 3D image warping (3D-warping), which maps the pixels of the reference view to the target view through 3D transformation equations.

[0083] The disadvantages of DIBR include accuracy issues and limited viewpoints.

[0084] The accuracy issue refers to the fact that the accuracy of the depth map directly affects the quality of the synthesized viewpoint. The core of DIBR is the utilization of depth information. The 3D information of the reference viewpoint is constructed through the depth information, and then the 3D information of the target viewpoint is obtained through mapping transformation. When there is noise in the depth map, the quality of the synthesized viewpoint image is poor.

[0085] The viewpoint limitation problem refers to the difficulty in effectively filling the holes generated in the target viewpoint when the virtual viewpoint of the target is too far away from the reference viewpoint, which greatly affects the quality of the synthesized image from the target viewpoint.

[0086] The MBR (Multi-Rendering) approach is based on rendering images from a specified target perspective using 3D geometric models. The MBR workflow involves: first, creating explicit 3D models of 3D objects or scenes, such as point clouds, voxels, and meshes; then adding textures, lighting, and shadows to the 3D models; and finally, producing realistic 2D images based on rendering algorithms (such as rasterization or ray tracing).

[0087] In the MBR scheme, although using explicit geometric models can yield high-quality composite images from relatively few input images, accurately estimating the scene geometry is difficult in challenging scenarios such as textureless areas, specular highlights, reflections, and repetitive textures. When the reconstructed 3D model is of poor quality, it is often impossible to fully recover the target viewpoint image from unreconstructed or overreconstructed areas. While specialized tools can be used to scan existing real-world objects for 3D modeling, this is impractical for most ordinary users. Furthermore, the rendering process requires additional input of scene lighting, materials, and other physical properties, which are difficult to obtain and estimate.

[0088] Neural rendering is a rapidly emerging field that allows for compact representations of scenes by using neural networks to learn information about the scene, such as spatial density and color, from existing data. The main idea behind neural rendering is to combine knowledge from physically based classical computer graphics and deep learning to generate realistic images from the target viewpoint.

[0089] Figure 2 shows a flowchart of a neural rendering implementation according to an embodiment of this application.

[0090] As shown in Figure 2, the input to the neural rendering method is a sequence of 2D images (input image 201), and other additional inputs can also be added (other inputs 202 (such as depth, optical flow, geometry, etc.)). The input information is uniformly transformed into spatial representation data 210. The spatial representation data 210 is input into the neural network 220.

[0091] Neural network 220 is supervised to represent the shape or appearance of a specific scene and is rendered using a preset rendering algorithm 230 (such as rasterization 231 or ray tracing 232) to generate a target viewpoint image 240 (a new viewpoint image). Neural network 220 utilizes learnable elements in the scene (such as the density and color of objects or scenes) to express scene information, thereby rendering an image from the target viewpoint. Currently, neural rendering technology has made significant progress in both the quality of new viewpoint synthesis and the real-time performance of rendering, making it possible to apply neural rendering technology to consumer-grade VR glasses. Neural rendering technology can achieve high-fidelity target viewpoint image synthesis at the static or dynamic, object-level or large-scene level.

[0092] For example, in one feasible technical solution, a reference viewpoint image and depth are obtained based on the DIBR method, and the reference viewpoint image is distorted to the target viewpoint according to the depth value.

[0093] Figure 3 shows a flowchart of a DIBR implementation according to an embodiment of this application.

[0094] As shown in Figure 3:

[0095] S310, acquire the reference viewpoint video and the reference viewpoint depth map video corresponding to the reference viewpoint video, decompose the reference viewpoint video into a frame sequence of reference viewpoint images, and decompose the reference viewpoint depth map video into a frame sequence of reference viewpoint depth maps.

[0096] S320: Map the reference viewpoint images of each frame to the virtual viewpoint to generate the original virtual viewpoint images of each frame;

[0097] S330, repair the reference viewpoint image and reference viewpoint depth map of each frame, and map the repaired reference viewpoint image of each frame to the virtual viewpoint to generate the virtual viewpoint auxiliary image of each frame.

[0098] S340, Repair the holes in the original virtual viewpoint image based on the virtual viewpoint auxiliary image of each frame to generate the final virtual viewpoint image;

[0099] S350, synthesizes the final images of each virtual viewpoint to generate a virtual viewpoint video;

[0100] S360 synthesizes virtual viewpoint video and reference viewpoint video to generate multi-viewpoint 3D video.

[0101] Based on the above scheme, each frame of the image can be repaired by extracting the reference image, expanding the image boundary, and repairing the abrupt change region. This can effectively solve the boundary holes and holes in the abrupt change region of the internal reference viewpoint depth map in the original image of the virtual viewpoint.

[0102] However, based on the above scheme, only target viewpoint images with small deviations from the reference viewpoint can be obtained, and the range of new viewpoints is limited to a certain extent. When the target viewpoint is far away from the reference viewpoint, the new viewpoint has large holes that are difficult to repair, resulting in poor image quality.

[0103] Furthermore, the above solution relies on depth and requires converting 2D camera images into 3D point clouds using depth values, thus necessitating the addition of a depth sensor or depth computing unit.

[0104] For example, in another feasible technical solution, based on the DIBR method, a camera device is used to acquire a reference viewpoint image without providing depth information, and the 2D video is converted into 3D video through a depth estimation method.

[0105] Figure 4 shows a flowchart of a DIBR implementation according to an embodiment of this application.

[0106] As shown in Figure 4:

[0107] In the depth information extraction stage, an existing 3D movie is used as the source dataset to train a U-shaped convolutional network, resulting in a high-performance network model that performs frame-by-frame depth estimation on the 2D video. A small neural network is then used to optimize the depth map by preserving edges and smoothing. In the viewpoint synthesis stage, a depth map-based viewpoint synthesis algorithm without camera parameters is proposed, employing a symmetrical rendering strategy from the center outwards to synthesize the left and right virtual viewpoints. Finally, in the image inpainting stage, a block-matching-based image inpainting algorithm incorporating temporal information is proposed to fill and repair cracks and holes in the left and right viewpoints. This method enables 2D-to-3D video conversion without any relevant parameter information from the original 2D video, effectively handling high-resolution images with good conversion results and high speed.

[0108] Based on the above approach, although it does not require an additional depth sensor and uses a pre-trained depth estimation network to estimate depth information, making data acquisition more convenient, the problem of limited target viewpoint still exists.

[0109] For example, in another feasible technical solution, based on the MBR method, a multi-view depth camera is used to acquire the corresponding 3D model of the target scene, such as the point cloud or polygon representation of the target scene (and objects in the scene). The model can reflect the 3D geometric structure of the scene (and objects in the scene). Based on the acquired interaction parameters, the 3D video model is used to draw the target view image.

[0110] Figure 5 shows a flowchart of an MBR implementation according to an embodiment of this application.

[0111] As shown in Figure 5:

[0112] S410, acquire depth video streams from at least three camera perspectives of the same scene;

[0113] S420, determine the target foreground point cloud and target background point cloud corresponding to the depth video streams from the at least three camera views;

[0114] S430, the target foreground point cloud and the target background point cloud are processed according to the target point cloud processing method corresponding to the depth video stream to obtain a 3D video model corresponding to the depth video stream.

[0115] Based on the above scheme, a 3D video model is obtained from the input of a multi-view depth camera. Based on the 3D explicit model, a new perspective image with a large angle can be obtained. However, this method requires the input image to come from at least a binocular depth camera, which does not conform to the daily settings of people shooting videos with mobile phones or ordinary camera devices.

[0116] To facilitate the acquisition of 3D videos, one embodiment of this application also provides a 3D video generation apparatus.

[0117] Figure 6 shows a schematic diagram of a 3D video generation apparatus according to an embodiment of this application.

[0118] As shown in Figure 6, the device includes a data acquisition unit 601, a computing unit 602, and a storage unit 603.

[0119] The acquisition unit 601 is used to acquire 3D video material data for generating 3D video.

[0120] In one embodiment, the acquisition unit 601 directly receives 3D video material data output by other electronic devices.

[0121] In another embodiment, the acquisition unit 601 receives captured data output by other electronic devices, processes the captured data, and generates 3D video material data, wherein the captured data includes 2D video and / or 2D images.

[0122] In another embodiment, the acquisition unit 601 includes a data acquisition component (e.g., a vision sensor, an inertial measurement device, a depth sensor, etc.). The acquisition unit 601 acquires and captures data, processes the captured data, and generates 3D video material data.

[0123] For example, in one embodiment, the acquisition unit 601 includes: a visual sensor (such as one or more ordinary cameras or depth cameras) for acquiring 2D video footage; an inertial measurement unit for acquiring the relative motion of the camera device or the relative motion of the VR glasses device; and a depth sensor, such as a ToF or laser device, for acquiring scene point cloud depth information.

[0124] The calculation unit 602 is used to generate 3D video data for 3D video playback based on the 3D video material data.

[0125] For example, in one embodiment, the computing unit 602 includes a CPU, GPU, cache, registers, etc., for running an operating system and processing various algorithm modules involved in generating 3D video data, such as real-time analysis of acquired data and new perspective generation algorithms.

[0126] Storage unit 603 includes memory and external storage, used for data acquisition, new perspective generation algorithm data, reading and writing temporary data, etc.

[0127] Optionally, in one embodiment, the device further includes a display unit 604. The display unit 604 is used to play 3D video data and present a 3D video playback effect. For example, the display unit 604 may be a VR glasses device.

[0128] In the description of the embodiments of this application, for the sake of convenience, the coloring device is described by function as various modules. The division of each module is only a logical functional division. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware.

[0129] Specifically, the apparatus proposed in this application can be fully or partially integrated onto a single physical entity (e.g., a GPU or other type of processor), or it can be physically separated. These modules can be implemented entirely in software via processing element calls; entirely in hardware; or some modules can be implemented in software via processing element calls, while others are implemented in hardware. For example, the detection module can be a separate processing element or integrated into a chip in an electronic device. The implementation of other modules is similar. Furthermore, these modules can be fully or partially integrated together or implemented independently. During implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0130] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs). Alternatively, these modules can be integrated together as a system-on-a-chip (SOC).

[0131] One embodiment of this application also provides a 3D video generation method. In the method of this embodiment, a 3D video that meets the 6 degrees of freedom (DoF) viewing requirements is generated from 2D video and / or 2D images.

[0132] Figure 7 shows a flowchart of a 3D video generation method according to an embodiment of this application.

[0133] S500 acquires 3D video material data based on captured data, including but not limited to 2D video and / or 2D images.

[0134] In addition to 2D video and / or 2D image data, the captured data may also include 3D point data from depth sensors and device motion data from inertial measurement units.

[0135] For example, in one embodiment, only 2D video and / or 2D images are used as 3D video material data.

[0136] For example, in one embodiment, 2D video and / or 2D images, along with depth information, are used as 3D video material data.

[0137] For example, in one embodiment, 2D video and / or 2D images, along with device motion information, are used as 3D video material data.

[0138] For example, in one embodiment, 2D video and / or 2D images, depth information, and device motion information are used as 3D video material data.

[0139] S510 analyzes 3D video scene types based on 3D video material data.

[0140] For example:

[0141] Scene size classification: Determine the scene size based on the actual spatial volume of the scene point cloud;

[0142] Scene shooting method classification: The shooting method of the scene is determined based on the data acquisition trajectory of the captured data (such as surround shooting, frontal shooting, etc.);

[0143] Scene motion / static attribute classification: motion / static detection is performed based on the image content in the captured data;

[0144] Object classification: Object detection is performed based on the image content in the captured data.

[0145] S511, based on the 3D video scene type, obtain the corresponding Novel View Synthesis (NVS) model for the 3D video scene type.

[0146] For example, in one implementation, the corresponding 3D video scene type NVS model is called from a pre-generated NVS model, based on the 3D video scene type.

[0147] For example, in another implementation, the model algorithm corresponding to the 3D video scene type is determined based on the 3D video scene type. Based on the determined model algorithm, an NVS model corresponding to the 3D video scene type is generated.

[0148] Specifically, in one embodiment, the novel perspective generation model for creating 3D video data includes various algorithms, such as:

[0149] For scenarios involving the shooting of dynamic long video sequences, including but not limited to using dynamic neural network models based on image-based rendering (IBR);

[0150] For scenarios involving dynamic, small-scene surround shooting, including but not limited to models that use tensor decomposition or voxel decomposition methods;

[0151] For scenarios involving handheld portrait selfies, models including but not limited to deformable field methods are available.

[0152] For scenes captured in static settings, including but not limited to models using rasterization rendering methods based on explicit geometric primitives.

[0153] The S520 uses a Novel View Synthesis (NVS) model corresponding to the 3D video scene type to generate 3D video data based on 3D video material data.

[0154] Different novel perspective generative models are used to generate 3D video data for different types of 3D video scenes. 3D video data exists in, but is not limited to, implicit model parameter forms, explicit geometric forms, or hybrid forms.

[0155] In one embodiment, S520 is followed by S530.

[0156] S530 enables interactive 3D video viewing based on 3D video data generated by S520.

[0157] Specifically, the real-time pose of the user's current viewing angle is calculated, and the new viewpoint generation model renders the corresponding scene of the viewing angle in real time based on the current observation pose.

[0158] In one embodiment, the current viewing pose is derived from the IMU motion information of the viewing device or a Simultaneous Localization and Mapping (SLAM) algorithm. For example, when a user wears VR glasses, the 3D video space is stationary relative to the real physical world. The user moves within the 3D video space, and based on the current position, a 3D video image is rendered at that position.

[0159] In another embodiment, the current observation viewpoint pose is determined by the user. For example, the user is wearing VR glasses and is stationary relative to the real physical world. The user selects different observation angles and positions in the 3D video space, and a corresponding 3D video image is generated based on the user's currently selected observation angle and position.

[0160] According to one embodiment of the method of this application, 6DoF 3D video is generated based on input captured data (2D video and / or 2D images). The image / video format of the input data is not limited, monocular or multi-view video is supported, and arbitrary new perspective images can be generated as the user's viewing angle changes.

[0161] According to one embodiment of this application, a method generates a 3D video that meets the 6DoF viewing requirement based on captured data (2D video and / or 2D images) collected by devices such as mobile phones, cameras, or VR glasses. When a user watches the 3D video while wearing VR glasses, a high-fidelity image at the corresponding target viewpoint is rendered in real time as the viewing viewpoint changes, as if the user is immersed in a 3D video virtual environment, providing the user with a sense of stereoscopic depth and immersion that surpasses ordinary 3D videos based on the principle of binocular parallax.

[0162] Furthermore, according to the method of this application embodiment, the generated 3D video data is not limited to applications for 3D video playback. Those skilled in the art can use the 3D video data according to actual needs (e.g., using 3D video data for 3D modeling) to realize a variety of different application scenarios (e.g., Augmented Reality (AR), Virtual Reality, Mixed Reality (MR)).

[0163] Specifically, in one embodiment, the new perspective generation model obtained in S511 is a neural network model.

[0164] According to an embodiment of this application, the method for generating new perspectives for neural network models has two advantages: first, the input is unrestricted, and video data can be collected using only ordinary shooting devices such as mobile phones and cameras, which is convenient and easy to operate for ordinary users; second, the target perspective is unrestricted, and target viewpoint images that are far from the reference perspective can still be generated even without explicit 3D model priors.

[0165] For example, when watching 3D videos with VR glasses, the user's 3D experience comes only from the depth perception caused by binocular parallax.

[0166] Figure 8 shows a 3D image left and right eye view according to an embodiment of this application.

[0167] As shown in Figure 8, there is a slight difference in the viewing angle between the left eye view 701 and the right eye view 702. The cube with a closer depth of field has a large parallax, while the triangle with a farther depth of field has a small parallax. When the left eye views the left view 701 and the right eye views the right view 702, the 2D images with different parallaxes of the left and right eyes are fused in the brain and appear to be at different distances, thus creating a 3D effect.

[0168] However, the content of the left and right viewpoints shown in Figure 8 is fixed. The user's 3D experience mainly comes from the principle of binocular parallax, which creates a 3D stereoscopic effect by perceiving the different distances of objects. Regardless of the user's actual viewing position in the physical environment, the 3D content they see remains unchanged.

[0169] According to one embodiment of this application, a neural network model can generate images with a wider range of new perspectives. According to one embodiment of this application, as the user's viewing angle changes, different angles of the 3D scene can be viewed, providing the user with a more immersive experience.

[0170] Figure 9 shows the left and right eye views of 3D images from two different viewing angles according to an embodiment of this application.

[0171] As shown in Figure 9, when a user views the cube from the right side, the left-eye view is shown as 801, and the right-eye view is shown as 802. There is a slight difference in the viewing angle between the left-eye view 801 and the right-eye view 802. The cube, with its closer depth of field, has a larger parallax, while the triangle, with its farther depth of field, has a smaller parallax. When the left eye views the left-eye view 801 and the right eye views the right-eye view 802, the 2D images with different parallaxes from the left and right eyes are fused in the brain, resulting in a sense of distance and thus creating a 3D effect.

[0172] According to one embodiment of the method of this application, when the user's viewing angle moves from the right side to the left side of the cube, the left-eye view is shown as 803, and the right-eye view is shown as 804. The left-eye view 903 and the right-eye view 904 have slightly different viewing angles. When the left eye views the left-eye view 903 and the right eye views the right-eye view 904, the 2D images with different parallaxes of the left and right eyes are fused in the brain, resulting in a sense of distance and thus creating a 3D effect. Simultaneously, as the viewing angle changes, the user sees different sides of the cube, providing richer 3D information and a more immersive 3D experience.

[0173] The current drawback of neural rendering-based methods for synthesizing new viewpoints is that when the input reference viewpoint does not fully cover the angle, the generated target viewpoint image contains significant noise, affecting the user's viewing experience. In dynamic scenes, high-quality viewpoint synthesis relies on data captured simultaneously by multiple cameras from different angles. However, in daily life, people mostly use mobile phones and camera devices to shoot monocular 2D videos, or use dual-camera phones and VR glasses to shoot binocular videos with a small baseline.

[0174] In response to the above situation, in one embodiment, based on 2D monocular video or small baseline binocular video, a 6DoF immersive 3D video viewing experience is provided to the user through methods such as guided acquisition and observable area recommendation.

[0175] Specifically, in one embodiment, during the acquisition of capture data, the acquisition of capture data is guided based on the completeness of the acquired capture data (e.g., 2D video and / or 2D images).

[0176] In one embodiment, the acquisition of data is guided by the completeness of the acquisition trajectory and / or acquisition point locations. Specifically, the acquisition trajectory and / or acquisition point locations during data acquisition are monitored to ensure that the user acquires data at all expected acquisition trajectories and / or acquisition point locations.

[0177] In one embodiment, the scene spatial geometry is restored based on the collected data. The user determines which parts of the scene need to continue collecting data or no longer need to be collected based on the completeness of the current scene space. For example, on a data acquisition device equipped with a depth sensor, the scene spatial structure is restored and displayed based on the collected depth information (including but not limited to geometric representations such as 3D point clouds and meshes). The user judges the completeness of the current scene spatial structure by observing it, continuing to collect data on areas where the spatial structure is poorly restored, and stopping collection on areas that are relatively complete, ultimately achieving complete collection of the target scene.

[0178] Figure 10 shows a flowchart of data acquisition according to an embodiment of this application.

[0179] The electronic device executes the following process shown in Figure 10.

[0180] S711: Determine the ideal acquisition trajectory and / or the ideal acquisition point location.

[0181] In one implementation, the target scene and / or target object type are identified, and based on the target scene and / or target object, an ideal acquisition trajectory and / or ideal acquisition point location are selected from preset acquisition trajectories and / or acquisition point locations.

[0182] In another implementation, a preset collection trajectory and / or collection point location is displayed to the user, and based on the user's selected operation, the ideal collection trajectory and / or ideal collection point location selected by the user is determined from the preset collection trajectory and / or collection point location.

[0183] For example, in one embodiment, the user actively selects (e.g., based on handle buttons, gesture interaction, eye tracking, etc.) the target scene and / or target object in the current shooting space, and the system virtually displays the ideal acquisition trajectory and / or acquisition point position in the space according to the selected target scene and / or target object.

[0184] For example, in another embodiment, the user can actively select a certain acquisition trajectory directly from the system's built-in acquisition trajectory library, and the system will simulate a virtual reality trajectory route in space based on the ideal acquisition trajectory actively selected by the user.

[0185] Figure 11 shows a schematic diagram of a recommended data collection method according to an embodiment of this application.

[0186] In one embodiment, the screen shown in FIG11 is displayed to indicate different recommended data collection methods to the user and to request the user to select a recommended data collection method.

[0187] As shown in Figure 11, based on common types of daily photography, the following data collection trajectory types are provided to users: surround object shooting (surround acquisition), front-facing object shooting (front-facing acquisition), and hemispherical shooting (hemispherical acquisition).

[0188] Image 900 illustrates the acquisition trajectory for a surround-type acquisition. Image 902 shows the subject being captured, and image 903 shows the acquisition trajectory. As shown in image 900, the acquisition trajectory 903 surrounds the subject 902. Image 901 is a top view of the surround-type acquisition. The arrows in image 901 indicate the direction the camera lens is pointing. As shown in image 901, the camera lens is pointing towards the subject 902.

[0189] Image 910 shows the acquisition trajectory for a frontal view. Image 912 shows the subject being captured, and image 913 shows the acquisition trajectory. As shown in image 910, the acquisition trajectory 913 covers the front of the subject 912. Image 911 is a top view of the frontal view acquisition. The arrows in image 911 indicate the direction the camera lens is facing. As shown in image 911, the camera lens is facing the front of the subject 912.

[0190] Image 920 illustrates the acquisition trajectory for hemispherical acquisition. Image 922 shows the subject being captured, and image 923 shows the acquisition trajectory. As shown in image 920, the acquisition trajectory 903 surrounds the subject 922, covering its perimeter and top, forming a hemisphere that encloses the subject 922. Image 921 is a top view of the hemispherical acquisition. The arrows in image 901 indicate the direction the camera lens is pointing. As shown in image 921, the camera lens position forms a hemisphere enclosing the subject 922, with the lens pointing towards the center of the hemisphere, the subject 902.

[0191] S712: Acquires capture data for a target scene and / or target object, including but not limited to 2D video and / or images and IMU information.

[0192] For example, if VR glasses are used to film a target scene, the VR glasses' multi-module cameras, inertial measurement unit, and depth sensor are utilized during the filming process to record camera motion and depth information while shooting 2D video. If a mobile phone is used to film the target scene, the phone's camera module and IMU are utilized. If the phone has a depth sensor, it is used simultaneously; otherwise, only 2D video and the relative motion information of the camera are recorded.

[0193] Register and fuse information from multiple sensors.

[0194] For example, the capture information from multiple camera 2D image frame sequences, IMU, LiDAR, or ToF devices can be synchronized to obtain the spatial pose and depth information corresponding to each RGB image.

[0195] The S720 guides the user to collect data based on the ideal acquisition trajectory and / or acquisition point location, according to the acquired trajectory and / or acquisition point location of the image acquisition device.

[0196] By comparing the recommended acquisition method with the acquisition trajectory and / or acquisition point locations of the image acquisition device, the missing acquisition trajectory and / or recommended acquisition point locations compared to the recommended acquisition method are determined. This guides the user to perform 2D video and / or 2D image acquisition based on the currently missing acquisition trajectory and / or recommended acquisition point locations.

[0197] Based on real-time feedback data from the IMU, the system calculates the user's current location, confirms the collected trajectory and / or location, and then calculates the uncollected trajectory and / or location, displaying the area that needs to be collected on the VR glasses' UI.

[0198] Specifically, in one embodiment, the user is guided to collect data via text, such as "Please move left or right" or "Please change the shooting angle".

[0199] In another embodiment, during the process of a user acquiring 2D video and / or 2D images, completed and pending acquisition trajectories and / or acquisition point locations are displayed via virtual markers.

[0200] For example, you can set up a virtual camera, such as placing a virtual camera icon in a virtual space, marking the completed virtual camera locations, and guiding the user to move to the marked incomplete virtual camera locations.

[0201] For example, a virtual collection area can be displayed, marking the areas where data collection has been completed and those where it has not, guiding users to collect data in the uncollected areas.

[0202] For example, a virtual data collection trajectory can be set up in a virtual space, marking the trajectories that have been walked and those that have not, guiding users to complete data collection by following the arrows on the virtual data collection trajectory.

[0203] In another embodiment, the collection of capture data is guided by the completeness of the collected data. Specifically, the collection results are verified after the capture data is collected to ensure that the user has collected enough data.

[0204] Figure 12 shows a flowchart of data acquisition according to an embodiment of this application.

[0205] The electronic device executes the following process shown in Figure 12.

[0206] S910: Acquire capture data for a target scene and / or target object. The acquired capture data includes, but is not limited to, 2D video and / or images and IMU information.

[0207] For details, please refer to the implementation of S712.

[0208] S920: Calculate the acquisition integrity based on the acquired data, including but not limited to judging by the coverage of the acquired location and the integrity of the spatial geometry.

[0209] The system calculates the collected location or scene geometry based on the collected data, and then determines the coverage of the collected data. If the collected data coverage is complete, the system prompts the user to stop collecting data. If the coverage does not meet the threshold requirements, the system prompts the user to continue collecting data.

[0210] In one embodiment, the coverage of the collected locations is analyzed. For example, it is determined whether the number of collected locations has reached a threshold.

[0211] In another embodiment, the scene spatial geometry (e.g., scene point cloud, mesh, etc.) is restored, and the integrity of the data acquisition is analyzed by determining the spatial geometry. Restoring the scene spatial structure requires the use of scene reconstruction techniques.

[0212] For example, in one embodiment, the spatial point cloud structure of the scene is recovered using a laser SLAM algorithm based on the collected IMU data and Lidar data.

[0213] For example, in one embodiment, the spatial point cloud structure of the scene is recovered using a visual SLAM algorithm based on the acquired image data and IMU data.

[0214] For example, in one embodiment, the spatial point cloud structure of the scene is recovered using a laser-visual SLAM algorithm based on the acquired image data, IMU data, and Lidar data.

[0215] For example, in one embodiment, three-dimensional modeling is performed based on the acquired 2D video and / or 2D images, spatial pose information, and depth information.

[0216] S930: Performs integrity checks on the calculation results of S920 to guide users in data collection.

[0217] Specifically, in one embodiment, in S930, the scene's coverage is confirmed based on the collected locations calculated in S920. For example, if the number of collected location points is small, the user is instructed to continue collecting until the number of collected location points reaches a threshold requirement, at which point the user is prompted to terminate the collection.

[0218] In another embodiment, in S930, the spatial geometry restored in S920 is displayed on the UI interface of the image acquisition device. By viewing the scene geometry, the user can confirm which parts are missing and need further acquisition, and which parts are structurally complete and acquisition can be stopped.

[0219] Figure 13 shows a schematic diagram of the geometric structure of a scene point cloud according to an embodiment of this application.

[0220] The UI of the image acquisition device is shown in Figure 13. The dashed line 1001 is the virtual trajectory of the acquisition path of the corresponding image acquisition device. 1002, 1003, 1004 and 1005 are four virtual acquisition points on the virtual trajectory 1001 (the triangle on the acquisition point represents the acquisition view of the acquisition point), which correspond to the four real acquisition points of the image acquisition device.

[0221] Data was captured at the actual acquisition points corresponding to 1002, 1003, 1004 and 1005 respectively. The scene point cloud structure was restored based on the acquired data (2D video and / or 2D image, spatial pose information and depth information). The restored scene point cloud structure is shown in Figure 11, 1010.

[0222] Users can identify the missing state of the scene point cloud structure based on the scene point cloud structure shown in Figure 13, and then perform supplementary data acquisition.

[0223] Figure 14 shows a schematic diagram of a scene point cloud structure according to an embodiment of this application.

[0224] Following the scene point cloud structure shown in Figure 13, the user performs supplementary capture data acquisition. The UI display of the image acquisition device after the supplementary capture data acquisition is shown in Figure 14. The dashed line 1101 represents the virtual trajectory (corresponding to 1001) of the acquisition trajectory of the corresponding image acquisition device. Virtual acquisition points 1102, 1103, 1104, and 1105 correspond to virtual acquisition points 1002, 1003, 1004, and 1005.

[0225] Virtual acquisition points 1106, 1107, 1108, 1109, 1110, 1111, and 1112 correspond to seven real acquisition points supplemented by the user on the acquisition trajectory of the corresponding image acquisition device.

[0226] Following the scene point cloud structure shown in Figure 13, the user performs supplementary capture data collection at seven real acquisition points corresponding to 1106, 1107, 1108, 1109, 1110, 1111, and 1112. Based on the obtained capture data (2D video and / or 2D image, spatial pose information, and depth information), the scene point cloud structure is restored. The restored scene point cloud structure is shown at 1120 in Figure 14.

[0227] According to one embodiment of the method of this application, in the data acquisition stage, a device including but not limited to a camera module, IMU, depth sensor, etc. is used to capture the target scene, the data already acquired by the sensor is analyzed in real time, the integrity of the target scene acquisition is determined, and the user is guided to capture the area where the acquisition has not reached the integrity threshold, thereby obtaining complete capture data.

[0228] According to an embodiment of this application, the method can ensure the integrity of the generated 3D video, allowing users to immerse themselves in watching 3D video content from different angles, providing a stronger 3D feel and richer 3D video content, and improving the user's 3D video viewing experience.

[0229] In one embodiment, in S500, 3D video material data for the target scene and / or target object is generated based on one or more 2D videos containing the target scene and / or target object.

[0230] Figure 15 shows a flowchart of 3D video material data generation according to an embodiment of this application.

[0231] The electronic device executes the following process shown in Figure 15 to achieve S500.

[0232] S610, acquire one or more 2D videos and / or images containing the target scene and / or target object.

[0233] Specifically, in one embodiment, one or more 2D videos can be 2D videos and / or videos taken by any device (e.g., VR glasses, mobile phone or camera), by any user at any time (at the same time or at different times).

[0234] That is, 2D footage shot by different users and on different devices at different times can be used as source data to generate the same 3D video.

[0235] S620 identifies video frames and / or images related to a target scene and / or target object in one or more 2D videos and / or images.

[0236] The 2D material acquired in S610 contains content that is irrelevant to the key scene (main scene). Therefore, in S620, the video frames and / or images of the key scene are identified in the 2D material acquired in S610, and the video frames and / or images of the key scene are used as material data for subsequent generation of 3D video.

[0237] Key scenario identification methods include, but are not limited to, user-specified methods and automated detection.

[0238] For example, in one embodiment, the user specifies a key scene (including but not limited to manually selecting the target scene or object in the 2D material acquired in S610, or providing a description of the key scene), and identifies video frames and / or images matching the key scene in the 2D material acquired in S610.

[0239] For example, in another embodiment, the 2D video and / or image acquired by the automatic identification and detection S610 (based on target recognition and other algorithms) is determined to be a key scene when the number of times a scene appears exceeds a certain threshold, and the relevant video frames and / or images of the key scene are saved.

[0240] S630 generates 3D video material data based on video frames and / or images related to the target scene and / or target object.

[0241] 3D video material data includes, but is not limited to, key scene video frames and / or images acquired by the S620.

[0242] Specifically, in one embodiment, in S630, the key scene video frames and / or images identified in S620 are used to obtain the corresponding poses and scene spatial point clouds of the key scene video frames and / or images through scene reconstruction algorithms (such as SLAM, colmap, etc.), and are used together with the key scene video frames and / or images as 3D video material data.

[0243] In another embodiment, in S610, depth information synchronously recorded by the acquisition device during the capture of 2D video can be obtained. In S630, the depth information can be used as a priori for the scene reconstruction algorithm, or it can be directly used together with key scene video frames and / or images as 3D video material data.

[0244] Optionally, in one embodiment, the video content of the 3D video is supplemented and expanded in S520.

[0245] Specifically, in one embodiment, the video content of 3D videos is supplemented and expanded using artificial intelligence (AI). Intelligent recognition is used to supplement missing 3D video content with AI-generated supplements, providing users with a 3D video viewing experience beyond the actual shooting range. For example, when the scene is identified as a grassland, the scene is intelligently expanded to an infinitely distant grassland; when the scene is identified as a classroom, AI is used to generate environments such as school corridors and playgrounds outside the classroom; when the scene is identified as a famous historical site, as the user moves through the view, pre-stored materials from that site are used to generate 3D content beyond the shooting range.

[0246] According to a method of one embodiment of this application, the video content of a 3D video is supplemented and expanded, which can ensure the integrity of the generated 3D video and improve the user's 3D video viewing experience.

[0247] Optionally, in one embodiment, the observable area of ​​the 3D video virtual space is calculated in S520, and the best viewing route and area are recommended to the user based on the observable area of ​​the 3D video virtual space to avoid the user's line of sight / viewpoint moving to areas outside the 3D virtual video space, thereby improving the user's 3D video viewing experience.

[0248] Specifically, in one embodiment, the depth point cloud structure of the scene during the data acquisition process or the 3D point cloud structure derived from the new perspective generation model (the new perspective generation model called by S520) is obtained, the density of the point cloud is determined, and when the point cloud density exceeds a certain threshold, it means that this area can be viewed from any angle, avoiding areas with low point cloud density.

[0249] In another embodiment, explicit geometric structures, such as point clouds, meshes, and voxels, are derived from the new perspective generation model (called by S520). The observable region is determined based on the integrity of the explicit geometry of the model, ensuring that the geometry within the observable region is complete and dense. For example, after deriving the scene mesh from the new perspective generation model called by S520, a mesh integrity detection algorithm is used to set complete mesh regions as observable regions and incomplete mesh regions as unobservable regions.

[0250] Optionally, in one embodiment, in S530, the acquisition trajectory is directly used as the viewing path of the 3D video virtual space. The optimal viewing trajectory is provided to the user through the placement of virtual indicator icons in the UI interface. The interpolated and smoothed acquisition trajectory is used as the recommended observation trajectory, and the recommended observable range is expanded in the 6DoF direction based on each acquisition point.

[0251] According to one embodiment of the method of this application, the acquisition trajectory is used as the viewing path of the 3D video virtual space, which can effectively prevent the user's line of sight / viewpoint from moving to areas outside the 3D virtual video space, thereby improving the user's 3D video viewing experience.

[0252] Optionally, in one embodiment, in S530, when playing 3D video data, the playback environment is detected; the 3D video virtual space corresponding to the 3D video data is matched with the playback environment to obtain the matching result; and the 3D video data is guided to be played according to the matching result.

[0253] Specifically, in one embodiment, the spatial correspondence between the playback environment and the 3D video virtual space is determined based on the matching result between the 3D video virtual space and the playback environment; based on the spatial correspondence, an active area is set in the 3D video virtual space, and / or the position of the viewing point in the 3D video virtual space is determined.

[0254] For example, detecting the playback environment focuses on the ground area, walls, and typical objects that exist in both the real playback environment and the virtual 3D video. This includes aligning the virtual 3D video floor with the real playback environment floor, and aligning the walls in the virtual 3D video with the walls in the real physical environment. Electronic fences can be set up to mark areas in the 3D video that are impassable in the real playback environment outside the fences.

[0255] Specifically, explicit spatial structure information, such as meshes, point clouds, and voxels, is derived from the novel perspective generative model. This explicit spatial structure is then used as the virtual spatial environment for the 3D video. The camera and depth sensor of the MR glasses are utilized to detect the structure of the current viewing environment. The geometric structure of the viewing environment is identified through methods such as object detection and semantic segmentation, thus serving as the real-world environment for 3D video viewing.

[0256] According to one embodiment of the method in this application, 3D video data is guided for playback based on the matching result between the 3D video virtual space and the playback environment, and an active area is set in the 3D video virtual space. This can effectively prevent users from colliding with actual objects in the playback environment when changing their viewing angle or position while watching 3D videos.

[0257] For example, comparing the virtual environment with the actual viewing environment can identify whether there are identical objects in the two environments. Objects of the same type are then superimposed. For instance, if the 3D video content shows a family camping on the grass, and the actual viewing environment is the living room, then the virtual grass plane is superimposed on the living room floor. Similarly, if the 3D video content shows an exciting basketball game, and the user is actually watching from a sofa, then the audience seating in the 3D video is superimposed on the actual sofa position.

[0258] According to one embodiment of the method in this application, the playback of 3D video data is guided based on the matching result between the 3D video virtual space and the playback environment to determine the position of the viewing point in the 3D video virtual space. This allows the 3D video virtual space to be integrated with the playback environment, improving the user's viewing experience of 3D videos.

[0259] For example, in one application scenario, based on the 3D video generation device shown in Figure 6, the method flow shown in Figure 7 is implemented, and the specific implementation process is as follows.

[0260] The acquisition unit 601 uses VR glasses or a mobile phone to access available camera modules to acquire 2D video sequences or images. If the acquisition device is VR glasses, it simultaneously utilizes the VR glasses' inertial measurement unit and depth sensor to synchronously record the camera's relative motion and depth information corresponding to the 2D video frame sequence or image. If the acquisition device is a mobile phone, it simultaneously utilizes the mobile phone's inertial measurement unit to obtain the camera's motion information. If the mobile phone device has a LiDAR or ToF depth sensor, it utilizes it as well. The acquisition unit 601 analyzes the integrity of the acquired data in real time and provides guided acquisition prompts on the acquisition device's UI interface.

[0261] The computing unit 602 analyzes the scene type based on the collected data, calls the corresponding new perspective generation algorithm to generate 3D video data, and calculates the observable area.

[0262] The display unit 604 uses VR glasses to play 3D video data, analyzes the consistency between the real playback physical environment and the 3D video virtual environment, merges the 3D video virtual environment with the real playback physical environment, tracks the changes in the user's viewing point in real time, and generates the corresponding 3D video image from the target perspective.

[0263] Optionally, in one embodiment, in S530, in addition to providing an immersive viewing experience for the user, the system can also provide the function of editing the scene in the 3D video virtual space. For example, the user can modify the ambient light, stylize the scene, delete objects, place objects from a personal or official material library, and generate new scene content using AI, etc. The interaction methods for the editing function include, but are not limited to, using virtual buttons in the 3D virtual space, using physical controllers, using gesture recognition, and using eye-tracking recognition.

[0264] An embodiment of this application also proposes an electronic device. This electronic device is used to execute the method flow or part of the method flow described in the embodiments of this application.

[0265] Figure 16 is a schematic diagram of an electronic device structure according to an embodiment of this application.

[0266] As shown in FIG16, the electronic device 2500 includes a memory 2502 for storing computer program instructions and a processor 2501 for executing the program instructions. When the computer program instructions are executed by the processor 2501, the electronic device 2500 is triggered to execute the steps of the method described in the embodiments of this application.

[0267] Specifically, in one embodiment of this application, the aforementioned one or more computer programs are stored in the aforementioned memory 2502. The aforementioned one or more computer programs include instructions that, when executed by the aforementioned electronic device 2500, cause the aforementioned electronic device 2500 to perform the method steps described in the embodiments of this application.

[0268] It is understood that the structural description of the electronic device 2500 in this application does not constitute a specific limitation on the electronic device 2500. In other embodiments of this application, the electronic device 2500 may include other components besides the processor 2501 and the memory 2502.

[0269] The processor 2501 may be an on-chip device (SOC) that may include a central processing unit (CPU) and may further include other types of processors.

[0270] The processor 2501 may include, for example, a CPU, DSP, microcontroller, or digital signal processor, and may also include a GPU, embedded neural network processing units (NPUs), and image signal processors (ISPs). The processor may also include necessary hardware accelerators or logic processing hardware circuitry, such as an ASIC, or one or more integrated circuits for controlling the execution of the program in this application. Furthermore, the processor may have the function of operating one or more software programs, which may be stored in a storage medium.

[0271] Processor 2501 may include one or more processing units. For example, a processor may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent components or integrated into one or more processors. In some embodiments, electronic device 2500 may also include one or more processors 2501. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0272] In some embodiments, the processor 2501 may include one or more interfaces. These interfaces may include an inter-integrated circuit (I2C) interface, an integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. The USB interface is a USB standard-compliant interface, specifically a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface can be used to connect a charger to charge the electronic device, and can also be used for data transfer between the electronic device and peripheral devices.

[0273] Electronic device 2500 may also include an external memory interface for connecting an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with processor 2501 through the external memory interface to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0274] The memory 2502 may include a code storage area and a data storage area. The code storage area may store the operating system. The data storage area may store data created during the use of the electronic device 2500. Furthermore, the memory 2502 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage components, flash memory components, universal flash storage (UFS), etc.

[0275] The memory 2502 may be a read-only memory (ROM), other types of static storage devices that can store static information and instructions, random access memory (RAM), or other types of dynamic storage devices that can store information and instructions. It may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices. Alternatively, it may be any computer-readable medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer.

[0276] Processor 2501 and memory 2502 can be combined into a single processing device, but more commonly they are separate components.

[0277] An embodiment of this application also proposes an electronic chip. This electronic chip is used to execute the method flow or part of the method flow described in the embodiments of this application. For example, the electronic chip can be a GPU.

[0278] Specifically, the electronic chip includes a processor for executing program instructions. When the computer program instructions are executed by the processor, the electronic chip is triggered to perform the steps described in the embodiments of this application. The processor of the electronic chip can refer to the processor of the above-described electronic device.

[0279] Optionally, the devices, apparatuses, and modules described in the embodiments of this application may be implemented by computer chips or physical entities, or by products with certain functions.

[0280] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media containing computer-usable program code.

[0281] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0282] Specifically, one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to execute the method provided in the embodiment of this application.

[0283] An embodiment of this application also provides a computer program product, which includes a computer program that, when run on a computer, causes the computer to perform the method provided in the embodiment of this application.

[0284] The embodiments described in this application are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.

[0285] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0286] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0287] It should also be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0288] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0289] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0290] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0291] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments of this application can be implemented using electronic hardware, computer software, or a combination of electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0292] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0293] [Revised according to Rule 91, Amended 22.01.2025] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application shall be determined by the protection scope of the claims.

Claims

1. A method for generating 3D video, characterized in that, The method is applied to an electronic device, and the method includes: Based on the captured data, 3D video material data is obtained, including 2D video and / or 2D images; Analyze 3D video scene types based on 3D video material data; Based on the type of 3D video scene, obtain the corresponding new perspective generation model for the 3D video scene type; Using the new perspective generation model corresponding to the 3D video scene type, 3D video data is generated based on the 3D video material data.

2. The method of claim 1, wherein, The captured data includes one or more 2D videos and / or 2D images containing the target scene and / or target object; the acquisition of 3D video material data based on the captured data includes: Generate 3D video material data for the target scene and / or the target object based on one or more 2D videos and / or 2D images containing the target scene and / or the target object.

3. The method of claim 2, wherein, The step of generating 3D video material data for the target scene and / or the target object based on one or more 2D videos and / or 2D images containing the target scene and / or target object includes: Identify video frames and / or images in one or more 2D videos and / or 2D images that are related to the target scene and / or the target object; The 3D video material data is generated based on the video frames and / or images related to the target scene and / or the target object.

4. The method of claim 1, wherein, The method further includes: Collecting the captured data includes guiding the collection of the captured data based on data collection integrity.

5. The method of claim 4, wherein, The method of guiding the acquisition of captured data based on data acquisition integrity includes: The acquisition of the captured data is guided by the completeness of the acquisition trajectory and / or acquisition point location of the captured data.

6. The method of claim 5, wherein, The step of guiding the acquisition of the captured data based on the acquisition trajectory and / or acquisition point location of the captured data includes: Obtain the ideal acquisition method that matches the target scene and / or target object, wherein the ideal acquisition method includes the ideal acquisition trajectory and / or the ideal acquisition point position; By comparing the acquisition trajectory and / or acquisition point location of the captured data with the ideal acquisition trajectory and / or ideal acquisition point location, the user is guided to supplement the acquisition data at the missing acquisition trajectory and / or acquisition point location based on the comparison results.

7. The method of claim 4, wherein, The method of guiding the acquisition of captured data based on data acquisition integrity includes: The acquisition of the captured data is guided by the completeness of the acquired data.

8. The method of claim 7, wherein, The step of guiding the acquisition of the captured data based on the completeness of the acquired data includes: Based on the captured data, calculate the captured location and determine the coverage of the captured data on the target scene and / or target object; Based on the coverage of the target scene and / or target object by the already collected capture data, guide the user to supplement the collection of capture data for the corresponding missing coverage areas.

9. The method of claim 7, wherein, The step of guiding the acquisition of the captured data based on the completeness of the acquired data includes: Based on the captured data that has been collected, calculate the spatial geometric structure corresponding to the captured data that has been collected; Based on the completeness of the spatial geometry corresponding to the captured data, the user is guided to supplement the captured data of the missing spatial geometry.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: Calculate the observable area of ​​the 3D video virtual space corresponding to the 3D video data.

11. The method according to any one of claims 1-9, characterized in that, The method further includes playing the 3D video data, wherein playing the 3D video data includes: The playback environment is detected while the 3D video data is being played; Match the 3D video virtual space corresponding to the 3D video data with the playback environment to obtain the matching result; Based on the matching results, the playback of the 3D video data is guided.

12. The method of claim 11, wherein, The guidance for playing the 3D video data includes: Determine the spatial correspondence between the playback environment and the 3D video virtual space; Based on the spatial location correspondence, an active area is set in the 3D video virtual space, and / or the location of the viewing point in the 3D video virtual space is determined.

13. An electronic device, comprising: The electronic device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to perform the method steps as described in any one of claims 1-12.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1-12.