3D video generation method, 3D video viewing method, and electronic device
Through deep learning and artificial intelligence's new perspective synthesis algorithm, combined with depth sensors and inertial measurement units to guide data collection, high-fidelity 3D video is generated, solving the accuracy and viewpoint limitations of 3D video generation in existing technologies, and achieving an immersive 6-degree-of-freedom 3D video experience.
Patent Information
- Application Number
- PCT/CN2024/143731
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-31
- Filing Date
- 2024-12-30
- Publication Date
- 2025-10-09
AI Technical Summary
Existing technologies have difficulty in effectively generating high-quality 3D videos, especially when converting monocular 2D videos or small-baseline binocular videos to 3D videos. There are accuracy issues and limited viewpoints, and it is difficult for ordinary users to obtain depth information and the modeling cost is high.
A new perspective synthesis algorithm based on deep learning and artificial intelligence is adopted. By analyzing the 3D video scene type, the corresponding new perspective generation model is used to generate 3D video data from 2D videos and images. The depth sensor and inertial measurement unit are combined to guide data collection to ensure the integrity of the collection, and neural rendering technology is used to generate high-fidelity 3D video.
It realizes the generation of high-quality 3D videos from 2D videos shot by ordinary devices, providing an immersive viewing experience and supporting 6-degree-of-freedom viewing. Users can change the viewing angle at will, providing a stronger sense of 3D and immersion, and is suitable for devices such as VR glasses.
Smart Images

Figure CN2024143731_09102025_PF_FP_ABST
Abstract
Description
3D video generation method, viewing method and electronic device
[0001] This application claims priority to Chinese patent application number 2024103860747, filed with the State Intellectual Property Office of China on March 31, 2024, entitled “A method for generating, viewing and electronic device for 3D video”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a 3D video generation method, a viewing method, and an electronic device. Background Art
[0003] With the rapid development of multimedia-related technologies and hardware equipment, various 3D devices have been launched on the market. Among them, virtual reality (VR) glasses have gradually been commercialized and are increasingly accepted by ordinary consumers.
[0004] VR glasses can play 3D videos, providing users with a stereoscopic and immersive viewing experience. In order to obtain 3D videos, a 3D video generation method is required. Summary of the Invention
[0005] In response to the problem of how to generate 3D video, the present application provides a 3D video generation method, viewing method and electronic device, and the present application also provides a computer-readable storage medium.
[0006] The embodiments of this application adopt the following technical solutions:
[0007] In a first aspect, the present application provides a 3D video generation method, which is applied to an electronic device and includes:
[0008] Acquiring 3D video material data based on the captured data, where the captured data includes 2D video and / or 2D image;
[0009] Analyze 3D video scene types based on 3D video material data;
[0010] According to the 3D video scene type, a new perspective generation model corresponding to the 3D video scene type is obtained;
[0011] The new viewing angle generation model corresponding to the 3D video scene type is used to generate 3D video data according to the 3D video material data.
[0012] According to the method of the first aspect, in the 3D video data generation stage, the corresponding new perspective generation model is called according to the 3D video scene type to generate 3D video data, so that the new perspective synthesis algorithm based on deep learning and artificial intelligence can obtain a more complete scene synthesis perspective, so that when playing 3D video data, the user can immersively watch different angles of the 3D video content, providing a stronger 3D sense and richer 3D video content.
[0013] In an implementation of the first aspect, the captured data includes one or more 2D videos and / or 2D images containing a target scene and / or a target object; and acquiring 3D video material data based on the captured data includes:
[0014] 3D video material data for the target scene and / or the target object is generated according to one or more 2D videos and / or 2D images containing the target scene and / or the target object.
[0015] According to the above-mentioned implementation method, the source of 3D video material data can be 2D video and / or 2D images shot using any image acquisition device (for example, VR glasses or mobile phone devices), and the 2D video and / or 2D images can come from the same scene material shot by different users at different times.
[0016] According to the above implementation method, the difficulty of obtaining 3D video material data is reduced.
[0017] In an implementation of the first aspect, generating 3D video material data for the target scene and / or the target object based on one or more 2D videos and / or 2D images containing the target scene and / or the target object includes:
[0018] identifying video frames and / or images related to the target scene and / or the target object in the one or more 2D videos and / or 2D images;
[0019] The 3D video material data is generated based on the video frames and / or images related to the target scene and / or the target object.
[0020] In an implementation of the first aspect, the method further includes:
[0021] Collecting the captured data includes directing the collection of the captured data based on data collection completeness.
[0022] According to the method of the above implementation mode, during the data acquisition stage, the target scene is photographed using a device including but not limited to a camera module, an inertial measurement unit, a depth sensor and other sensors, and the data collected by the sensor is analyzed in real time to determine the completeness of the target scene acquisition. The user is guided to photograph the areas where the acquisition does not reach the completeness threshold, thereby obtaining complete captured data.
[0023] According to the above implementation method, the integrity of the generated 3D video can be ensured, allowing users to immersively watch different angles of the 3D video content, providing a stronger 3D sense and richer 3D video content, and improving the user's 3D video viewing experience.
[0024] In an implementation of the first aspect, guiding the collection of the captured data based on data collection integrity includes:
[0025] The collection of the captured data is guided according to the completeness of the collection trajectory and / or the collection point positions of the captured data.
[0026] In an implementation of the first aspect, guiding the collection of the captured data according to a collection trajectory and / or a collection point position of the captured data includes:
[0027] Acquire an ideal acquisition method that matches the target scene and / or target object, wherein the ideal acquisition method includes an ideal acquisition trajectory and / or an ideal acquisition point position;
[0028] Compare the acquisition trajectory and / or acquisition point position of the captured data with the ideal acquisition trajectory and / or ideal acquisition point position, and guide the user to supplement the acquisition data at the missing acquisition trajectory and / or acquisition point position based on the comparison result.
[0029] In an implementation of the first aspect, guiding the collection of the captured data based on data collection integrity includes:
[0030] According to the completeness of the collected captured data, the collection of the captured data is guided.
[0031] In an implementation of the first aspect, guiding the collection of the captured data based on the integrity of the collected captured data includes:
[0032] Calculating the collected position based on the collected capture data, and determining coverage of the target scene and / or target object by the collected capture data;
[0033] According to the coverage of the target scene and / or target object by the collected capture data, the user is guided to supplement the capture data corresponding to the missing coverage part.
[0034] In an implementation of the first aspect, guiding the collection of the captured data based on the integrity of the collected captured data includes:
[0035] Calculating a spatial geometric structure corresponding to the collected capture data based on the collected capture data;
[0036] According to the completeness of the spatial geometric structure corresponding to the captured data that has been collected, the user is guided to supplement the collection of captured data corresponding to the missing spatial geometric structure.
[0037] In an implementation of the first aspect, the method further includes:
[0038] An observable area of a 3D video virtual space corresponding to the 3D video data is calculated.
[0039] According to the above implementation method, during the 3D video playback stage, users can be guided to watch 3D videos based on the observable area of the 3D video virtual space, avoiding the user's line of sight / viewing angle moving to the area outside the 3D virtual video space, thereby improving the user's 3D video viewing experience.
[0040] In an implementation of the first aspect, the method further includes playing the 3D video data, and the playing the 3D video data includes:
[0041] When playing the 3D video data, detecting a playback environment;
[0042] Matching the 3D video virtual space corresponding to the 3D video data with the playback environment to obtain a matching result;
[0043] According to the matching result, the 3D video data is guided to be played.
[0044] In an implementation of the first aspect, the guiding the playback of the 3D video data includes:
[0045] Determining a spatial position correspondence between the playback environment and the 3D video virtual space;
[0046] According to the spatial position correspondence, an active area is set in the 3D video virtual space, and / or the position of the viewing point in the 3D video virtual space is determined.
[0047] According to the above implementation method, based on the matching result between the 3D video virtual space and the playback environment, the playback of 3D video data is guided and a movable area is set in the 3D video virtual space. This can effectively prevent users from colliding with real objects in the playback environment when changing the viewing angle or viewing position of the 3D video.
[0048] According to the above implementation method, the 3D video data is played back and the position of the viewing point in the 3D video virtual space is determined based on the matching result between the 3D video virtual space and the playback environment. This can integrate the 3D video virtual space with the playback environment, improving the user's viewing experience of 3D videos.
[0049] In a second aspect, the present application provides an electronic device, comprising a memory for storing computer program instructions and a processor for executing computer program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to execute the method steps described in the first aspect.
[0050] In a third aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] FIG1 is a flowchart of a DIBR implementation according to an embodiment of the present application;
[0052] FIG2 shows a flowchart of a neural rendering implementation according to an embodiment of the present application;
[0053] FIG3 is a flowchart illustrating a DIBR implementation according to an embodiment of the present application;
[0054] FIG4 is a flowchart illustrating a DIBR implementation according to an embodiment of the present application;
[0055] FIG5 is a flowchart illustrating an MBR implementation according to an embodiment of the present application;
[0056] FIG6 is a schematic structural diagram of a 3D video generating device according to an embodiment of the present application;
[0057] FIG7 is a flow chart of a 3D video generation method according to an embodiment of the present application;
[0058] FIG8 shows left and right eye views of a 3D image according to an embodiment of the present application;
[0059] FIG9 shows left and right eye views of a 3D image according to an embodiment of the present application;
[0060] FIG10 is a flowchart showing a captured data collection process according to an embodiment of the present application;
[0061] FIG11 is a schematic diagram showing a recommended collection method according to an embodiment of the present application;
[0062] FIG12 is a flowchart showing a captured data collection process according to an embodiment of the present application;
[0063] FIG13 is a schematic diagram showing a scene point cloud structure according to an embodiment of the present application;
[0064] FIG14 is a schematic diagram showing a scene point cloud structure according to an embodiment of the present application;
[0065] FIG15 is a flowchart showing 3D video material data generation according to an embodiment of the present application;
[0066] FIG16 is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0067] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0068] The terms used in the implementation section of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.
[0069] 3D video can provide users with a three-dimensional and immersive viewing experience. 3D video is primarily based on a rich variety of binocular 3D video resources. It utilizes the parallax between the left and right eyes, which fuse in the brain to create a three-dimensional 3D image. This allows users to perceive objects as being far and near, resembling a real three-dimensional scene.
[0070] One technical solution for capturing 3D video is active stereo capture, using binocular cameras to ensure consistent lens, aperture, and colorimetry, and synchronized signals. However, common mobile camera devices, such as mobile phones and cameras, can only capture monocular 2D video or small-baseline binocular video. Active stereo capture requires specialized equipment, making this method relatively expensive.
[0071] Another technical solution for obtaining 3D video is passive video generation, which converts monocular 2D video into 3D video, mainly by estimating binocular video (two-channel video) through computer vision and computer graphics methods.
[0072] Generally, the key technology in converting monocular 2D video to 3D video lies in virtual viewpoint synthesis based on 2D video. Virtual viewpoint synthesis technology can be divided into depth image based rendering (DIBR), model based rendering (MBR), and neural rendering technology based on the implementation method and auxiliary tools.
[0073] DIBR is a method of rendering virtual viewpoint images based on a reference viewpoint texture map and a corresponding depth map through a 3D mapping equation. Due to its fast speed and low complexity, it is generally widely used in the field of new perspective synthesis.
[0074] FIG1 is a flowchart showing a DIBR implementation process according to an embodiment of the present application.
[0075] In one embodiment, the basic flow chart of DIBR is shown in FIG1 .
[0076] The DIBR process can be simply described as:
[0077] S110, reversely projecting one or more reference viewpoint images (e.g., reference image 101, depth image 102) into a world coordinate system;
[0078] S120, re-projecting to the virtual viewpoint plane;
[0079] S130, generating a target viewpoint virtual image (virtual viewpoint image) through image fusion;
[0080] S140, performing post-processing (e.g., S141, removing artifacts and overlaps; S142, filling holes);
[0081] S150: Obtain the final image and output it.
[0082] The core technology of DIBR is 3D image warping (3D-warping), which maps the reference view pixels to the target view through a 3D transformation equation.
[0083] The disadvantages of DIBR include accuracy issues and viewpoint limitations.
[0084] The accuracy problem refers to the fact that the accuracy of the depth map directly affects the quality of the synthesized perspective. The core of DIBR is the utilization of depth information. The three-dimensional information of the reference viewpoint is constructed through depth information, and then the three-dimensional information of the target viewpoint is obtained through mapping transformation. When there is noise in the depth map, the quality of the synthesized perspective image is poor.
[0085] The viewpoint limitation problem refers to the situation that when the target virtual viewpoint is too far away from the reference viewpoint, the holes generated in the target viewpoint are difficult to be effectively filled, which greatly affects the quality of the target perspective synthetic image.
[0086] The MBR solution renders images from a specified target perspective based on a 3D geometric model. In the MBR process, an explicit 3D model of the 3D object or scene is first created, such as a point cloud, voxel, or mesh. Textures, lighting, and shadows are then added to the 3D model. Finally, a realistic 2D image is produced using a rendering algorithm (such as rasterization or ray tracing).
[0087] In the MBR scheme, although the use of explicit geometric models can produce high-quality synthetic images from relatively few input images, it is difficult to accurately estimate the scene geometry model in difficult scenes such as textureless areas, highlights, reflections, and repeated textures. When the reconstructed 3D model quality is poor, it is usually impossible to fully recover the target perspective image from unreconstructed areas or over-reconstructed areas. Although specialized tools can be used to scan existing real objects for 3D modeling, this is not realistic for most ordinary users. At the same time, the rendering process requires additional input of physical properties such as lighting and materials of the scene, which are relatively difficult to obtain and estimate.
[0088] Neural rendering is a rapidly emerging field that allows for compact scene representations, leveraging neural networks to learn information about the scene, such as spatial density and color, from existing data. The key idea behind neural rendering is to combine classical physics-based computer graphics with deep learning to generate realistic images from a target viewpoint.
[0089] FIG2 shows a flowchart of a neural rendering implementation according to an embodiment of the present application.
[0090] As shown in FIG2 , the input of the neural rendering method is a 2D image sequence (input image 201 ), and other additional inputs (other inputs 202 (e.g., depth, optical flow, geometry, etc.)) may also be added. The input information is uniformly converted into spatial representation data 210. Spatial representation data 210 is input to a neural network 220.
[0091] The neural network 220 is supervised to represent the shape or appearance of a specific scene, and a preset rendering algorithm 230 is used for rendering (such as rasterization 231 or ray tracing 232) to generate a target perspective image 240 (new perspective image). The neural network 220 uses learnable elements in the scene (such as the density and color of objects or scenes) to express scene information, thereby rendering the image under the target viewpoint. At present, neural rendering technology has made significant progress in both the quality of new perspective synthesis and the real-time rendering, which also makes it possible to apply neural rendering technology to consumer-grade VR glasses. Neural rendering technology can achieve high-fidelity target perspective image synthesis at the static or dynamic, object level or large scene level.
[0092] For example, in a feasible technical solution, based on the DIBR method, a reference viewpoint image and depth are obtained, and the reference viewpoint image is warped to the target viewpoint according to the depth value.
[0093] FIG3 is a flowchart showing a DIBR implementation according to an embodiment of the present application.
[0094] As shown in Figure 3:
[0095] S310, obtaining a reference viewpoint video and a reference viewpoint depth map video corresponding to the reference viewpoint video, decomposing the reference viewpoint video into a frame sequence of reference viewpoint images, and decomposing the reference viewpoint depth map video into a frame sequence of reference viewpoint depth maps;
[0096] S320, mapping the reference viewpoint image of each frame to the virtual viewpoint to generate the virtual viewpoint original image of each frame;
[0097] S330, repairing the reference viewpoint image and the reference viewpoint depth map of each frame, and mapping the repaired reference viewpoint image of each frame to a virtual viewpoint to generate a virtual viewpoint auxiliary image of each frame;
[0098] S340 , repairing holes in the virtual viewpoint original image according to the virtual viewpoint auxiliary images of each frame to generate a virtual viewpoint final image;
[0099] S350, synthesizing the final virtual viewpoint images of each frame to generate a virtual viewpoint video;
[0100] S360: synthesize the virtual viewpoint video and the reference viewpoint video to generate a multi-viewpoint 3D video.
[0101] Based on the above scheme, each frame image can be repaired by extracting the reference image, expanding the image boundary, and repairing the mutation area, which can effectively solve the boundary holes in the original image of the virtual viewpoint and the holes in the mutation area of the internal reference viewpoint depth map.
[0102] However, based on the above scheme, only the target viewpoint image with a small deviation from the reference viewpoint can be obtained, and the new viewing angle range is subject to certain restrictions. When the target viewpoint is far away from the reference viewpoint, the new viewing angle hole is large and difficult to repair, and the image quality is poor.
[0103] Furthermore, the above solution relies on depth and needs to convert 2D camera images into 3D point clouds through depth values, so an additional depth sensor or depth calculation unit is required.
[0104] For another example, in another feasible technical solution, based on the DIBR method, a camera device is used to obtain a reference viewpoint image without providing depth information, and a 2D video is converted into a 3D video through a depth estimation method.
[0105] FIG4 is a flowchart showing a DIBR implementation according to an embodiment of the present application.
[0106] As shown in Figure 4:
[0107] In the depth information extraction stage, a U-shaped convolutional network is trained using existing 3D movies as the source dataset. This results in a high-performance network model for frame-by-frame depth estimation of 2D videos, and a small neural network is used to optimize the edge-preserving and smoothing of the depth map. In the viewpoint synthesis stage, a depth map-based viewpoint synthesis algorithm without camera parameters is proposed, and a symmetrical rendering strategy from the center to the sides is used to synthesize left and right virtual viewpoints. Finally, in the image restoration stage, a block matching-based image restoration algorithm combined with time domain information is proposed to fill and repair cracks and holes in the left and right viewpoints. It is possible to convert 2D to 3D videos without any relevant parameter information of the original 2D video. This not only effectively processes high-resolution images, but also achieves good conversion results and is fast.
[0108] Based on the above solution, although there is no need to provide additional depth sensors and the depth information is estimated using a pre-trained depth estimation network, data acquisition is more convenient, but there is still the problem of limited target viewing angle.
[0109] For example, in another feasible technical solution, based on the MBR method, multi-eye depth cameras are used to obtain a three-dimensional model corresponding to the target scene, such as a point cloud or polygonal representation of the target scene (and objects in the scene). The model can reflect the three-dimensional geometric structure of the scene (and objects in the scene), and the target perspective image of the 3D video model is drawn based on the obtained interaction parameters.
[0110] FIG5 is a flowchart showing an MBR implementation according to an embodiment of the present application.
[0111] As shown in Figure 5:
[0112] S410, respectively obtaining depth video streams of at least three camera perspectives of the same scene;
[0113] S420, determining a target foreground point cloud and a target background point cloud corresponding to the depth video streams of the at least three camera perspectives;
[0114] S430 , processing the target foreground point cloud and the target background point cloud according to a target point cloud processing method corresponding to the depth video stream to obtain a 3D video model corresponding to the depth video stream.
[0115] Based on the above scheme, a 3D video model is obtained according to the input of multiple depth cameras, and a new perspective image with a large angle can be obtained based on the 3D explicit model. However, this method requires that the input image comes from at least a binocular depth camera, which does not conform to the daily settings of people using mobile phones or ordinary camera devices to shoot videos.
[0116] In order to more conveniently obtain 3D video, an embodiment of the present application further provides a 3D video generating device.
[0117] FIG6 is a schematic structural diagram of a 3D video generating device according to an embodiment of the present application.
[0118] As shown in FIG6 , the device includes a collection unit 601 , a calculation unit 602 , and a storage unit 603 .
[0119] The acquisition unit 601 is used to acquire 3D video material data for generating a 3D video.
[0120] In one embodiment, the acquisition unit 601 directly receives 3D video material data output by other electronic devices.
[0121] In another embodiment, the acquisition unit 601 receives capture data output by other electronic devices, processes the capture data, and generates 3D video material data, where the capture data includes 2D video and / or 2D image.
[0122] In another embodiment, the acquisition unit 601 includes a data acquisition component (eg, a visual sensor, an inertial measurement device, a depth sensor, etc.), and the acquisition unit 601 acquires captured data, processes the captured data, and generates 3D video material data.
[0123] For example, in one embodiment, the acquisition unit 601 includes: a visual sensor (such as one or more ordinary cameras or depth cameras) for acquiring 2D video materials; an inertial measurement unit for obtaining the relative motion of the camera device or the relative motion of the VR glasses device; a depth sensor, such as ToF, laser equipment, etc., for obtaining scene point cloud depth information.
[0124] The calculation unit 602 is configured to generate 3D video data for 3D video playback according to the 3D video material data.
[0125] For example, in one embodiment, the computing unit 602 includes a CPU, a GPU, a cache, registers, etc., which are used to run an operating system and process various algorithm modules involved in generating 3D video data, such as real-time analysis of collected data and new perspective generation algorithms.
[0126] The storage unit 603 includes internal memory and external storage, and is used for collecting data, generating new perspective algorithm data, reading and writing temporary data, etc.
[0127] Optionally, in one embodiment, the apparatus further includes a display unit 604. The display unit 604 is configured to play 3D video data and present a 3D video playback effect. For example, the display unit 604 may be a VR glasses device.
[0128] In the description of the embodiments of the present application, for the convenience of description, the description of the coloring device is divided into various modules according to their functions. The division of each module is only a division of logical functions. When implementing the embodiments of the present application, the functions of each module can be implemented in the same or one or more software and / or hardware.
[0129] Specifically, the device proposed in the embodiment of the present application can be fully or partially integrated into a physical entity (for example, a GPU or other type of processor) during actual implementation, or it can be physically separated. And these modules can all be implemented in the form of software calling through a processing element; they can also all be implemented in the form of hardware; some modules can also be implemented in the form of software calling through a processing element, and some modules can be implemented in the form of hardware. For example, the detection module can be a separately established processing element, or it can be integrated in a chip of an electronic device. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. During the implementation process, each step of the above method or each of the above modules can be completed by an integrated logic circuit of hardware in the processor element or an instruction in the form of software.
[0130] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0131] An embodiment of the present application further provides a method for generating a 3D video. In the method of the embodiment of the present application, a 3D video meeting a 6-degree-of-freedom (DoF) viewing requirement is generated based on a 2D video and / or a 2D image.
[0132] FIG7 is a flow chart of a 3D video generation method according to an embodiment of the present application.
[0133] S500: Acquire 3D video material data based on captured data, where the captured data includes but is not limited to 2D video and / or 2D image.
[0134] The captured data may include 3D point data from a depth sensor, device motion data from an inertial measurement unit, in addition to 2D video and / or 2D image data.
[0135] For example, in one embodiment, only 2D video and / or 2D images are used as 3D video material data.
[0136] For example, in one embodiment, 2D video and / or 2D image and depth information are used as 3D video material data.
[0137] For example, in one embodiment, 2D video and / or 2D image and device motion information are used as 3D video material data.
[0138] For example, in one embodiment, 2D video and / or 2D image, depth information, and device motion information are used as 3D video material data.
[0139] S510: Analyze the 3D video scene type based on the 3D video material data.
[0140] For example:
[0141] Scene size classification: Determine the scene size based on the actual spatial volume of the scene point cloud;
[0142] Classification of scene shooting mode: confirming the shooting mode of the scene (such as surround shooting, front-facing shooting, etc.) based on the data acquisition trajectory of the captured data;
[0143] Scene motion and stillness attribute classification: motion and stillness detection based on the image content in the captured data;
[0144] Object classification: Detecting objects based on the image content in the captured data.
[0145] S511 , according to the 3D video scene type, obtaining a Novel View Synthesis (NVS) model corresponding to the 3D video scene type.
[0146] For example, in one implementation, based on the 3D video scene type, an NVS model corresponding to the 3D video scene type is called from pre-generated NVS models.
[0147] For another example, in another implementation, a model algorithm corresponding to the 3D video scene type is determined according to the 3D video scene type, and an NVS model corresponding to the 3D video scene type is generated based on the determined model algorithm.
[0148] Specifically, in one embodiment, the new perspective generation model for producing 3D video data includes multiple algorithms, such as:
[0149] For scenes shot with dynamic long video sequences, including but not limited to using dynamic neural network models based on image-based rendering (IBR);
[0150] For scenes with dynamic small-scale surround shooting, including but not limited to models using tensor decomposition or voxel decomposition methods;
[0151] For handheld selfie scenes, including but not limited to models of deformable field methods;
[0152] Scenes shot for static scenes, including but not limited to models using explicit geometric primitives based rasterization rendering methods.
[0153] S520 : Generate 3D video data according to the 3D video material data using a Novel View Synthesis (NVS) model corresponding to the 3D video scene type.
[0154] For different types of 3D video scenes, different new perspective generation models are used to generate 3D video data. 3D video data includes but is not limited to implicit model parameter form, explicit geometric form, or hybrid form.
[0155] In one embodiment, after S520, S530 is further included.
[0156] S530: Perform interactive 3D video viewing based on the 3D video data generated in S520.
[0157] Specifically, the real-time posture of the user's current viewing angle is calculated, and the new perspective generation model renders the corresponding picture of the observation angle in real time according to the current observation posture.
[0158] In one embodiment, the current viewing angle pose is derived from the viewing device's IMU motion information or a Simultaneous Localization and Mapping (SLAM) algorithm. For example, if a user wears VR glasses, the 3D video space is stationary relative to the real physical world. The user moves within the 3D video space, and the 3D video image at the current position is rendered based on the user's current movement position.
[0159] In another embodiment, the current observation angle and position is manually specified by the user. For example, the user wears VR glasses and is stationary relative to the real physical world. The user selects different observation angles and positions in the 3D video space, and the corresponding 3D video images are generated based on the user's currently selected observation angles and positions.
[0160] According to the method of one embodiment of the present application, 6DoF 3D video is generated based on input captured data (2D video and / or 2D images). The input data can be in any image / video format, supporting both monocular and multi-view video. As the user's viewing angle changes, images from any new perspective can be generated.
[0161] According to a method in one embodiment of the present application, a 3D video that meets 6DoF viewing requirements is generated based on captured data (2D video and / or 2D images) collected by a device such as a mobile phone, camera, or VR glasses. When a user watches the 3D video wearing VR glasses, a high-fidelity image corresponding to the target viewpoint is rendered in real time as the viewing point changes, creating a sense of being in a virtual 3D video environment, providing the user with a sense of three-dimensionality and immersion that surpasses that of conventional 3D videos based on binocular parallax.
[0162] Furthermore, according to the method of the embodiment of the present application, the generated 3D video data is not limited to application and 3D video playback. Those skilled in the art can use 3D video data according to actual needs (for example, using 3D video data for 3D modeling), thereby realizing a variety of different application scenarios (for example, augmented reality (AR), virtual reality, and mixed reality (MR)).
[0163] Specifically, in one embodiment, the new perspective generation model obtained in S511 is a neural network model.
[0164] According to the method of one embodiment of the present application, there are two advantages of using a new perspective to generate a neural network model: first, the input is not restricted, and only ordinary shooting equipment such as mobile phones and cameras can be used to collect video data, which is convenient and operable for ordinary users; second, the target perspective is not restricted, and in the absence of an explicit 3D model prior, a target viewpoint image that is far away from the reference perspective can still be generated.
[0165] For example, when watching 3D videos using VR glasses, the user's 3D experience comes only from the depth perception caused by binocular parallax.
[0166] FIG8 shows left and right eye views of a 3D image according to an embodiment of the present application.
[0167] As shown in FIG8 , there is a slight difference in perspective between the left-eye view 701 and the right-eye view 702 . The parallax of the cube with a closer depth of field is larger, and the parallax of the triangle with a farther depth of field is smaller. When the left eye views the left view 701 and the right eye views the right view 702 , the 2D image contents with different parallaxes of the left and right eyes are fused in the brain at different distances, thereby producing a 3D sense.
[0168] However, the left and right eye views shown in Figure 8 are fixed. The user's 3D experience is primarily based on binocular parallax, which creates a 3D sense by perceiving objects at different distances. Regardless of where the user is viewing the content within the actual physical environment, the 3D content they see remains the same.
[0169] According to an embodiment of the present application, a neural network model can be used to generate a wider range of new perspective images. According to an embodiment of the present application, as the user's viewing angle changes, different angles of the 3D scene can be viewed, providing the user with a more immersive experience.
[0170] FIG9 shows left and right visual views of a 3D image under two different observation angles according to an embodiment of the present application.
[0171] As shown in Figure 9, when the user is viewing from the right side of the cube, the left eye view is shown as 801, and the right eye view is shown as 802. There is a slight difference in perspective between left-eye view 801 and right-eye view 802. The parallax of the cube with a closer depth of field is larger, while the parallax of the triangle with a farther depth of field is smaller. When the left eye views left view 801 and the right eye views right view 802, the 2D images with different parallaxes for the left and right eyes are fused in the brain at different distances, thus creating a 3D perception.
[0172] According to the method of one embodiment of the present application, when the user's observation perspective moves from the right side of the cube to the left side, the left eye view is shown in Figure 803, and the right eye view is shown in Figure 804. There is a slight difference in perspective between the left eye view 903 and the right eye view 904. When the left eye views the left view 903 and the right eye views the right view 904, the 2D image content with different parallaxes of the left and right eyes is fused in the brain at different distances, thus creating a 3D perception. At the same time, as the observation perspective changes, the user sees different sides of the cube, and richer 3D information can provide the user with a more immersive 3D experience.
[0173] A drawback of current neural rendering-based methods for synthesizing new perspectives is that when the input reference viewpoints don't cover all angles, the generated target viewpoint images can be noisy, impacting the user's perception. In dynamic scenes, high-quality perspective synthesis relies on multiple cameras capturing data from different angles simultaneously. In daily life, people often use mobile phones and camera devices to capture monocular 2D videos, or dual-camera phones and VR glasses to capture binocular videos with smaller baselines.
[0174] In response to the above situation, in one embodiment, based on 2D monocular video or small baseline binocular video, a 6DoF immersive 3D video viewing experience is provided to users through methods such as guided acquisition and observable area recommendation.
[0175] Specifically, in one embodiment, during the process of collecting the captured data, the collection of the captured data is guided based on the collection completeness of the collected captured data (eg, 2D video and / or 2D image).
[0176] In one embodiment, the collection of captured data is guided based on the integrity of the collection trajectory and / or collection point locations of the captured data. Specifically, the collection trajectory and / or collection point locations during the capture of the captured data are monitored to ensure that the user has collected data at all expected collection trajectories and / or collection point locations.
[0177] In one embodiment, the scene spatial geometry is restored based on the collected data, and the user determines which parts of the scene need to be further collected or no longer need to be collected based on the current scene spatial integrity. For example, on a collection device equipped with a depth sensor, the scene spatial structure is restored and displayed based on the collected depth information (including but not limited to 3D point clouds, meshes, and other geometric expressions). The user can judge the integrity of the current scene spatial structure by observing it, and can continue collecting in areas where the spatial structure is poorly restored, and stop collecting in areas where it is relatively complete, ultimately achieving complete collection of the target scene.
[0178] FIG10 is a flowchart showing a capture data collection process according to an embodiment of the present application.
[0179] The electronic device executes the following process shown in FIG. 10 .
[0180] S711: Determine an ideal acquisition trajectory and / or an ideal acquisition point position.
[0181] In one implementation, a target scene and / or target object type is identified, and an ideal acquisition trajectory and / or ideal acquisition point position is selected from preset acquisition trajectories and / or acquisition point positions according to the target scene and / or target object.
[0182] In another implementation, preset collection trajectories and / or collection point positions are displayed to the user, and according to the user's selection operation, the ideal collection trajectory and / or ideal collection point position selected by the user among the preset collection trajectories and / or collection point positions is determined.
[0183] For example, in one embodiment, the user actively selects (for example, based on handle buttons, gesture interaction, eye tracking, etc.) the target scene and / or target object in the current shooting space, and the system virtually displays the ideal acquisition trajectory and / or acquisition point position in the space based on the selected target scene and / or target object.
[0184] For example, in another embodiment, the user actively selects a certain acquisition trajectory from the system's built-in acquisition trajectory library, and the system virtualizes the trajectory route in space based on the ideal acquisition trajectory actively selected by the user.
[0185] FIG11 is a schematic diagram showing a recommended collection method according to an embodiment of the present application.
[0186] In one embodiment, the screen shown in FIG. 11 is displayed to indicate different recommendation collection methods to the user and request the user to select a recommendation collection method.
[0187] As shown in FIG11 , based on common types of daily photography, user acquisition trajectory types are provided: shooting around the object (surrounding acquisition), shooting facing the object (front-facing acquisition), and hemispherical shooting (hemispherical acquisition).
[0188] 900 illustrates a capture trajectory for surround capture. 902 represents the subject, and 903 represents the capture trajectory. As shown in 900, capture trajectory 903 surrounds subject 902. 901 illustrates a top-down view of surround capture. The arrow in 901 indicates the orientation of the camera lens. As shown in 901, the camera lens is facing subject 902.
[0189] 910 illustrates a capture trajectory for frontal capture. 912 represents the subject, and 913 represents the capture trajectory. As shown in 910, capture trajectory 913 covers the front of subject 912. 911 illustrates a top view of frontal capture. The arrow in 911 indicates the direction of the camera lens. As shown in 911, the camera lens is facing the front of subject 912.
[0190] 920 is a schematic diagram of the acquisition trajectory for hemispherical acquisition. 922 is the subject, and 923 is the capture trajectory. As shown in 920, capture trajectory 903 surrounds subject 922, covering all sides and the top of subject 922, forming a hemisphere encompassing subject 922. 921 is a top view of hemispherical acquisition. The arrow in 901 indicates the orientation of the camera lens. As shown in 921, the camera lens is positioned to form a hemisphere encompassing subject 922, with the camera lens facing towards subject 902 at the center of the hemisphere.
[0191] S712: Capture data is collected for the target scene and / or target object. The captured data includes but is not limited to 2D video and / or image and IMU information.
[0192] For example, if you use VR glasses to capture a target scene, the VR glasses' multi-module camera, inertial measurement unit, and depth sensor are used during the shooting process to record camera motion and depth information while capturing 2D video. If you use a mobile phone to capture a target scene, the phone's camera module and IMU are used. If the phone has a depth sensor, both are used simultaneously. If not, only the relative motion information between the 2D video and the camera is recorded.
[0193] Register and fuse multi-sensor information.
[0194] For example, synchronize the captured information of multi-camera 2D image frame sequences, IMU, Lidar or ToF devices, and obtain the spatial pose information and depth information corresponding to each RGB image.
[0195] S720 , based on the ideal acquisition trajectory and / or acquisition point positions, and according to the acquired trajectory and / or acquisition point positions of the image acquisition device, guide the user to perform data acquisition.
[0196] By comparing the recommended acquisition method with the acquisition trajectory and / or acquisition point positions of the image acquisition device, the currently missing acquisition trajectory and / or recommended acquisition point positions compared to the recommended acquisition method are determined, and the user is guided to perform 2D video acquisition and / or 2D image acquisition based on the currently missing acquisition trajectory and / or recommended acquisition point positions.
[0197] Based on the real-time feedback data from the IMU, the user's current moving position is calculated, the collected trajectory and / or position is confirmed, and the uncollected trajectory and / or position are calculated, and the user is prompted on the UI interface of the VR glasses to indicate the areas that need to be re-collected.
[0198] Specifically, in one embodiment, the user is guided to capture data through text, for example, "Please move left and right", "Please change the shooting angle", etc.
[0199] In another embodiment, when the user is performing 2D video capture and / or 2D image capture, completed and to-be-completed capture trajectories and / or capture point positions are displayed through virtual markers.
[0200] For example, setting a virtual camera, such as placing a virtual camera pattern in a virtual space, marking completed virtual camera position points, and guiding the user to move to virtual camera positions marked as unfinished.
[0201] For another example, a virtual collection area is displayed, and areas where collection has been completed and areas where collection has not been completed are marked in the virtual collection area, guiding the user to perform supplementary data collection in the uncollected areas.
[0202] For another example, a virtual collection track is set in the virtual space, marking the track that has been walked and the track that has not been walked, and guiding the user to complete data collection by following the arrows on the virtual collection track.
[0203] In another embodiment, the collection of captured data is guided based on the integrity of the captured data. Specifically, the collection result after the capture data collection is performed is verified to ensure that the user has collected enough data.
[0204] FIG12 is a flowchart showing a capture data collection process according to an embodiment of the present application.
[0205] The electronic device executes the following process shown in FIG12 .
[0206] S910: Capture data is collected for the target scene and / or target object. The captured data includes but is not limited to 2D video and / or image and IMU information.
[0207] For details, please refer to the implementation of S712.
[0208] S920: Calculate the acquisition completeness based on the acquired capture data, including but not limited to judging by the coverage of the acquired positions and the completeness of the spatial geometric structure.
[0209] The collected position or scene geometry is calculated based on the collected data to determine the coverage of the collected data. If the collected data is fully covered, the user is prompted to terminate the collection. If the coverage does not meet the threshold requirement, the user is prompted to continue collecting.
[0210] In one embodiment, the coverage of the collected positions is analyzed, for example, to determine whether the number of collected positions reaches a threshold.
[0211] In another embodiment, the scene spatial geometry (eg, scene point cloud, mesh, etc.) is restored, and the data acquisition integrity is analyzed by determining the spatial geometry. Restoring the scene spatial structure requires the use of scene reconstruction related technologies.
[0212] For example, in one embodiment, based on the collected IMU data and Lidar data, the scene space point cloud structure is restored through the laser SLAM algorithm.
[0213] For example, in one embodiment, based on the collected image data and IMU data, the scene space point cloud structure is restored through a visual SLAM algorithm.
[0214] For example, in one embodiment, based on the collected image data, IMU data, and Lidar data, the scene space point cloud structure is restored through a laser-vision SLAM algorithm.
[0215] For example, in one embodiment, three-dimensional modeling is performed based on the collected 2D video and / or 2D image, spatial pose information, and depth information.
[0216] S930: Check the integrity of the calculation result of S920 and guide the user to collect data.
[0217] Specifically, in one embodiment, in S930, the scene coverage is confirmed based on the collected locations calculated in S920. For example, if the number of collected locations is small, the user is instructed to continue collecting until the number of collected locations reaches a threshold, at which point the user is prompted to terminate the collection.
[0218] In another embodiment, in S930, the spatial geometric structure restored in S920 is displayed in the UI of the image acquisition device. The user can view the scene geometric structure to confirm which parts are correct and need to be further acquired and improved, and which parts are complete and can be stopped.
[0219] FIG13 is a schematic diagram showing the geometric structure of a scene point cloud according to an embodiment of the present application.
[0220] The UI display of the image acquisition device is shown in Figure 13. Dashed line 1001 is a virtual trajectory corresponding to the acquisition trajectory of the image acquisition device. 1002, 1003, 1004, and 1005 are four virtual acquisition points on virtual trajectory 1001 (the triangle on the acquisition point represents the acquisition angle of that acquisition point), which correspond to the four real acquisition points of the image acquisition device.
[0221] Capture data is collected at the real collection points corresponding to 1002, 1003, 1004 and 1005 respectively, and the scene point cloud structure is restored based on the obtained capture data (2D video and / or 2D image, spatial pose information and depth information). The restored scene point cloud structure is shown as 1010 in Figure 11.
[0222] The user confirms the missing state of the scene point cloud structure based on the scene point cloud structure shown in FIG13 , and performs supplementary capture data collection.
[0223] FIG14 is a schematic diagram showing a scene point cloud structure according to an embodiment of the present application.
[0224] After the scene point cloud structure shown in Figure 13, the user performs additional capture data. After the additional capture data is collected, the UI interface display of the image acquisition device is shown in Figure 14. Dashed line 1101 is a virtual track corresponding to the capture track of the image acquisition device (corresponding to 1001). Virtual capture points 1102, 1103, 1104, and 1105 correspond to virtual capture points 1002, 1003, 1004, and 1005.
[0225] Virtual collection points 1106 , 1107 , 1108 , 1109 , 1110 , 1111 , and 1112 correspond to seven real collection points additionally collected by the user on the collection trajectory of the image collection device.
[0226] After the scene point cloud structure shown in FIG13 , the user performs supplementary capture data collection at the seven real acquisition points corresponding to 1106 , 1107 , 1108 , 1109 , 1110 , 1111 and 1112 , and restores the scene point cloud structure based on the obtained capture data (2D video and / or 2D image, spatial pose information and depth information). The restored scene point cloud structure is shown in 1120 of FIG14 .
[0227] According to the method of one embodiment of the present application, during the data acquisition stage, a device including but not limited to a camera module, an IMU, a depth sensor, etc. is used to shoot the target scene, analyze the data collected by the sensor in real time, judge the integrity of the target scene acquisition, and guide the user to shoot the area where the acquisition does not reach the integrity threshold, so as to obtain complete captured data.
[0228] According to the method of one embodiment of the present application, the integrity of the generated 3D video can be ensured, allowing users to immersively watch different angles of the 3D video content, providing a stronger 3D sense and richer 3D video content, and improving the user's 3D video viewing experience.
[0229] In one embodiment, in S500 , 3D video material data for the target scene and / or target object is generated based on one or more 2D videos containing the target scene and / or target object.
[0230] FIG. 15 is a flowchart showing 3D video material data generation according to an embodiment of the present application.
[0231] The electronic device executes the following process shown in FIG. 15 to implement S500 .
[0232] S610: Acquire one or more 2D videos and / or images containing a target scene and / or a target object.
[0233] Specifically, in one embodiment, one or more 2D videos may be 2D videos and / or videos shot by any device (e.g., VR glasses, mobile phones, or cameras) and any user at any time (the same time or different times).
[0234] That is, 2D materials shot by different users, different devices, and at different times can be used as material data for generating the same 3D video.
[0235] S620: Identify video frames and / or images related to the target scene and / or target object in one or more 2D videos and / or images.
[0236] The 2D material obtained in S610 contains content that is irrelevant to the key scene (main scene). Therefore, in S620, the video frames and / or images of the key scene are identified in the 2D material obtained in S610, so that the video frames and / or images of the key scene are used as material data for subsequent generation of 3D video.
[0237] Methods for identifying key scenarios include, but are not limited to, user-specified methods, automated detection, and other methods.
[0238] For example, in one embodiment, a user specifies a key scene (including but not limited to manually selecting a target scene or object in the 2D material obtained in S610, or providing a description of the key scene), and identifies video frames and / or images matching the key scene in the 2D material obtained in S610.
[0239] For example, in another embodiment, the 2D video and / or image acquired by automatic identification detection S610 (based on algorithms such as target recognition) is determined to be a key scene when the number of occurrences of a certain scene exceeds a certain threshold, and the relevant video frames and / or images of the key scene are saved.
[0240] S630 : Generate 3D video material data based on video frames and / or images related to the target scene and / or target object.
[0241] The 3D video material data includes but is not limited to the key scene video frames and / or images acquired in S620 .
[0242] Specifically, in one embodiment, in S630, the key scene video frames and / or images identified in S620 are used through a scene reconstruction algorithm (such as SLAM, colmap, etc.) to obtain the corresponding poses and scene space point clouds of the key scene video frames and / or images, and use them together with the key scene video frames and / or images as 3D video material data.
[0243] In another embodiment, depth information synchronously recorded by the capture device during the 2D video shooting process can be obtained in S610. In S630, the depth information can be used as a priori for the scene reconstruction algorithm or directly used as 3D video material data together with the key scene video frames and / or images.
[0244] Optionally, in one embodiment, in S520 , the video content of the 3D video is supplemented and expanded.
[0245] Specifically, in one embodiment, the video content of a 3D video is supplemented and expanded through artificial intelligence (AI) generation. Through intelligent recognition, AI is used to generate and supplement missing 3D video content, providing users with a 3D video viewing experience beyond the actual shooting range. For example, when the scene is identified as a meadow, the scene is intelligently expanded to an infinitely distant meadow; when the scene is identified as a classroom, AI is used to generate the school corridor, playground, and other environments outside the classroom; when the scene is identified as a scenic spot, as the user moves around, 3D content beyond the shooting range is generated using pre-stored materials of the scenic spot.
[0246] According to the method of an embodiment of the present application, the video content of a 3D video is supplemented and expanded, which can ensure the integrity of the generated 3D video and improve the user's 3D video viewing experience.
[0247] Optionally, in one embodiment, the observable area of the 3D video virtual space is calculated in S520, and the best viewing route and area for the user is recommended based on the observable area of the 3D video virtual space to prevent the user's line of sight / viewing angle from moving to an area outside the 3D virtual video space, thereby improving the user's 3D video viewing experience.
[0248] Specifically, in one embodiment, the depth point cloud structure of the scene during the data acquisition process or the 3D point cloud structure derived from the new perspective generation model (the new perspective generation model called by S520) is obtained to determine the density of the point cloud. When the point cloud density exceeds a certain threshold, it means that this area can be viewed from any angle, avoiding areas with lower point cloud density.
[0249] In another embodiment, explicit geometric structures, such as point clouds, meshes, and voxels, are derived from the new perspective generation model (the new perspective generation model invoked in S520 ). Observable regions are determined based on the integrity of the model's explicit geometry, ensuring that the geometry within the observable region is complete and dense. For example, after the scene mesh is derived from the new perspective generation model invoked in S520 , it is run through a mesh integrity check algorithm, and regions with complete meshes are set as observable regions, while regions with incomplete meshes are set as unobservable regions.
[0250] Optionally, in one embodiment, in S530, the captured trajectory is directly used as the viewing path in the 3D video virtual space. A virtual indicator icon is placed in the UI to provide the user with the optimal viewing trajectory. The interpolated and smoothed captured trajectory is used as the recommended observation trajectory, and the recommended observable range is expanded in 6DoF directions based on each captured point.
[0251] According to the method of one embodiment of the present application, the captured trajectory is used as the viewing path of the 3D video virtual space, which can effectively prevent the user's line of sight / viewing angle from moving to an area outside the 3D virtual video space, thereby improving the user's 3D video viewing experience.
[0252] Optionally, in one embodiment, in S530, when playing 3D video data, the playing environment is detected; the 3D video virtual space corresponding to the 3D video data is matched with the playing environment to obtain a matching result; and the playing of the 3D video data is guided according to the matching result.
[0253] Specifically, in one embodiment, based on the matching result between the 3D video virtual space and the playback environment, the spatial position correspondence between the playback environment and the 3D video virtual space is determined; based on the spatial position correspondence, an active area is set in the 3D video virtual space, and / or the position of the viewing point in the 3D video virtual space is determined.
[0254] For example, the system detects the playback environment, focusing on the floor, walls, and typical objects that exist in both the real playback environment and the virtual 3D video. For example, the system aligns the floor of the virtual 3D video space with the floor of the real playback environment, and aligns the walls in the virtual 3D video space with the walls in the real physical environment. An electronic fence is set up to mark inaccessible areas in the real playback environment outside the fence in the 3D video.
[0255] Specifically, the new perspective generation model is used to derive explicit spatial structure information, such as meshes, point clouds, and voxels, and use this explicit spatial structure as the virtual spatial environment for the 3D video. The MR glasses' cameras and depth sensors are then used to detect the structure of the current viewing environment. Object detection, semantic segmentation, and other methods are used to identify the geometric structure of the viewing environment and use it as the real viewing environment for the 3D video.
[0256] According to a method of an embodiment of the present application, based on the matching result between the 3D video virtual space and the playback environment, the playback of 3D video data is guided and a movable area is set in the 3D video virtual space. This can effectively prevent users from colliding with real objects in the playback environment when changing the viewing angle or viewing position while watching the 3D video.
[0257] For example, the system compares the virtual environment with the actual viewing environment, identifies objects that are identical in the virtual and real environments, and then overlaps objects of the same type. For example, if the 3D video content shows a family camping on the grass, and the actual viewing environment is the living room at home, the virtual grass plane can be aligned with the living room floor. For example, if the 3D video content shows an exciting basketball game, and the user is actually watching from a sofa, the audience seats in the 3D video can be aligned with the actual sofa position.
[0258] According to a method according to an embodiment of the present application, based on the matching result between the 3D video virtual space and the playback environment, the playback of 3D video data is guided and the position of the viewing point in the 3D video virtual space is determined. This can integrate the 3D video virtual space with the playback environment, thereby improving the user's viewing experience of 3D videos.
[0259] For example, in one application scenario, based on the 3D video generation device shown in FIG6 , the method flow shown in FIG7 is implemented, and the specific implementation process is as follows.
[0260] The acquisition unit 601 uses VR glasses or a mobile phone to call an available camera module to capture 2D video sequences or images. If the acquisition device used is VR glasses, the inertial measurement unit and depth sensor of the VR glasses are called at the same time to synchronously record the camera relative motion information and depth information corresponding to the 2D video frame sequence or image. If the acquisition device used is a mobile phone, the inertial measurement unit on the mobile phone is called at the same time to obtain the camera motion information. If the mobile phone device has a depth sensor such as Lidar or ToF, it is called at the same time. The acquisition unit 601 analyzes the integrity of the collected data in real time and provides guided acquisition instructions on the UI interface of the acquisition device.
[0261] The calculation unit 602 analyzes the scene type based on the collected data, calls the corresponding new viewing angle generation algorithm to generate 3D video data, and calculates the observable area.
[0262] The display unit 604 uses VR glasses to play 3D video data, analyzes the consistency between the real playback physical environment and the 3D video virtual environment, integrates the 3D video virtual environment with the real playback physical environment, tracks the user's observation viewpoint changes in real time, and generates the corresponding 3D video picture under the target perspective.
[0263] Optionally, in one embodiment, in addition to providing an immersive viewing experience, step S530 also allows users to edit scenes within the 3D video virtual space. For example, users can modify ambient lighting, stylize scenes, delete objects, place objects from personal or official libraries, and use AI to generate new scene content. Interaction options for editing functions include, but are not limited to, using virtual buttons within the 3D virtual space, physical controllers, gesture recognition, and eye movement recognition.
[0264] An embodiment of the present application further provides an electronic device, which is used to execute the method flow or part of the method flow described in the embodiment of the present application.
[0265] FIG16 is a schematic structural diagram of an electronic device according to an embodiment of the present application.
[0266] As shown in Figure 16, the electronic device 2500 includes a memory 2502 for storing computer program instructions and a processor 2501 for executing program instructions, wherein, when the computer program instructions are executed by the processor 2501, the electronic device 2500 is triggered to execute the method steps described in the embodiment of the present application.
[0267] Specifically, in one embodiment of the present application, the above-mentioned one or more computer programs are stored in the above-mentioned memory 2502, and the above-mentioned one or more computer programs include instructions. When the above-mentioned instructions are executed by the above-mentioned electronic device 2500, the above-mentioned electronic device 2500 executes the method steps described in the embodiment of the present application.
[0268] It is understood that the structural description of the electronic device 2500 in the embodiment of the present application does not constitute a specific limitation on the electronic device 2500. In other embodiments of the present application, the electronic device 2500 may include other components besides the processor 2501 and the memory 2502.
[0269] The processor 2501 may be a device on a chip (SOC), and the processor 2501 may include a central processing unit (CPU), and may further include other types of processors.
[0270] The processor involved in processor 2501 may include, for example, a CPU, a DSP, a microcontroller, or a digital signal processor, and may also include a GPU, an embedded neural network processor (NPU), and an image signal processor (ISP). The processor may also include necessary hardware accelerators or logic processing hardware circuits, such as ASICs, or one or more integrated circuits for controlling the execution of the program of the technical solution of this application. In addition, the processor may have the function of operating one or more software programs, and the software programs may be stored in a storage medium.
[0271] The processor 2501 may include one or more processing units. For example, the processor may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent components or integrated into one or more processors. In some embodiments, the electronic device 2500 may also include one or more processors 2501. The controller may generate an operation control signal based on the instruction opcode and the timing signal to complete the control of instruction fetching and execution.
[0272] In some embodiments, the processor 2501 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. Among them, the USB interface is an interface that complies with the USB standard specification, and specifically can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface can be used to connect a charger to charge the electronic device, and can also be used to transmit data between the electronic device and peripheral devices.
[0273] Electronic device 2500 may also include an external memory interface for connecting an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with processor 2501 via the external memory interface to implement data storage. For example, files such as music and videos can be stored on the external memory card.
[0274] Memory 2502 may include a code storage area and a data storage area. The code storage area may store an operating system. The data storage area may store data created during the use of electronic device 2500. Furthermore, memory 2502 may include high-speed random access memory and non-volatile memory, such as one or more disk storage components, flash memory components, and universal flash storage (UFS).
[0275] The memory 2502 may be a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any computer-readable medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer.
[0276] The processor 2501 and the memory 2502 may be combined into one processing device, or more commonly, they may be independent components.
[0277] One embodiment of the present application further provides an electronic chip. The electronic chip is used to execute the method flow or part of the method flow described in the embodiment of the present application. For example, the electronic chip can be a GPU.
[0278] Specifically, the electronic chip includes a processor for executing program instructions. When the computer program instructions are executed by the processor, the electronic chip is triggered to execute the method steps described in the embodiments of the present application. The processor of the electronic chip can refer to the processor of the above-mentioned electronic device.
[0279] Optionally, the devices, apparatuses, and modules described in the embodiments of the present application may be implemented by computer chips or entities, or by products having certain functions.
[0280] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0281] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of this application.
[0282] Specifically, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer-readable storage medium is run on a computer, the computer executes the method provided in the embodiment of the present application.
[0283] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program product is run on a computer, it enables the computer to execute the method provided in the embodiment of the present application.
[0284] The embodiment description in this application is described with reference to the flow chart and / or block diagram according to the method, equipment (device) and computer program product of embodiment of the present application.It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of general-purpose computer, special-purpose computer, embedded processing machine or other programmable data processing equipment to produce a machine, so that the instruction executed by the processor of computer or other programmable data processing equipment produces the device for realizing the function specified in one flow chart flow chart or multiple flow charts and / or one block or multiple blocks of block diagram.
[0285] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0286] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0287] It should also be noted that, in the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a and b and c, where a, b, c can be single or multiple.
[0288] In the embodiments of the present application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, commodity, or apparatus comprising the element.
[0289] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0290] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiments.
[0291] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments of the present application can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0292] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, apparatuses and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0293] [Corrected 22.01.2025 in accordance with Rule 91] The above description is merely a specific embodiment of the present application. Any modifications or substitutions that a person skilled in the art could easily conceive within the technical scope disclosed in this application should be included within the scope of protection of this application. The scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A 3D video generation method, characterized in that: The method is applied to an electronic device, and includes: Acquiring 3D video material data based on the captured data, where the captured data includes 2D video and / or 2D image; Analyze 3D video scene types based on 3D video material data; According to the 3D video scene type, a new perspective generation model corresponding to the 3D video scene type is obtained; The new viewing angle generation model corresponding to the 3D video scene type is used to generate 3D video data according to the 3D video material data.
2. The method according to claim 1, characterized in that The captured data includes one or more 2D videos and / or 2D images of a target scene and / or a target object; and obtaining 3D video material data based on the captured data includes: 3D video material data for the target scene and / or the target object is generated according to one or more 2D videos and / or 2D images containing the target scene and / or the target object.
3. The method according to claim 2, characterized in that The step of generating 3D video material data for the target scene and / or the target object based on one or more 2D videos and / or 2D images containing the target scene and / or the target object comprises: identifying video frames and / or images related to the target scene and / or the target object in the one or more 2D videos and / or 2D images; The 3D video material data is generated based on the video frames and / or images related to the target scene and / or the target object.
4. The method according to claim 1, wherein The method further comprises: Collecting the captured data includes directing the collection of the captured data based on data collection completeness.
5. The method according to claim 4, characterized in that The guiding the collection of the captured data based on the data collection integrity includes: The collection of the captured data is guided according to the completeness of the collection trajectory and / or the collection point positions of the captured data.
6. The method according to claim 5, characterized in that The step of guiding the collection of the captured data according to the collection trajectory and / or the collection point position of the captured data includes: Acquire an ideal acquisition method that matches the target scene and / or target object, wherein the ideal acquisition method includes an ideal acquisition trajectory and / or an ideal acquisition point position; Compare the acquisition trajectory and / or acquisition point position of the captured data with the ideal acquisition trajectory and / or ideal acquisition point position, and guide the user to supplement the acquisition data at the missing acquisition trajectory and / or acquisition point position based on the comparison result.
7. The method according to claim 4, characterized in that The guiding the collection of the captured data based on the data collection integrity includes: According to the completeness of the collected captured data, the collection of the captured data is guided.
8. The method according to claim 7, characterized in that The step of guiding the collection of the captured data according to the integrity of the captured data includes: Calculating the collected position based on the collected capture data, and determining coverage of the target scene and / or target object by the collected capture data; According to the coverage of the target scene and / or target object by the collected capture data, the user is guided to supplement the capture data corresponding to the missing coverage part.
9. The method according to claim 7, characterized in that The step of guiding the collection of the captured data according to the integrity of the captured data includes: Calculating a spatial geometric structure corresponding to the collected capture data based on the collected capture data; According to the completeness of the spatial geometric structure corresponding to the captured data that has been collected, the user is guided to supplement the collection of captured data corresponding to the missing spatial geometric structure.
10. The method according to any one of claims 1 to 9, characterized in that The method further comprises: An observable area of a 3D video virtual space corresponding to the 3D video data is calculated.
11. The method according to any one of claims 1 to 9, characterized in that The method further includes playing the 3D video data, wherein playing the 3D video data includes: When playing the 3D video data, detecting a playback environment; Matching the 3D video virtual space corresponding to the 3D video data with the playback environment to obtain a matching result; According to the matching result, the 3D video data is guided to be played.
12. The method according to claim 11, characterized in that The guiding the playing of the 3D video data includes: Determining a spatial position correspondence between the playback environment and the 3D video virtual space; According to the spatial position correspondence, an active area is set in the 3D video virtual space, and / or the position of the viewing point in the 3D video virtual space is determined.
13. An electronic device, characterized in that: The electronic device comprises a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to perform the method steps according to any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Method and system to convert 2d video into 3d video
CN101563935A
Method and system for converting two-dimensional video of complex scene into three-dimensional video
CN101917636A
Photographing method and electronic device
CN106605403A
Information generation method and device, computer equipment and storage medium
CN113141498A
Generation of stereo image sequence from 2D image sequence
CN1582457A
Cited By
Gesture detection method and device based on application scene and storage medium
CN121214500A
Gesture detection method and device based on application scenario, and storage medium
CN121214500B