A method and apparatus for three-dimensional light field generation based on a visual enclosure
By using a three-dimensional light field generation method based on a visual shell, and combining multiple RGB cameras and a virtual camera array with light field multi-viewpoint encoding and a visual shell algorithm, the problem of long generation time in three-dimensional light field display technology is solved, and a fast and realistic three-dimensional light field display is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2022-11-28
- Publication Date
- 2026-04-17
AI Technical Summary
Existing 3D light field display technologies struggle to rapidly generate realistic scenes. Traditional methods are time-consuming, costly, or lack robustness, failing to meet the real-time display requirements of dynamic scenes.
A three-dimensional light field generation method based on a visual shell is adopted. Continuous synchronous video streams are acquired by multiple RGB cameras. Combined with a virtual camera array and multi-view coding of the light field, the visual shell algorithm is used to determine the intersection and non-intersection points of the rendering light rays with the three-dimensional plane, generate a composite map and display the three-dimensional light field.
It enables rapid generation of 3D light fields, displays content that closely resembles reality, exhibits good transferability and robustness, saves time and computational costs, and supports flexible display of dynamic scenes.
Smart Images

Figure CN115841539B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional imaging technology, and in particular to a method and apparatus for generating three-dimensional light fields based on a visible shell. Background Technology
[0002] In recent years, with the development of three-dimensional light field display technology, it has been widely used in education, medical care, national defense, augmented reality and other fields because it does not require wearing additional equipment and can provide viewers with a wide viewing angle, high resolution and rich details. It has attracted widespread attention from researchers at home and abroad.
[0003] However, methods for 3D light field display still struggle to achieve real-time dynamic display of realistic scenes. Specifically, traditional light field content generation methods fall into three categories: 3D reconstruction and multi-view rendering of real scenes, multi-view rendering of simulation models, and dense viewpoint generation. The first type of method, which involves 3D reconstruction, multi-view rendering, and composite image calculation of acquired multi-view images, can achieve 3D display of realistic scenes. However, modeling and rendering operations are time-consuming and cannot meet the demand for rapid 3D light field generation. The second type of method can achieve real-time 3D display of static scenes, but its simulation models are pre-made, lacking realism. Furthermore, for dynamic scenes, model creation is required for each frame, resulting in high costs and time consumption, also failing to meet the demand for rapid 3D light field generation. The third type of method is based on deep learning for dense viewpoint generation, often requiring a training process. It is highly sensitive to changes in the environment and characters, lacking robustness and stability, and thus cannot meet the demand for rapid 3D light field generation. Summary of the Invention
[0004] This invention provides a method and apparatus for generating three-dimensional light fields based on a visible shell, which solves the problem of long generation time for three-dimensional light fields in real scenes in the prior art and realizes rapid generation of three-dimensional light fields.
[0005] This invention provides a method for generating a 3D light field based on a visual shell, comprising: acquiring a continuous synchronous video stream of a target model to be reconstructed in a 3D scene from multiple RGB cameras, wherein the continuous synchronous video stream includes multiple synchronous video frames; determining the arrangement of a virtual camera array based on parameters of a 3D light field display device, wherein the virtual camera array includes multiple virtual cameras; determining all sub-pixels required for generating a composite image and their corresponding virtual camera information using a light field multi-viewpoint encoding method based on the arrangement of the virtual camera array and the parameters of the 3D light field display device; controlling the corresponding multiple virtual cameras to emit rendering rays based on all sub-pixels required for generating the composite image and their corresponding virtual camera information; determining the 3D intersection points and non-intersection points between the rendering rays and the target model to be reconstructed based on a visual shell algorithm and the multiple synchronous video frames, and determining the color and coordinates of the intersection points and the background color of the non-intersection points; generating a composite image based on the color and coordinates of the 3D intersection points and the background color of the non-intersection points; and generating a 3D light field based on the composite image.
[0006] In some embodiments, obtaining a continuous synchronous video stream of the target model to be reconstructed in a 3D scene based on the acquisition of multiple RGB cameras includes: synchronously acquiring and saving a video stream of the continuous motion of the target model to be reconstructed in a 3D scene based on the multiple RGB cameras, and performing distortion correction, color correction and instance segmentation to obtain a continuous video stream of the target model to be reconstructed, wherein the multiple RGB cameras are calibrated in the 3D scene before synchronous acquisition.
[0007] In some embodiments, the arrangement of the virtual camera array includes the number, position, and orientation of the virtual cameras.
[0008] In some embodiments, determining all sub-pixels and their corresponding virtual camera information required to generate a composite image using a light field multi-viewpoint coding method based on the arrangement of the virtual camera array and the parameters of the three-dimensional light field display device includes: determining the viewing angle range of the composite image based on the maximum viewing angle range of the three-dimensional light field display device; and determining the sub-pixels and their corresponding virtual camera information required to generate the composite image based on the arrangement of the virtual camera array and the viewing angle range of the composite image.
[0009] In some embodiments, controlling the corresponding plurality of virtual cameras to emit rendering rays based on all sub-pixels required for generating the composite image and their corresponding virtual camera information includes: controlling the virtual camera corresponding to the virtual camera number to emit corresponding rendering rays according to the virtual camera number corresponding to the sub-pixel required for generating the composite image.
[0010] In some embodiments, determining the 3D intersection and non-intersection points of the rendered ray and the target model to be reconstructed based on the visual shell algorithm and the multi-channel synchronous video frames, and determining the color and coordinates of the intersection points and the background color of the non-intersection points, includes: projecting the coordinates of the ray tip during the forward search of the target model surface to be reconstructed onto the multi-channel synchronous video frames, and calculating the coordinates of all 2D projection points; determining whether all the 2D projection point coordinates are within the contour map of the target model to be reconstructed during the forward search; if all the 2D projection point coordinates are within the contour map of the target model to be reconstructed, then determining that the coordinates of the ray tip have reached the target model to be reconstructed. The ray is directed to the surface of the target model and intersects with it. The coordinates of the ray's tip are determined as the coordinates of the 3D intersection point. The color of the 3D intersection point is obtained based on the single video frame corresponding to the 2D projection point. If the coordinates of the 2D projection point are not within the outline of the target model to be reconstructed, it is determined that the coordinates of the ray's tip have not yet reached the surface of the target model to be reconstructed. The rendering ray continues to search forward. If the coordinates of the ray's tip exceed the 3D scene range of the target model to be reconstructed, the forward search stops. It is determined that this rendering ray does not intersect with the target model to be reconstructed, and the coordinates of the ray's tip are determined as a non-intersection point. The color of the non-intersection point is determined as the background color.
[0011] This invention also provides a three-dimensional light field generation device based on a visual shell, comprising: multiple RGB cameras for synchronously acquiring a continuous synchronous video stream of a target model to be reconstructed in a three-dimensional scene, the continuous synchronous video stream including multiple synchronous video frames; an optical computing device for determining the arrangement of a virtual camera array based on parameters of a three-dimensional light field display device, the virtual camera array including multiple virtual cameras; further configured to determine all sub-pixels required for generating a composite image and their corresponding virtual camera information using a light field multi-viewpoint encoding method based on the arrangement of the virtual camera array and the parameters of the three-dimensional light field display device; further configured to control the corresponding multiple virtual cameras to emit rendering rays based on all sub-pixels required for generating the composite image and their corresponding virtual camera information; further configured to determine the three-dimensional intersection points and non-intersection points of the rendering rays and the target model to be reconstructed based on a visual shell algorithm and the multiple synchronous video frames, and determine the color, coordinates of the intersection points and the background color of the non-intersection points; further configured to generate a composite image based on the color, coordinates of the three-dimensional intersection points and the background color of the non-intersection points; and a three-dimensional light field display device for generating a three-dimensional light field based on the composite image.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the above-described methods for generating a three-dimensional light field based on a visible shell.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the three-dimensional light field generation method based on a visual shell as described above.
[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for generating a three-dimensional light field based on a visual shell.
[0015] This invention provides a method and apparatus for generating three-dimensional light fields based on a visual shell. It determines the necessary sub-pixel information in the composite image through a multi-viewpoint encoding method for the light field, and then reconstructs the necessary three-dimensional model points and colors corresponding to the sub-pixels in the composite image based on the visual shell algorithm. This completes the calculation of the necessary sub-pixel information in the composite image, improving the calculation speed of the composite image, achieving real-time display, and making the displayed content more realistic. It also exhibits good transferability and robustness, improving the perception and utilization of three-dimensional information in three-dimensional light field display, making light field display for dynamic scenes more flexible, and saving time and computational costs. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a three-dimensional light field generation method based on a visible shell provided by the present invention;
[0018] Figure 2 This is a schematic diagram of the arrangement of virtual camera arrays provided in some embodiments of the present invention;
[0019] Figure 3 A schematic diagram of the GPU multi-threaded computing process provided by the present invention.
[0020] Figure 4 A schematic diagram of a three-dimensional light field generation device based on a visible shell provided by the present invention.
[0021] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] The following is combined Figures 1-3 This invention describes a method for generating a three-dimensional light field based on a visible shell.
[0024] Figure 1 This is one of the flowcharts of a three-dimensional light field generation method based on a visible shell provided by the present invention, including steps S101 to S107:
[0025] Step S101: A continuous synchronous video stream of the target model to be reconstructed in the three-dimensional scene is obtained based on the acquisition of multiple RGB cameras. The continuous synchronous video stream includes multiple synchronous video frames.
[0026] A multi-channel RGB camera acquisition device was set up in a real-world 3D scene. The multi-channel RGB cameras were calibrated, and the cameras synchronously acquired and saved a video stream of continuous motion of the target model to be reconstructed within the 3D scene. After distortion correction, color correction, and instance segmentation, a continuous synchronous video stream was obtained. The target model to be reconstructed is a 3D model.
[0027] In some embodiments, a multi-channel RGB camera synchronous acquisition array can be built to achieve synchronous acquisition of light fields with a wide field of view and multiple viewpoints in three-dimensional space. The multi-channel RGB camera array does not force the adjustment of the position and orientation of each camera, but follows an arbitrary distribution.
[0028] In some embodiments, the OBS (Open Broadcaster Software) open-source software can be used to acquire and locally save multiple synchronous video streams of a continuously moving target model to be reconstructed. Incremental camera calibration methods can also be used to calibrate the intrinsic and extrinsic parameters of multiple cameras and correct distortion.
[0029] In some embodiments, color correction can be performed on the acquired video stream based on a color calibration chart;
[0030] In some embodiments, instance segmentation can be performed on the target model to be reconstructed in each frame of each video stream based on an instance segmentation algorithm; instance segmentation algorithms include deep learning algorithms, green screen keying methods, etc.
[0031] Step S102: Determine the arrangement of the virtual camera array based on the parameters of the three-dimensional light field display device. The virtual camera array includes multiple virtual cameras.
[0032] Three-dimensional light field display devices are used to display three-dimensional light fields. Different three-dimensional light field display devices have different parameters, such as different display viewing angles, closest viewing distances, farthest viewing distances, and optimal viewing distances. In some embodiments, the arrangement of the virtual camera array can be determined based on parameters such as the number of viewpoints and the optimal viewing distance of the three-dimensional light field display device.
[0033] In some embodiments, the arrangement of the virtual camera array includes the number, position, and orientation of the virtual cameras. For example, the virtual cameras may be oriented in a parallel, convergent, or off-axis configuration. As an example, Figure 2 This is a schematic diagram illustrating the arrangement of virtual camera arrays according to some embodiments of the present invention. For example... Figure 2 As shown, for the position of each virtual camera, the vertical distance d between the virtual camera array and the center of the target model to be reconstructed and the distance l between every two virtual cameras can be set according to the optimal viewing distance of the 3D light field display device and the viewing angle selected by the user.
[0034] In some embodiments, the conditions to be met when arranging virtual cameras include that the overall viewing angle range of the virtual camera array is smaller than the overall acquisition viewing angle range of the RGB camera array, the total rendering viewing angle range of the virtual camera array is within the viewing angle range provided by the three-dimensional light field display device, and the position of the virtual camera array is between the nearest viewing distance and the farthest viewing distance provided by the three-dimensional light field display device.
[0035] Step S103: Based on the arrangement of the virtual camera array and the parameters of the three-dimensional light field display device, the light field multi-viewpoint coding method is used to determine all sub-pixels required to generate the composite image and their corresponding virtual camera information.
[0036] The sub-pixels required for the composite image and their corresponding virtual camera information represent each sub-pixel needed for the composite image, and which virtual camera acquired each sub-pixel. For example, the corresponding virtual camera information can be a sub-image index. As an example, for a composite image with a resolution of 4K, it requires 3840 x 2160 x 3 sub-pixels, where the sub-image index of a certain sub-pixel is 10, indicating that the color of this sub-pixel comes from virtual camera number 10. Each sub-pixel required for the composite image and its corresponding virtual camera information can be determined using a light field multi-view coding method.
[0037] By employing a light field multi-viewpoint encoding method, multiple sub-pixels and their corresponding virtual camera information required for the composite image are determined, thus avoiding the complete target model reconstruction task and allowing only the necessary multi-viewpoint information to be rendered in the composite image. The composite image obtained through this method avoids the computation of redundant information. For example, to generate a 4K resolution composite image, traditional methods require rendering multiple 4K sub-images from different viewpoints after completing the 3D reconstruction of the target model, and then synthesizing a single 4K composite image. The complete 3D reconstruction of the target model contains a significant amount of redundant information. Furthermore, the resolution of the composite image is consistent with the resolution of the sub-images from multiple viewpoints, both being 4K. Therefore, only a small portion of the 4K resolution sub-pixel information for each viewpoint is utilized. For instance, if there are 60 virtual viewpoints, only 1 / 60 of the sub-pixels in the 4K sub-image for each viewpoint are used. The composite image is generated based on this 1 / 60 of the sub-pixel information corresponding to the 60 viewpoints, while the remaining 59 / 60 of redundant sub-pixel information cannot be used to generate the composite image, thus slowing down the computation speed when calculating the composite image.
[0038] This invention, from the perspective of generating composite images, first determines all the sub-pixels required for generating the composite image and their corresponding virtual camera information. Subsequently, it only calculates the required 3D model point information to obtain the sub-pixels required for generating the composite image, thereby greatly improving the efficiency of composite image generation and avoiding the waste of other redundant pixel information when calculating the composite image.
[0039] In some embodiments, the viewing angle range of the composite image is determined based on the maximum viewing angle range of the three-dimensional light field display device; based on the arrangement of the virtual camera array and the viewing angle range of the composite image, all sub-pixels required to generate the composite image and their corresponding virtual camera information are determined. For example, if the maximum viewing angle range of the three-dimensional light field display device is 60 degrees, then this invention only displays the information required to calculate the 60-degree viewing angle, that is, it calculates the multiple sub-pixel indices required to generate the composite image and their corresponding virtual camera information. By only calculating the target model to be reconstructed within the effective viewing angle range, and not calculating the complete three-dimensional model, computational efficiency is improved, and the real-time performance of the overall computation process can be guaranteed.
[0040] Step S104: Based on all the sub-pixels required to generate the composite image and their corresponding virtual camera information, control the corresponding multiple virtual cameras to emit rendering rays.
[0041] In some embodiments, based on the virtual camera number corresponding to the sub-pixel required for generating the composite image, the virtual camera corresponding to the virtual camera number is controlled to emit a corresponding rendering ray. For example, based on the viewpoint index number corresponding to the sub-pixel required for generating the composite image, the corresponding virtual camera is selected to emit a corresponding rendering ray, thereby detecting the coordinates and color of the 3D model point corresponding to the sub-pixel, i.e., the color of the sub-pixel. For example, if the sub-image index of sub-pixel number 1 is 10, indicating that the color of sub-pixel number 1 comes from virtual camera number 10, the sub-pixel number 1 and its corresponding virtual camera information required for the composite image can be represented as sub-pixel number 1 and its corresponding virtual camera number 10. Then, virtual camera number 10 is controlled to emit a rendering ray to detect the coordinates and color of the 3D model point corresponding to sub-pixel number 1, and this color is the color of sub-pixel number 1.
[0042] The sub-image index represents the viewpoint number of the sub-image displayed by the 3D light field display device at a certain viewing position, and also represents the virtual camera number corresponding to the generation of that sub-pixel. The sub-image index is calculated by the light field multi-viewpoint coding method.
[0043] Controlling multiple virtual cameras to emit rendering rays can represent multi-view rendering of the target model to be reconstructed from multiple viewpoints. Multi-view rendering is a real-time rendering method for 3D models that can generate virtual viewpoints of the target object under arbitrary observation conditions.
[0044] Step S105: Based on the visual shell algorithm and the multi-channel synchronous video frames, determine the three-dimensional intersection points and non-intersection points between the rendered light rays and the target model to be reconstructed, and determine the color and coordinates of the intersection points and the background color of the non-intersection points.
[0045] The visual hull algorithm is a real-time 3D model reconstruction method. It acquires multi-view images and extracted contour information of the target model to be reconstructed, and through mathematical calculations of perspective projection, determines the intersection convex hull of the target objects in 3D space, thereby achieving the reconstruction of the target model.
[0046] In some embodiments, the coordinates of the leading edge of the rendering ray during the forward search of the surface of the target model to be reconstructed are projected into multiple synchronous video frames, and the coordinates of all two-dimensional projection points are calculated.
[0047] Determine whether the coordinates of all the two-dimensional projection points are within the outline of the target model to be reconstructed during the forward search process;
[0048] If all the coordinates of the two-dimensional projection points are within the outline of the target model to be reconstructed, then the coordinates of the front end of the light ray are determined to reach the surface of the target model and intersect. The coordinates of the front end of the rendering light ray are determined as the coordinates of the three-dimensional intersection point. Based on the single video frame corresponding to the two-dimensional projection point, the color of the three-dimensional intersection point is obtained.
[0049] If the coordinates of the two-dimensional projection points are not all within the outline of the target model to be reconstructed, it is determined that the coordinates of the rendering ray's leading edge have not yet reached the surface of the target model to be reconstructed. The rendering ray continues to search forward. If the coordinates of the leading edge of the ray exceed the three-dimensional scene range of the target model to be reconstructed, the forward search stops, it is determined that the rendering ray does not intersect with the target model, and the leading edge of the ray is determined as a non-intersecting point. The color of the non-intersecting point is determined as the background color.
[0050] Step S106: Generate a composite image based on the color and coordinates of the three-dimensional intersection points and the background color of the non-intersection points.
[0051] Once the color of each sub-pixel in the composite image has been solved, the composite image can be generated based on the color and coordinates of the three-dimensional intersection points and the background color of the non-intersection points.
[0052] Step S107: Generate a three-dimensional light field based on the synthesized image.
[0053] A 3D light field display device generates a 3D light field based on a composite image, thus displaying one frame of an image. Through continuous calculation of the composite image, the 3D light field can generate and display multiple frames of images in real time, thereby achieving the purpose of real-time video display.
[0054] To improve computational efficiency, this 3D light field generation method uses GPU parallel computing. Each GPU thread is responsible for the computation of one virtual ray. The computation process is also accelerated in parallel using CUDA (Compute Unified Device Architecture), thereby achieving real-time processing of continuous video streams captured by RGB real cameras to generate real-time 3D light fields. Figure 3 This is a schematic diagram of the GPU multi-threaded computing process provided by the present invention. Figure 3As shown, after the GPU begins computation, it first reads data and initializes parameters, then starts the GPU kernel function for parallel computation. Through synchronous computation by multiple GPU threads, it calculates the coordinates of all model points and their corresponding textures required to generate the composite image, thus completing the calculation of the color of each sub-pixel in the composite image. The composite image is then sent to a 3D light field display device for display. Each GPU thread is responsible for calculating one virtual ray. The calculation task for each virtual ray includes: projecting a rendering ray through a virtual camera in the virtual camera array; calculating the coordinates of the 3D points where the rendering ray intersects with the 3D model to be reconstructed based on the visual shell algorithm; calculating the texture corresponding to the 3D model points; and if the rendering ray does not intersect with the 3D model to be reconstructed, assigning the corresponding non-intersecting point the background color. This achieves the reconstruction of the 3D model points corresponding to this virtual ray and the calculation of the corresponding sub-pixel colors in the composite image. Finally, by synchronously calculating the computation tasks of multiple virtual rays through multi-threaded GPU computation, the coordinates and textures of multiple 3D model points are obtained, and a composite image is generated.
[0055] Based on the same inventive concept, embodiments of the present invention provide a three-dimensional light field generation device based on a visible shell, such as... Figure 4 As shown, the three-dimensional light field generation device provided by the present invention will be described below. The three-dimensional light field generation device based on a visible shell described below can be referred to in correspondence with the three-dimensional light field generation method described above:
[0056] A three-dimensional light field generation device based on a visual shell includes: a multi-channel RGB camera 41, used to synchronously acquire a continuous synchronous video stream of a target model to be reconstructed in a three-dimensional scene, the continuous synchronous video stream including multiple synchronous video frames; an optical computing device 42, used to determine the arrangement of a virtual camera array based on parameters of a three-dimensional light field display device, the virtual camera array including multiple virtual cameras; further used to determine multiple sub-pixels and their corresponding virtual camera information required to generate a composite image using a light field multi-viewpoint encoding method based on the arrangement of the virtual camera array and the parameters of the three-dimensional light field display device; further used to control the corresponding multiple virtual cameras to emit rendering rays based on all the sub-pixels and their corresponding virtual camera information required to generate the composite image; further used to determine the three-dimensional intersection points and non-intersection points of the rendering rays and the target model to be reconstructed based on a visual shell algorithm and the multiple synchronous video frames, and determine the color, coordinates of the intersection points and the background color of the non-intersection points; further used to generate a composite image based on the color, coordinates of the three-dimensional intersection points and the background color of the non-intersection points; and a three-dimensional light field display device 43, used to generate a three-dimensional light field based on the composite image.
[0057] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a three-dimensional light field generation method. This method includes: acquiring a continuous synchronous video stream of a target model to be reconstructed in a three-dimensional scene from multiple RGB cameras, the continuous synchronous video stream including multiple synchronous video frames; determining the arrangement of a virtual camera array based on parameters of a three-dimensional light field display device, the virtual camera array including multiple virtual cameras; determining all sub-pixels required to generate a composite image and their corresponding virtual camera information using a light field multi-viewpoint encoding method based on the arrangement of the virtual camera array and the parameters of the three-dimensional light field display device; controlling the corresponding multiple virtual cameras to emit rendering rays based on all sub-pixels required to generate the composite image and their corresponding virtual camera information; determining the three-dimensional intersection points and non-intersection points of the rendering rays and the target model to be reconstructed based on a visual shell algorithm and the multiple synchronous video frames, and determining the color and coordinates of the intersection points and the background color of the non-intersection points; generating a composite image based on the color and coordinates of the three-dimensional intersection points and the background color of the non-intersection points; and generating a three-dimensional light field based on the composite image.
[0058] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0059] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the three-dimensional light field generation method provided by the above methods. The method includes: acquiring a continuous synchronous video stream of a target model to be reconstructed in a three-dimensional scene based on multiple RGB cameras, wherein the continuous synchronous video stream includes multiple synchronous video frames; determining the arrangement of a virtual camera array based on parameters of a three-dimensional light field display device, wherein the virtual camera array includes multiple virtual cameras; and determining the arrangement of the virtual camera array based on the parameters of the three-dimensional light field display device. The parameters of the three-dimensional light field display device are used to determine all sub-pixels and their corresponding virtual camera information required to generate a composite image using a light field multi-viewpoint encoding method. Based on all sub-pixels and their corresponding virtual camera information required to generate the composite image, the corresponding multiple virtual cameras are controlled to emit rendering rays. Based on the visual shell algorithm and the multi-channel synchronous video frames, the three-dimensional intersection points and non-intersection points of the rendering rays and the target model to be reconstructed are determined, and the color and coordinates of the intersection points and the background color of the non-intersection points are determined. A composite image is generated based on the color and coordinates of the three-dimensional intersection points and the background color of the non-intersection points. A three-dimensional light field is generated based on the composite image.
[0060] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the three-dimensional light field generation method provided by the above methods. This method includes: obtaining a continuous synchronous video stream of a target model to be reconstructed in a three-dimensional scene based on multiple RGB cameras, the continuous synchronous video stream including multiple synchronous video frames; determining the arrangement of a virtual camera array based on parameters of a three-dimensional light field display device, the virtual camera array including multiple virtual cameras; determining all sub-pixels required for generating a composite image and their corresponding virtual camera information using a light field multi-viewpoint encoding method based on the arrangement of the virtual camera array and the parameters of the three-dimensional light field display device; controlling the corresponding multiple virtual cameras to emit rendering rays based on all sub-pixels required for generating the composite image and their corresponding virtual camera information; determining the three-dimensional intersection points and non-intersection points of the rendering rays and the target model to be reconstructed based on a visual shell algorithm and the multiple synchronous video frames, and determining the color and coordinates of the intersection points and the background color of the non-intersection points; generating a composite image based on the color and coordinates of the three-dimensional intersection points and the background color of the non-intersection points; and generating a three-dimensional light field based on the composite image.
[0061] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0062] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating a three-dimensional light field based on a visible shell, characterized in that, include: A continuous synchronous video stream of the target model to be reconstructed in a 3D scene is acquired by multiple RGB cameras, and the continuous synchronous video stream includes multiple synchronous video frames. The arrangement of the virtual camera array is determined based on the parameters of the three-dimensional light field display device, wherein the virtual camera array includes multiple virtual cameras; Based on the arrangement of the virtual camera array and the parameters of the three-dimensional light field display device, the light field multi-viewpoint coding method is used to determine all the sub-pixels and their corresponding virtual camera information required to generate the composite image. Based on all the sub-pixels required to generate the composite image and their corresponding virtual camera information, control the corresponding multiple virtual cameras to emit rendering rays. Based on the visual shell algorithm and the multi-channel synchronous video frames, the three-dimensional intersection points and non-intersection points of the rendered light rays and the target model to be reconstructed are determined, and the color and coordinates of the intersection points and the background color of the non-intersection points are determined. A composite image is generated based on the color and coordinates of the three-dimensional intersection points and the background color of the non-intersection points; A three-dimensional light field is generated based on the synthesized image; The method based on the visual shell algorithm and the multi-channel synchronized video frames determines the 3D intersection points and non-intersection points between the rendered light rays and the target model to be reconstructed, and determines the color and coordinates of the intersection points and the background color of the non-intersection points, including: The coordinates of the leading edge of the rendering ray during its forward search of the surface of the target model to be reconstructed are projected onto multiple synchronous video frames, and the coordinates of all two-dimensional projection points are calculated. Determine whether the coordinates of all the two-dimensional projection points are within the outline of the target model to be reconstructed during the forward search process. If all the coordinates of the two-dimensional projection points are within the outline of the target model to be reconstructed, then the coordinates of the light source tip are determined to reach the surface of the target model to be reconstructed and intersect. The coordinates of the light source tip are determined as the coordinates of the three-dimensional intersection point, and the color of the three-dimensional intersection point is obtained according to the single video frame corresponding to the two-dimensional projection point. If the coordinates of the two-dimensional projection points are not all within the outline of the target model to be reconstructed, it is determined that the coordinates of the ray tip have not yet reached the surface of the target model to be reconstructed. The rendering ray continues to search forward. If the coordinates of the ray tip exceed the range of the three-dimensional scene where the target model to be reconstructed is located, the forward search stops. It is determined that this rendering ray does not intersect with the target model to be reconstructed, and the coordinates of the ray tip are determined as non-intersecting points. The color of the non-intersecting point is determined as the background color.
2. The method for generating a three-dimensional light field based on a visible shell according to claim 1, characterized in that, The method of obtaining a continuous synchronous video stream of the target model to be reconstructed in a 3D scene based on the acquisition of multiple RGB cameras includes: synchronously acquiring and saving the video stream of the continuous motion of the target model to be reconstructed in the 3D scene based on the multiple RGB cameras, and performing distortion correction, color correction and instance segmentation to obtain a continuous video stream of the target model to be reconstructed, wherein the multiple RGB cameras are calibrated in the 3D scene before synchronous acquisition.
3. The method for generating a three-dimensional light field based on a visible shell according to claim 1, characterized in that, The arrangement of the virtual camera array includes the number, position, and orientation of the virtual cameras.
4. The method for generating a three-dimensional light field based on a visible shell according to claim 1, characterized in that, The step of determining all sub-pixels and their corresponding virtual camera information required to generate a composite image using a light field multi-viewpoint coding method based on the arrangement of the virtual camera array and the parameters of the three-dimensional light field display device includes: determining the viewing angle range of the composite image based on the maximum viewing angle range of the three-dimensional light field display device; and determining the sub-pixels and their corresponding virtual camera information required to generate the composite image based on the arrangement of the virtual camera array and the viewing angle range of the composite image.
5. The method for generating a three-dimensional light field based on a visible shell according to claim 1, characterized in that, The step of controlling the corresponding multiple virtual cameras to emit rendering rays based on all sub-pixels required for generating the composite image and their corresponding virtual camera information includes: controlling the virtual camera corresponding to the virtual camera number to emit the corresponding rendering ray according to the virtual camera number corresponding to the sub-pixel required for generating the composite image.
6. A three-dimensional light field generation device based on a visible shell, characterized in that, For performing the method as described in any one of claims 1-5, comprising: A multi-channel RGB camera is used to synchronously acquire a continuous synchronous video stream of the target model to be reconstructed in a 3D scene, wherein the continuous synchronous video stream includes multiple synchronous video frames. An optical computing device for determining the arrangement of a virtual camera array based on parameters of a three-dimensional light field display device, wherein the virtual camera array includes multiple virtual cameras; Based on the arrangement of the virtual camera array and the parameters of the three-dimensional light field display device, the light field multi-viewpoint coding method is used to determine all the sub-pixels and their corresponding virtual camera information required to generate the composite image. Based on all the sub-pixels required to generate the composite image and their corresponding virtual camera information, control the corresponding multiple virtual cameras to emit rendering rays. Based on the visual shell algorithm and the multi-channel synchronous video frames, the three-dimensional intersection points and non-intersection points of the rendered light rays and the target model to be reconstructed are determined, and the color and coordinates of the intersection points and the background color of the non-intersection points are determined. A composite image is generated based on the color and coordinates of the three-dimensional intersection points and the background color of the non-intersection points; A three-dimensional light field display device for generating a three-dimensional light field based on the synthesized image.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the three-dimensional light field generation method based on a visible shell as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the three-dimensional light field generation method based on a visual shell as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the three-dimensional light field generation method based on a visual shell as described in any one of claims 1 to 5.