Adaptive ray sampling three-dimensional light field content generation method and device

Through adaptive light sampling strategy and neural network model optimization, the problem of poor three-dimensional light field visual effects caused by inflexible sampling strategy in existing technologies is solved, and higher-quality three-dimensional light field content generation is achieved.

CN119722904BActive Publication Date: 2025-10-17BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411685002.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-10-17
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

In existing 3D light field content generation methods, the sampling strategy lacks flexibility, resulting in inaccurate distribution of sampling points in complex scenes, affecting the model's restoration ability and visual effects, and making it difficult to meet users' expectations for quality.

Method used

An adaptive light sampling strategy is adopted to establish a three-dimensional world coordinate system by acquiring multi-view images and camera calibration information, perform initial sampling and resampling, adjust the sampling method according to the scene geometry distribution, and combine neural network model optimization training to generate high-quality three-dimensional light field content.

Benefits of technology

The quality of 3D light field content generation has been improved, which can better adapt to the needs of complex scenes and enhance the visual effects and model restoration capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722904B_ABST
    Figure CN119722904B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of three-dimensional light field content generation, and provide a self-adaptive light ray sampling three-dimensional light field content generation method and device. The method uses a plurality of RGB cameras to collect images of a target three-dimensional scene, establishes a three-dimensional world coordinate system, and constructs a three-dimensional sampling space and a voxel space. Through a self-adaptive light ray sampling strategy, initial sampling is performed on the occupied voxels to form initial sampling points, and the resampling mode is adaptively adjusted according to the sufficiency of the initial sampling points. In addition, a neural network model is used for training to update the occupancy state of the voxel grid. Finally, a virtual rendering camera array is used to calculate a rendering light to generate a composite image, thereby generating three-dimensional light field content. Thus, the method flexibly adjusts the sampling strategy according to the geometric distribution of the scene, thereby improving the quality of the generated three-dimensional light field content.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of three-dimensional light field content generation, and in particular to a self-adaptive light ray sampling three-dimensional light field content generation method and device. BACKGROUND

[0002] With the rapid development of computer vision and computer graphics, its application in the field of three-dimensional light field display has attracted widespread attention. Learning-based methods can achieve efficient reconstruction of three-dimensional scenes by constructing neural network models and training with a large amount of data. However, this method still has some shortcomings.

[0003] Learning-based methods often rely on reliable light ray sampling strategies to search and determine the positions of target objects in three-dimensional space, so as to accurately represent the scene distribution. Traditional three-dimensional light field content generation methods are usually based on a defined sampling space or a fixed sampling strategy, which fails to fully consider the complexity of scene geometry. In high complexity scenes, some areas may be rich in details, while the geometric structure of other areas is relatively simple. If the sampling of these areas lacks flexibility, it may lead to sparse samples or excessive sampling points, thereby affecting the model's recovery ability in these areas. Under a fixed sampling strategy, the sampling of light rays cannot be adaptively adjusted, which may lead to inaccurate distribution of sampling points, thereby reducing the effect of light field content generation. For example, uniform sampling or voxel space-based sampling strategies may result in insufficient sampling point density in some complex areas, leading to missing image details. In open areas, excessive sampling may result in waste of computing resources.

[0004] Therefore, the existing sampling strategy not only affects the restoration ability of the neural network model for three-dimensional scenes, but also reduces the visual effect of the finally generated three-dimensional light field, making it difficult to meet users' expectations for quality in actual applications. Although some improved light ray sampling methods have been proposed in the prior art, there are still problems such as low sampling efficiency, severe detail loss, and low quality of generated images. SUMMARY

[0005] The present application provides a self-adaptive light ray sampling three-dimensional light field content generation method and device to solve the problem that the existing sampling strategy in the prior art not only affects the restoration ability of the neural network model for three-dimensional scenes, but also reduces the visual effect of the finally generated three-dimensional light field, making it difficult to meet users' expectations for quality in actual applications, and to achieve adaptive and flexible adjustment of the sampling method according to the geometric distribution and detail requirements of the scene, thereby improving the quality of three-dimensional light field content generation and better adapting to the needs of complex scenes.

[0006] The present application provides a self-adaptive light ray sampling three-dimensional light field content generation method, comprising:

[0007] acquire a plurality of original images of a target three-dimensional scene from different perspectives and camera calibration information of a multi-channel RGB camera used to acquire the plurality of original images;

[0008] establish a three-dimensional world coordinate system based on the camera calibration information of the multi-channel RGB camera, and construct a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system;

[0009] set a plurality of virtual rendering cameras in the three-dimensional world coordinate system based on parameters of a light field display device, emit a rendering light ray for calculating a composite image based on light field encoding, and each sub-pixel in the composite image corresponds to a rendering light ray;

[0010] perform initial sampling on the light rays emitted by the plurality of virtual rendering cameras in the three-dimensional voxel space based on a preset adaptive light ray sampling strategy, obtain a plurality of initial sampling points, calculate the point density of the plurality of initial sampling points, and calculate the transparency of the initial sampling points according to the point density;

[0011] adaptively adjust the resampling mode for resampling according to a set number threshold of initial sampling points;

[0012] The adaptive adjustment of the resampling mode for resampling according to the set number threshold of initial sampling points comprises:

[0013] When the number of initial sampling points reaches the number threshold, perform voxel space resampling according to the cumulative transparency distribution of the plurality of initial sampling points to obtain the coordinates and direction vectors of a plurality of resampling points;

[0014] When the number of initial sampling points does not reach the number threshold, reselect new sampling points in the range of the three-dimensional sampling space as resampling points to obtain the coordinates and direction vectors of a plurality of resampling points;

[0015] Input the coordinates and direction vectors of all resampling points on each rendering light ray into a neural network model after position encoding and direction encoding to calculate the color and point density of all resampling points on each rendering light ray, accumulate to obtain the sub-pixel color corresponding to the rendering light ray in the composite image, and generate three-dimensional light field content;

[0016] Generate a target three-dimensional light field based on the three-dimensional light field content and a three-dimensional light field display device.

[0017] In one possible implementation, the method further comprises:

[0018] construct the three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system, wherein the three-dimensional voxel space represents a spatial range in which initial sampling or resampling based on a three-dimensional voxel grid is performed, and the three-dimensional sampling space represents a spatial range in which new sampling points are selected when the number of initial sampling points does not reach the number threshold.

[0019] In one possible implementation, the method further includes:

[0020] In the neural network model optimization training process, the neural network model is optimized and trained based on pixel colors of the multi-path RGB camera and corresponding original image pixel colors in the original image, with the objective of minimizing a preset error function.

[0021] In the neural network model optimization training process, the occupancy state of the three-dimensional voxel grid is updated according to position information of the three-dimensional voxel grid in the three-dimensional voxel space.

[0022] In one possible implementation, the method further includes:

[0023] Determine a plurality of initial sampling line segments based on the plurality of initial sampling points.

[0024] Input the midpoint position of each initial sampling line segment into a neural network model to obtain a point density corresponding to each initial sampling point.

[0025] Calculate the transparency of each initial sampling point according to the point density of each initial sampling point.

[0026] Accumulate the transparency of each initial sampling point on each rendering light along the sampling light to obtain a transparency accumulation distribution of the initial sampling points on each rendering light.

[0027] Determine an initial sampling line segment area corresponding to each resampling point based on the transparency accumulation distribution of the plurality of initial sampling points.

[0028] Resample based on the initial sampling line segment area to obtain coordinates and direction vectors of a plurality of resampling points.

[0029] In one possible implementation, the method further includes:

[0030] Set the number of virtual rendering cameras according to the number of viewpoints of the three-dimensional light field display device.

[0031] Set the position of the virtual rendering cameras relative to the target three-dimensional scene according to the optimal viewing distance of the three-dimensional light field display device.

[0032] Set the spacing of the virtual rendering cameras according to the maximum viewing angle of the three-dimensional light field display device.

[0033] In one possible implementation, the method further comprises:

[0034] extracting viewpoint information corresponding to each sub-pixel from the composite image;

[0035] outputting the viewpoint information to the three-dimensional light field display device and generating target three-dimensional light field content based on the three-dimensional light field display device.

[0036] The application also provides an adaptive light ray sampling three-dimensional light field content generation device, comprising the following modules:

[0037] an acquisition module, configured to acquire multiple original images of a target three-dimensional scene under different viewpoints and camera calibration information of a multiple-channel RGB camera collecting the multiple original images;

[0038] a construction module, configured to construct a three-dimensional world coordinate system based on the camera calibration information of the multiple-channel RGB camera and construct a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system;

[0039] a setting module, configured to set multiple virtual rendering cameras in the three-dimensional world coordinate system based on parameters of a three-dimensional light field display device, emit rendering light rays for calculating a composite image based on light field encoding, and each sub-pixel in the composite image corresponds to a rendering light ray;

[0040] a sampling module, configured to perform initial sampling on the light rays emitted by the multiple virtual rendering cameras in the three-dimensional voxel space based on a preset adaptive light ray sampling strategy, obtain multiple initial sampling points, calculate point densities of the multiple initial sampling points, calculate the transparency of the initial sampling points according to the point densities, adaptively adjust a resampling mode to perform resampling according to a set number threshold of the initial sampling points, and the adaptively adjusting the resampling mode to perform resampling according to the set number threshold of the initial sampling points comprises: when the number of initial sampling points reaches the number threshold, performing voxel space resampling according to the cumulative transparency distribution of the multiple initial sampling points to obtain coordinates and direction vectors of multiple resampling points; and when the number of initial sampling points does not reach the number threshold, reselecting new sampling points in the range of the three-dimensional sampling space as resampling points to obtain coordinates and direction vectors of multiple resampling points;

[0041] a calculation module, configured to input coordinates and direction vectors of all resampling points on each rendering light ray after position encoding and direction encoding into a neural network model to calculate the color and point density of all resampling points on each rendering light ray, accumulate to obtain the color of a sub-pixel corresponding to the rendering light ray in the composite image, and generate three-dimensional light field content;

[0042] The generating module is configured to generate a target three-dimensional light field based on the three-dimensional light field content and a three-dimensional light field display device.

[0043] The application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the adaptive ray sampling three-dimensional light field content generation method according to any one of the above.

[0044] The application further provides a non-transitory computer readable storage medium, which stores a computer program executable by a processor to implement the adaptive ray sampling three-dimensional light field content generation method according to any one of the above.

[0045] The application further provides a computer program product, which includes a computer program executable by a processor to implement the adaptive ray sampling three-dimensional light field content generation method according to any one of the above.

[0046] The adaptive light sampling three-dimensional light field content generation method and device provided by the application, by acquiring multiple original images of a target three-dimensional scene under different viewing angles and collecting camera calibration information of a multi-channel RGB camera for the multiple original images, establishing a three-dimensional world coordinate system based on the camera calibration information of the multi-channel RGB camera, and constructing a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system, setting a multi-channel virtual rendering camera in the three-dimensional world coordinate system based on the parameters of a three-dimensional light field display device, emitting a rendering light for calculating a composite image based on light field coding, each sub-pixel in the composite image corresponding to a rendering light, performing initial sampling on the light emitted by the multi-channel virtual rendering camera in the three-dimensional voxel space based on a preset adaptive light sampling strategy, obtaining multiple initial sampling points, calculating the point density of the multiple initial sampling points, and calculating the transparency of the initial sampling points according to the point density; adaptively adjusting the resampling mode for resampling according to the set number threshold of the initial sampling points; the adaptive adjustment of the resampling mode for resampling according to the set number threshold of the initial sampling points comprises: when the number of initial sampling points reaches the number threshold, performing voxel space resampling according to the cumulative transparency distribution of the multiple initial sampling points to obtain the coordinates and direction vectors of multiple resampling points; when the number of initial sampling points does not reach the number threshold, newly selecting a new sampling point in the range of the three-dimensional sampling space as a resampling point to obtain the coordinates and direction vectors of multiple resampling points; inputting the position encoding and direction encoding of the coordinates and direction vectors of all resampling points on each rendering light into a neural network model to calculate the color and point density of all resampling points on each rendering light, and accumulating to obtain the sub-pixel color corresponding to the rendering light in the composite image, thereby generating three-dimensional light field content; and generating a target three-dimensional light field based on the three-dimensional light field content and a three-dimensional light field display device. Compared with the existing sampling strategy in the prior art, the network model's restoration ability for a three-dimensional scene is not only affected, but also the visual effect of the finally generated three-dimensional light field is reduced, making it difficult to meet the user's quality expectations in actual application. According to the present application, the sampling mode is flexibly adjusted according to the geometric distribution and detail requirements of the scene, thereby improving the quality of three-dimensional light field content generation and better adapting to the requirements of complex scenes. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0048] Figure 1 is one of the flowcharts of the adaptive light sampling three-dimensional light field content generation method provided by the application.

[0049] Figure 2 Figure 2 is a flowchart of a second embodiment of the adaptive light sampling three-dimensional light field content generation method provided by the present application.

[0050] Figure 3 Figure 3 is a schematic diagram of the adaptive light sampling method provided by the present application.

[0051] Figure 4 Figure 4 is a structural schematic diagram of the adaptive light sampling three-dimensional light field content generation device provided by the present application.

[0052] Figure 5 Figure 5 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0053] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0054] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0055] Figure 1 Figure 1 is a flowchart of the adaptive light sampling three-dimensional light field content generation method provided by the present application. As shown in Figure 1, the method specifically includes the following steps. Figure 1

[0056] S11, acquiring multiple original images of a target three-dimensional scene under different perspectives and camera calibration information of a multi-channel RGB camera collecting the multiple original images.

[0057] In the embodiments of the present application, the multi-channel RGB camera calibrated by the camera calibration method is used to capture the target three-dimensional scene from different perspectives to acquire multiple original image data under multiple perspectives. The camera calibration method can be Zhang Zhengyou calibration method, or open-source calibration software Colmap can be used. After the camera calibration, the multiple original images are corrected for distortion and color to ensure that the acquired images have good performance in terms of distortion removal and color consistency. The color correction of the multiple perspective images can be realized by collecting and color calibrating the 24-color standard color board in the target scene.

[0058] S12, establishing a three-dimensional world coordinate system based on the camera calibration information of the multi-channel RGB camera, and constructing a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system. ​

[0059] Based on the multi-view images collected by the multi-path RGB camera, the camera's internal and external parameters are calculated by a camera calibration method as the camera calibration information, so as to establish a three-dimensional world coordinate system. In this coordinate system, a three-dimensional sampling space and a three-dimensional voxel space are constructed. At the same time, a three-dimensional voxel grid is constructed in the three-dimensional voxel space, that is, the three-dimensional space is divided into a plurality of voxels, and each voxel records its occupancy state in the three-dimensional space.

[0060] S13, setting a plurality of virtual rendering cameras in the three-dimensional world coordinate system based on the parameters of the three-dimensional light field display device, and emitting rendering light rays for calculating a composite image based on the light field coding, each sub-pixel in the composite image corresponding to a rendering light ray.

[0061] According to the number of viewpoints of the three-dimensional light field display device, the number of virtual rendering cameras is set, according to the optimal viewing distance of the three-dimensional light field display device, the position of the virtual rendering camera from the target three-dimensional scene is set, and according to the maximum viewing angle of the three-dimensional light field display device, the spacing of the virtual rendering camera is set.

[0062] Further, according to the light field coding, rendering light rays for calculating a composite image are emitted, and each rendering light ray corresponds to each sub-pixel in the later composite image.

[0063] S14, based on a preset adaptive light sampling strategy, the light rays emitted by the plurality of virtual rendering cameras are initially sampled in the three-dimensional voxel space to obtain a plurality of initial sampling points, and the point density of the plurality of initial sampling points is calculated, and the transparency of the initial sampling points is calculated according to the point density.

[0064] The light rays emitted by the virtual rendering camera are sampled by using the adaptive light sampling strategy when passing through the three-dimensional voxel space.

[0065] The adaptive light sampling strategy first performs initial sampling in the occupied three-dimensional voxel space based on the three-dimensional voxel grid to form the left end point and the right end point of the initial sampling line segment, collectively referred to as the initial sampling point, and according to whether the actual number of initial sampling points reaches the set number threshold of initial sampling points (for example, the number threshold of initial sampling points is set to 64 or 128), the sampling method is adaptively adjusted, when the initial sampling points are sufficient (the number is greater than or equal to the number threshold of initial sampling points), the initial sampling point position is input into a neural network model to obtain an initial sampling point cumulative transparency distribution, and the voxel space is resampled based on the cumulative transparency distribution, when the initial sampling points are insufficient, sampling is performed in the three-dimensional sampling space range.

[0066] Specifically, the purpose of adaptive light sampling is to flexibly adjust the sampling strategy according to the geometric distribution of the scene. The specific implementation includes initial sampling and resampling process. First, based on the occupancy state of the three-dimensional voxel grid, initial sampling is performed in the three-dimensional voxel space to generate a series of initial sampling line segments. The left end point and the right end point of these line segments are called initial sampling points. By calculating the length and position of each initial sampling line segment, the influence range of the surrounding area is determined. Then, according to the sufficiency of the initial sampling points, the sampling method is adaptively adjusted. When the initial sampling points are sufficient, the position of these points is input into the neural network model to calculate the cumulative initial sampling point density. The transparency of the initial sampling points is calculated according to the point density of the initial sampling points. The cumulative distribution of the transparency is calculated according to the transparency of the initial sampling points, and the voxel space is resampled according to the cumulative distribution of the transparency. When the initial sampling points are insufficient, the sampling points are reselected in the three-dimensional sampling space. At this time, a method different from the initial sampling of the voxel space or the resampling of the voxel space is used to ensure that the resampling points have sufficient spatial span range in the empty area. The three-dimensional voxel space and the three-dimensional sampling space are both in the set three-dimensional world coordinate system, and the coverage areas set by them have an intersection.

[0067] S15, according to the set number threshold of the initial sampling points, adaptively adjusting the resampling method to resample.

[0068] In the three-dimensional voxel space resampling process, the area of each resampling point is determined according to the initial sampling line segment selected by the initial sampling point transparency cumulative distribution, and a front end point is defined based on the position of the initial sampling line segment where each resampling point is located. The position of the front end point is before the left end point of the initial sampling line segment, and the length from the front end point to the left end point of the initial sampling line segment is related to the length from the left end point to the right end point of the initial sampling line segment (for example, the distance is equal, or there is a related calculation relationship), and the position of the resampling point is obtained from the positions of the front end point, the left end point and the right end point, and then a plurality of resampling points are resampled.

[0069] Based on the initial sampling, the voxel space resampling is used to improve the accuracy of the light field content generation. Specifically, the area of each resampling point is determined according to the initial sampling line segment selected by the initial sampling point transparency cumulative distribution. The initial sampling line segment area where different resampling points are located may be the same or different, and the area where the resampling point is located is determined by the initial sampling point transparency cumulative distribution. A front end point is defined in front of each selected initial sampling line segment, and the position of the front end point is determined by the length relationship between the left end point and the right end point. The distance from the front end point to the left end point is used as the basis for resampling to ensure that the position of the resampling point is within a reasonable range. The distance from the front end point to the left end point can be selected to be slightly less than, equal to, or slightly greater than the distance from the left end point to the right end point. The resampling result is calculated through the position information of the front end point, the left end point and the right end point. This process will effectively limit the range of the resampling point and avoid the resampling point occupying the invalid space area.

[0070] S16, based on the position encoding and direction encoding of all the coordinates of the resampling points and the direction vectors on each rendering light, inputting the neural network model to calculate the color and point density of all the resampling points on each rendering light, and accumulating to obtain the sub-pixel color corresponding to the rendering light in the synthesized image, and generating three-dimensional light field content.

[0071] The coordinates and direction vectors of the resampling points on each rendering light emitted by the virtual rendering camera can be determined.

[0072] After encoding (for example, hash encoding, sine function encoding, etc.) the coordinates and direction vectors of each resampling point on each rendering light, the color and point density of all the resampling points are obtained by inputting the neural network model, so as to calculate the transparency of the resampling points, and the image pixel color corresponding to each rendering light can be calculated by weighted average of the color of the resampling points according to the transparency of the resampling points.

[0073] The coordinates and direction vectors of the resampling points on each light are encoded and input into the neural network model. In this way, the neural network model can learn how to generate corresponding color and point density information according to the position and direction of the sampling points during the training process.

[0074] Optionally, during the optimization and training of the neural network, the image pixel color of the image pixel corresponding to each rendering light can be calculated based on the neural network model, and compared with the original image pixel color of the corresponding image pixel coordinate in the original image, so as to calculate and minimize the error function as the target, and optimize and train the neural network model. During the optimization and training of the neural network, the occupancy state of the three-dimensional voxel grid is updated by the neural network model according to the position information of the three-dimensional voxel grid.

[0075] Specifically, during the optimization and training of the neural network, the occupancy state of the three-dimensional voxel grid needs to be dynamically updated. A certain intensity of Gaussian random noise is added to the center position of each three-dimensional voxel grid to obtain the corresponding position information. The position information after adding noise is input into the neural network to obtain the corresponding point density, and whether the voxel grid is occupied is judged according to whether the point density reaches the set point density threshold. During the training process, the occupancy state of the voxel is constantly updated according to the point density of the target scene, so as to ensure that the neural network model always accurately reflects the target scene. This method can improve the geometric distribution of the dynamically updated scene and improve the effect of adaptive ray sampling.

[0076] S17, generating a target three-dimensional light field based on the three-dimensional light field content and a three-dimensional light field display device.

[0077] In the rendering process, the multi-path virtual rendering camera is controlled to emit rendering light, and the coordinates of the resampling points are calculated using an adaptive light sampling strategy. This process will ensure that rich scene information is captured from various perspectives. Through a neural network model, a synthetic image is calculated to generate a three-dimensional light field content. The three-dimensional light field content integrates information from different perspectives, providing a basis for subsequent three-dimensional light field generation.

[0078] Further, according to the viewpoint information corresponding to each sub-pixel in the synthetic image and the color information of each sub-pixel as the three-dimensional light field content, a target three-dimensional light field is generated using a three-dimensional light field display device.

[0079] The adaptive light sampling three-dimensional light field content generation method provided by the application, by acquiring a plurality of original images of a target three-dimensional scene under different viewing angles and collecting camera calibration information of a plurality of RGB cameras for collecting the plurality of original images; based on the camera calibration information of the plurality of RGB cameras, a three-dimensional world coordinate system is established, and a three-dimensional voxel space and a three-dimensional sampling space are constructed based on the three-dimensional world coordinate system; in the three-dimensional world coordinate system, a plurality of virtual rendering cameras are set based on the parameters of a three-dimensional light field display device, a rendering light is emitted for calculating a composite image based on light field coding, and each sub-pixel in the composite image corresponds to a rendering light; based on a preset adaptive light sampling strategy, the light emitted by the plurality of virtual rendering cameras is initially sampled in the three-dimensional voxel space to obtain a plurality of initial sampling points, and the point density of the plurality of initial sampling points is calculated, and the transparency of the initial sampling points is calculated according to the point density; according to the set number threshold of the initial sampling points, the resampling mode is adaptively adjusted for resampling; the adaptive adjustment of the resampling mode according to the set number threshold of the initial sampling points includes: when the number of initial sampling points reaches the number threshold, the voxel space is resampled according to the cumulative transparency distribution of the plurality of initial sampling points to obtain the coordinates and direction vectors of a plurality of resampling points; when the number of initial sampling points does not reach the number threshold, new sampling points are selected in the range of the three-dimensional sampling space as resampling points to obtain the coordinates and direction vectors of a plurality of resampling points; based on the position encoding and direction encoding of all resampling point coordinates and direction vectors on each rendering light, the color and point density of all resampling points on each rendering light are calculated by inputting a neural network model, and the sub-pixel color corresponding to the rendering light in the composite image is obtained after accumulation, thereby generating three-dimensional light field content; and based on the three-dimensional light field content and a three-dimensional light field display device, a target three-dimensional light field is generated. Compared with the existing sampling strategy in the prior art, the restoration ability of the neural network model to the three-dimensional scene is not only limited, but also the visual effect of the finally generated three-dimensional light field is reduced, so that it is difficult to meet the user's expectation of the quality of the light field content in actual application. According to the method, the sampling mode is adaptively and flexibly adjusted according to the geometric distribution and detail requirements of the scene, thereby improving the quality of the three-dimensional light field content generation and better adapting to the requirements of complex scenes.

[0080] Figure 2 is a flowchart of the adaptive light sampling three-dimensional light field content generation method provided by the application, as shown in Figure 2 , the method specifically includes:

[0081] S21, acquiring a plurality of original images of a target three-dimensional scene under different viewing angles and collecting camera calibration information of a plurality of RGB cameras for collecting the plurality of original images.

[0082] In the embodiment of the present application, the multi-channel RGB camera is used to capture the target three-dimensional scene from different angles through camera calibration to obtain multiple original image data under multiple perspectives. The camera calibration method can be Zhang Zhengyou calibration method, or open source calibration software Colmap. After camera calibration, distortion correction and color correction are performed on multiple original images to ensure that the images have good performance in terms of distortion and color consistency. Color correction of multi-perspective images can be achieved by collecting and color calibrating 24-color standard color plates in the target scene.

[0083] S22, establishing a three-dimensional world coordinate system based on the camera calibration information of the multi-channel RGB camera, and constructing a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system.

[0084] Based on the multi-perspective images collected by the multi-channel RGB camera, the camera's intrinsic and extrinsic parameters are calculated as camera calibration information by using a camera calibration method, thereby establishing a three-dimensional world coordinate system. In this coordinate system, a three-dimensional sampling space and a three-dimensional voxel space are constructed. At the same time, a three-dimensional voxel grid is constructed in the three-dimensional voxel space, i.e. the three-dimensional space is divided into multiple voxels, and each voxel records its occupancy state in the three-dimensional space.

[0085] S23, setting multiple virtual rendering cameras in the three-dimensional world coordinate system based on the parameters of the three-dimensional light field display device, emitting rendering light rays for calculating a composite image based on light field encoding, and each sub-pixel in the composite image corresponds to a rendering light ray.

[0086] S24, based on a preset adaptive light sampling strategy, performing initial sampling on the light rays emitted by the multiple virtual rendering cameras in the three-dimensional voxel space to obtain multiple initial sampling points, calculating the point density of the multiple initial sampling points, and calculating the transparency of the initial sampling points according to the point density.

[0087] S25, determining multiple initial sampling line segments based on the multiple initial sampling points.

[0088] S26, inputting the midpoint position of each initial sampling line segment into a neural network model to obtain the point density corresponding to each initial sampling point.

[0089] S27, calculating the transparency of each initial sampling point according to the point density of each initial sampling point.

[0090] S28, accumulating the transparency of each initial sampling point on each rendering light ray along the sampling light ray to obtain the transparency accumulation distribution of the initial sampling points on each rendering light ray.

[0091] The following uniformly describes S23-S28:

[0092] The number of virtual rendering cameras is set according to the number of viewpoints of the three-dimensional light field display device, the positions of the virtual rendering cameras relative to the target three-dimensional scene are set according to the optimal viewing distance of the three-dimensional light field display device, and the spacing of the virtual rendering cameras is set according to the maximum viewing angle of the three-dimensional light field display device.

[0093] Further, rendering rays for calculating a composite image are emitted according to light field encoding, and each rendering ray corresponds to each sub-pixel in the post-composite image.

[0094] The rays emitted by the virtual rendering cameras are sampled using an adaptive ray sampling strategy when passing through the three-dimensional voxel space.

[0095] As shown in Figure 3 The adaptive ray sampling strategy first performs initial sampling in the occupied three-dimensional voxel space based on the three-dimensional voxel grid to form left endpoints and right endpoints of initial sampling line segments, collectively referred to as initial sampling points, and adaptively adjusts the sampling method according to whether the actual number of initial sampling points reaches a set initial sampling point number threshold (for example, the initial sampling point number threshold is set to 64 or 128), when the initial sampling points are sufficient (the number is greater than or equal to the initial sampling point number threshold), the initial sampling point position is input into a neural network model to obtain an initial sampling point cumulative transparency distribution, and the voxel space is resampled based on the cumulative transparency distribution, when the initial sampling points are insufficient, sampling is performed in the three-dimensional sampling space range.

[0096] Specifically, the purpose of adaptive ray sampling is to flexibly adjust the sampling strategy according to the geometric distribution of the scene. The specific implementation includes the initial sampling and resampling processes. First, initial sampling is performed in the three-dimensional voxel space based on the occupancy state of the three-dimensional voxel grid to generate a series of initial sampling line segments. The left endpoints and right endpoints of these line segments are called initial sampling points. By calculating the length and position of each initial sampling line segment, the influence range of the surrounding area is determined. Then, according to the sufficiency of the initial sampling points, the sampling method is adaptively adjusted. When the initial sampling points are sufficient, the positions of these points are input into a neural network model to obtain the cumulative initial sampling point density, the transparency of the initial sampling points is calculated according to the point density of the initial sampling points, the transparency cumulative distribution is calculated according to the transparency of the initial sampling points, and the voxel space is resampled according to the transparency cumulative distribution. When the initial sampling points are insufficient, the sampling points are reselected in the three-dimensional sampling space range. At this time, a method different from the initial sampling of the voxel space or the resampling of the voxel space is used to ensure that the resampling points have sufficient spatial span range in the empty area. The three-dimensional voxel space and the three-dimensional sampling space are both in a set three-dimensional world coordinate system, and the coverage areas of the three-dimensional voxel space and the three-dimensional sampling space have an intersection.

[0097] S29, determining the initial sampling line segment region corresponding to each resampling point based on the transparency cumulative distribution of the plurality of initial sampling points.

[0098] S210 : Resampling is performed based on the initial sampling line segment area to obtain coordinates and direction vectors of multiple resampling points.

[0099] During the three-dimensional voxel space resampling process, the area defined by the initial sampling segment where each resampling point is located is determined based on the cumulative distribution of transparency of the initial sampling points, and a front end point is defined based on the position of the initial sampling segment where each resampling point is located. The position of the front end point is before the left endpoint of the initial sampling segment, and the length from the front end point to the left endpoint of the initial sampling segment is related to the length from the left endpoint to the right endpoint of the initial sampling segment (for example, the distances are equal, or there is a related calculation relationship), and the position of the resampling point is obtained from the positions of the front end point, the left endpoint, and the right endpoint, and then resampling is performed to obtain multiple resampling points.

[0100] Voxel-space resampling is performed based on the initial sampling to improve the accuracy of light field content generation. Specifically, the region selected by the initial sampling segment for each resampling point is determined based on the cumulative transparency distribution of the initial sampling points. The regions of the initial sampling segment where different resampling points reside may be the same or different, and the region where the resampling point resides is determined by the cumulative transparency distribution of the initial sampling points. A leading point is defined in front of each selected initial sampling segment, its position determined by the length relationship between the left and right endpoints. The distance from the leading point to the left endpoint serves as the basis for resampling, ensuring that the position of the resampling point is within a reasonable range. The distance from the leading point to the left endpoint can be slightly less than, equal to, or slightly greater than the distance from the left to the right endpoint. The resampling result is calculated based on the position information of the leading point, left endpoint, and right endpoint. This process effectively limits the range of the resampling points and prevents them from reoccupying invalid spatial regions.

[0101] S211. Based on the coordinates and direction vectors of all resampled points on each rendering ray, after position encoding and direction encoding, the coordinates and direction vectors are input into a neural network model to calculate the color and point density of all resampled points on each rendering ray, and after accumulation, the sub-pixel color corresponding to the rendering ray in the synthetic image is obtained to generate three-dimensional light field content.

[0102] According to each rendering ray emitted by the virtual rendering camera, the coordinates and direction vector of the resampling point on the rendering ray can be determined.

[0103] The coordinates and direction vectors of each resampling point on each rendering ray are encoded (such as hash coding, sine and cosine function coding, etc.) and input into the neural network model to obtain the color and point density of all resampling points, thereby calculating the transparency of the resampling points. The image pixel color corresponding to each rendering ray can be calculated by taking the weighted average of the colors of the resampling points according to the transparency of the resampling points.

[0104] The coordinates of the resampling points and the direction vectors on each light ray are encoded and input into the neural network model. In this way, the neural network model can learn how to generate corresponding color and point density information according to the position and direction of the sampling points during the training process.

[0105] Optionally, during the optimization and training of the neural network, the image pixel color corresponding to the image pixel coordinates of each rendering light ray calculated based on the neural network model can be compared with the original image pixel color corresponding to the image pixel coordinates in the original image, so as to calculate and minimize the error function as the target for optimization and training of the neural network model. During the optimization and training of the neural network, the occupancy state of the three-dimensional voxel grid is updated by the neural network model according to the position information of the three-dimensional voxel grid.

[0106] Specifically, during the optimization and training of the neural network, the occupancy state of the three-dimensional voxel grid needs to be dynamically updated. A certain intensity of Gaussian random noise is added to the center position of each three-dimensional voxel grid to obtain the corresponding position information. The position information after adding the noise is input into the neural network to obtain the corresponding point density, and whether the voxel grid is occupied is determined according to whether the point density reaches a set point density threshold. During the training process, the occupancy state of the voxel is constantly updated according to the point density of the target scene to ensure that the neural network model always accurately reflects the target scene. This method can improve the geometric distribution of the dynamically updated scene and improve the effect of adaptive ray sampling.

[0107] S212, extracting the view point information corresponding to each sub-pixel from the composite image.

[0108] S213, outputting the view point information to the three-dimensional light field display device to generate a target three-dimensional light field.

[0109] During the rendering process, the multi-path virtual rendering camera is controlled to emit rendering light rays, and the adaptive ray sampling strategy is used to calculate the coordinates of the resampling points. This process will ensure that rich scene information is captured from various angles. Through the neural network model, a composite image is generated to realize three-dimensional light field content generation. The three-dimensional light field content integrates information from different angles to provide a basis for subsequent three-dimensional light field generation.

[0110] Further, the view point information corresponding to each sub-pixel in the composite image and the color information of each sub-pixel are used as three-dimensional light field content, and a target three-dimensional light field is generated using a three-dimensional light field display device.

[0111] The adaptive light ray sampling three-dimensional light field content generation method provided by the application, by acquiring a plurality of original images of a target three-dimensional scene under different viewing angles and collecting camera calibration information of a plurality of RGB cameras for the plurality of original images; establishing a three-dimensional world coordinate system based on the camera calibration information of the plurality of RGB cameras, and constructing a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system; setting a plurality of virtual rendering cameras in the three-dimensional world coordinate system based on the parameters of a three-dimensional light field display device, emitting a rendering light ray for calculating a composite image based on light field coding, each sub-pixel in the composite image corresponding to a rendering light ray; based on a preset adaptive light ray sampling strategy, the light rays emitted by the plurality of virtual rendering cameras are initially sampled in the three-dimensional voxel space to obtain a plurality of initial sampling points, and the point density of the plurality of initial sampling points is calculated, and the transparency of the initial sampling points is calculated according to the point density; according to the set number threshold of the initial sampling points, the resampling mode is adaptively adjusted for resampling; the adaptive adjustment of the resampling mode according to the set number threshold of the initial sampling points includes: when the number of initial sampling points reaches the number threshold, the voxel space is resampled according to the cumulative transparency distribution of the plurality of initial sampling points to obtain the coordinates and direction vectors of a plurality of resampling points; when the number of initial sampling points does not reach the number threshold, new sampling points are selected in the range of the three-dimensional sampling space as resampling points to obtain the coordinates and direction vectors of a plurality of resampling points; based on the position encoding and direction encoding of all resampling point coordinates and direction vectors on each rendering light ray, the color and point density of all resampling points on each rendering light ray are calculated by inputting a neural network model, and the sub-pixel color corresponding to the rendering light ray in the composite image is obtained after accumulation, thereby generating three-dimensional light field content; and generating a target three-dimensional light field based on the three-dimensional light field content and a three-dimensional light field display device. According to the method, the sampling mode is flexibly adjusted according to the geometric distribution and detail requirements of the scene, thereby improving the quality of three-dimensional light field content generation and better adapting to the requirements of complex scenes.

[0112] The adaptive light ray sampling three-dimensional light field content generation device provided by the application is described below, and the adaptive light ray sampling three-dimensional light field content generation device described below can be correspondingly referred to the adaptive light ray sampling three-dimensional light field content generation method described above.

[0113] Figure 4 FIG. 1 is a structural schematic diagram of the adaptive light ray sampling three-dimensional light field content generation device provided by the application, and specifically includes:

[0114] The acquisition module 401 is configured to acquire a plurality of original images of a target three-dimensional scene under different viewing angles and collect camera calibration information of a plurality of RGB cameras for the plurality of original images. For details, refer to the related description of the method embodiment described above, which will not be repeated here.

[0115] The establishing module 402 is configured to establish a three-dimensional world coordinate system based on camera calibration information of the multiple RGB cameras, and construct a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system. For details, refer to the related description of the corresponding method embodiment described above, which will not be repeated here.

[0116] The setting module 403 is configured to set multiple virtual rendering cameras in the three-dimensional world coordinate system based on parameters of a three-dimensional light field display device, emit a rendering light ray for calculating a composite image based on light field encoding, and each sub-pixel in the composite image corresponds to a rendering light ray. For details, refer to the related description of the corresponding method embodiment described above, which will not be repeated here.

[0117] The sampling module 404 is configured to perform initial sampling on the light rays emitted by the multiple virtual rendering cameras in the three-dimensional voxel space based on a preset adaptive light ray sampling strategy, obtain multiple initial sampling points, calculate the point density of the multiple initial sampling points, and calculate the transparency of the initial sampling points according to the point density; and adaptively adjust the resampling mode for resampling according to a set number threshold of initial sampling points. The adaptive adjustment of the resampling mode for resampling according to the set number threshold of initial sampling points includes: when the number of initial sampling points reaches the number threshold, performing voxel space resampling according to the cumulative transparency distribution of the multiple initial sampling points to obtain the coordinates and direction vectors of multiple resampling points; and when the number of initial sampling points does not reach the number threshold, selecting new sampling points in the range of the three-dimensional sampling space as resampling points to obtain the coordinates and direction vectors of multiple resampling points. For details, refer to the related description of the corresponding method embodiment described above, which will not be repeated here.

[0118] The calculating module 405 is configured to input the coordinates and direction vectors of all resampling points on each rendering light ray into a neural network model after position encoding and direction encoding to calculate the color and point density of all resampling points on each rendering light ray, and accumulate to obtain the sub-pixel color corresponding to the rendering light ray in the composite image, and generate three-dimensional light field content. For details, refer to the related description of the corresponding method embodiment described above, which will not be repeated here.

[0119] The generating module 406 is configured to generate a target three-dimensional light field based on the three-dimensional light field content and a three-dimensional light field display device. For details, refer to the related description of the corresponding method embodiment described above, which will not be repeated here.

[0120] Figure 5 An example of an electronic device is shown in the schematic diagram of the physical structure of the electronic device as shown in Figure 5As shown, the electronic device may include: a processor 810 , a communication interface 820 , a memory 830 and a communication bus 840 , wherein the processor 810 , the communication interface 820 and the memory 830 communicate with each other via the communication bus 840 . The processor 810 can call the logic instructions in the memory 830 to execute the adaptive light sampling three-dimensional light field content generation method, which includes: obtaining multiple original images of the target three-dimensional scene from different perspectives and camera calibration information of multiple RGB cameras that collect the multiple original images; establishing a three-dimensional world coordinate system based on the camera calibration information of the multiple RGB cameras, and constructing a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system; setting multiple virtual rendering cameras based on parameters of the three-dimensional light field display device in the three-dimensional world coordinate system, emitting rendering rays for calculating a synthetic image based on light field encoding, wherein each sub-pixel in the synthetic image corresponds to one rendering ray; performing initial sampling of the rays emitted by the multiple virtual rendering cameras in the three-dimensional voxel space based on a preset adaptive light sampling strategy to obtain multiple initial sampling points, and calculating the point density of the multiple initial sampling points, and calculating the transparency of the initial sampling points according to the point density of the initial sampling points. Adaptively adjusting the resampling method to perform resampling according to a set threshold number of initial sampling points; the adaptively adjusting the resampling method to perform resampling according to the set threshold number of initial sampling points comprises: when the number of initial sampling points reaches the threshold number, performing voxel space resampling according to the cumulative transparency distribution of the plurality of initial sampling points to obtain coordinates and direction vectors of the plurality of resampling points; when the number of initial sampling points does not reach the threshold number, reselecting new sampling points within the three-dimensional sampling space as resampling points to obtain coordinates and direction vectors of the plurality of resampling points; calculating the color and point density of all resampling points on each rendering ray based on the coordinates and direction vectors of all resampling points on each rendering ray after position encoding and direction encoding, and accumulating the colors of sub-pixels corresponding to the rendering ray in a composite image to generate three-dimensional light field content; and generating a target three-dimensional light field based on the three-dimensional light field content and the three-dimensional light field display device.

[0121] In addition, the logic instructions in the memory 830 described above can be implemented in the form of software function units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the parts that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0122] In another aspect, the present application also provides a computer program product comprising a computer program, which can be stored on a non-transitory computer readable storage medium, and the computer program is executable by a processor to enable a computer to perform the adaptive light sampling three-dimensional light field content generation method provided by the above-mentioned methods, which comprises: acquiring a plurality of original images of a target three-dimensional scene under different viewing angles and camera calibration information of a multi-channel RGB camera for collecting the plurality of original images; establishing a three-dimensional world coordinate system based on the camera calibration information of the multi-channel RGB camera, and constructing a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system; setting a plurality of virtual rendering cameras based on the parameters of a light field display device in the three-dimensional world coordinate system, emitting a rendering light for calculating a composite image based on light field encoding, wherein each sub-pixel in the composite image corresponds to a rendering light; performing initial sampling on the lights emitted by the plurality of virtual rendering cameras in the three-dimensional voxel space based on a preset adaptive light sampling strategy, obtaining a plurality of initial sampling points, and calculating the point density of the plurality of initial sampling points, and calculating the transparency of the initial sampling points according to the point density; adaptively adjusting the resampling mode for resampling according to the set number threshold of the initial sampling points; the adaptive adjustment of the resampling mode for resampling according to the set number threshold of the initial sampling points comprises: when the number of initial sampling points reaches the number threshold, performing voxel space resampling according to the cumulative transparency distribution of the plurality of initial sampling points to obtain the coordinates and direction vectors of a plurality of resampling points; when the number of initial sampling points does not reach the number threshold, reselecting new sampling points in the range of the three-dimensional sampling space as resampling points to obtain the coordinates and direction vectors of a plurality of resampling points; inputting the coordinates and direction vectors of all resampling points on each rendering light into a neural network model after position encoding and direction encoding to calculate the color and point density of all resampling points on each rendering light, and accumulating to obtain the sub-pixel color corresponding to the rendering light in the composite image, thereby generating three-dimensional light field content; and generating a target three-dimensional light field based on the three-dimensional light field content and the three-dimensional light field display device.

[0123] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the adaptive light sampling three-dimensional light field content generation method provided by the above method, and the method comprises: acquiring a plurality of original images of a target three-dimensional scene under different viewing angles and camera calibration information of a plurality of RGB cameras collecting the plurality of original images; establishing a three-dimensional world coordinate system based on the camera calibration information of the plurality of RGB cameras, and constructing a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system; setting a plurality of virtual rendering cameras based on the parameters of a three-dimensional light field display device in the three-dimensional world coordinate system, emitting a rendering light for calculating a composite image based on light field encoding, wherein each sub-pixel in the composite image corresponds to a rendering light; performing initial sampling on the lights emitted by the plurality of virtual rendering cameras in the three-dimensional voxel space based on a preset adaptive light sampling strategy to obtain a plurality of initial sampling points, and calculating the point density of the plurality of initial sampling points, and calculating the transparency of the initial sampling points according to the point density; adaptively adjusting the resampling mode to perform resampling according to the set number threshold of the initial sampling points; the adaptive adjustment of the resampling mode according to the set number threshold of the initial sampling points comprises: when the number of initial sampling points reaches the number threshold, performing voxel space resampling according to the cumulative transparency distribution of the plurality of initial sampling points to obtain the coordinates and direction vectors of a plurality of resampling points; when the number of initial sampling points does not reach the number threshold, selecting new sampling points in the range of the three-dimensional sampling space as resampling points to obtain the coordinates and direction vectors of a plurality of resampling points; inputting the coordinates and direction vectors of all resampling points on each rendering light into a neural network model after position encoding and direction encoding to calculate the color and point density of all resampling points on each rendering light, and accumulating to obtain the sub-pixel color corresponding to the rendering light in the composite image, thereby generating three-dimensional light field content; and generating a target three-dimensional light field based on the three-dimensional light field content and the three-dimensional light field display device.

[0124] The device embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0125] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0126] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for generating three-dimensional light field content using adaptive light sampling, characterized in that: include: Acquire multiple original images of a target three-dimensional scene at different viewing angles and camera calibration information of multiple RGB cameras that acquire the multiple original images; Establishing a three-dimensional world coordinate system based on camera calibration information of the multiple RGB cameras, and constructing a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system; Setting a multi-channel virtual rendering camera based on parameters of a three-dimensional light field display device in the three-dimensional world coordinate system, and emitting rendering rays for calculating a composite image based on light field coding, wherein each sub-pixel in the composite image corresponds to one of the rendering rays; performing initial sampling of light emitted by the multi-channel virtual rendering camera in the three-dimensional voxel space based on a preset adaptive light sampling strategy to obtain a plurality of initial sampling points, calculating a point density of the plurality of initial sampling points, and calculating the transparency of the initial sampling points according to the point density of the initial sampling points; According to the set threshold of the number of initial sampling points, the resampling method is adaptively adjusted to perform resampling; The step of adaptively adjusting the resampling mode to perform resampling according to the set threshold value of the number of initial sampling points includes: When the number of initial sampling points reaches the number threshold, performing voxel space resampling according to the cumulative transparency distribution of the multiple initial sampling points to obtain coordinates and direction vectors of the multiple resampling points; When the number of initial sampling points does not reach the threshold, new sampling points are selected as resampling points within the three-dimensional sampling space to obtain coordinates and direction vectors of the plurality of resampling points; The coordinates and direction vectors of all resampled points on each rendering ray are position-encoded and direction-encoded, and then input into a neural network model to calculate the color and point density of all resampled points on each rendering ray. After accumulation, the sub-pixel color corresponding to the rendering ray in the composite image is obtained to generate three-dimensional light field content. A target three-dimensional light field is generated based on the three-dimensional light field content and the three-dimensional light field display device.

2. The method according to claim 1, characterized in that The establishing of a three-dimensional world coordinate system based on the camera calibration information of the multiple RGB cameras, and constructing a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system include: The three-dimensional voxel space and the three-dimensional sampling space are constructed based on the three-dimensional world coordinate system, wherein the three-dimensional voxel space represents the spatial range of sampling points obtained by initial sampling or resampling based on the three-dimensional voxel grid, and the three-dimensional sampling space represents the spatial range of new sampling points selected when the number of the initial sampling points does not reach the number threshold.

3. The method according to claim 1, characterized in that The method further comprises: During the optimization training of the neural network model, the neural network model is optimized and trained based on the pixel colors of the multi-channel RGB cameras and the corresponding original image pixel colors in the original image with the goal of minimizing a preset error function; During the optimization training of the neural network model, the occupancy state of the three-dimensional voxel grid is updated according to the position information of the three-dimensional voxel grid in the three-dimensional voxel space.

4. The method according to claim 2, characterized in that When the number of initial sampling points reaches the number threshold, voxel space resampling is performed according to the cumulative transparency distribution of the multiple initial sampling points to obtain coordinates and direction vectors of the multiple resampling points, including: determining a plurality of initial sampling line segments based on the plurality of initial sampling points; The midpoint position of each initial sampling line segment is input into the neural network model to obtain the point density corresponding to each initial sampling point; Calculate the transparency of each initial sampling point according to the point density of each initial sampling point; Accumulate the transparency of each initial sampling point on each rendering ray along the sampling ray to obtain a cumulative distribution of the transparency of the initial sampling point on each rendering ray; Determine an initial sampling line segment area corresponding to each resampling point based on the cumulative transparency distribution of the multiple initial sampling points; Resampling is performed based on the initial sampling line segment area to obtain coordinates and direction vectors of multiple resampling points.

5. The method according to claim 4, characterized in that The method of setting a multi-channel virtual rendering camera based on parameters of a three-dimensional light field display device in the three-dimensional world coordinate system includes: Setting the number of virtual rendering cameras according to the number of viewpoints of the three-dimensional light field display device; Setting a position of the virtual rendering camera from the target three-dimensional scene according to an optimal viewing distance of the three-dimensional light field display device; The spacing of the virtual rendering cameras is set according to the maximum viewing angle of the three-dimensional light field display device.

6. The method according to claim 5, characterized in that Generating a target three-dimensional light field based on the three-dimensional light field content and the three-dimensional light field display device includes: Extracting viewpoint information corresponding to each sub-pixel from the synthesized image; The viewpoint information is output to the three-dimensional light field display device, and a target three-dimensional light field is generated based on the three-dimensional light field display apparatus.

7. An adaptive light sampling 3D light field content generation device, characterized in that: include: An acquisition module is used to acquire multiple original images of the target three-dimensional scene from different perspectives and camera calibration information of the multiple RGB cameras that acquire the multiple original images; An establishment module, configured to establish a three-dimensional world coordinate system based on the camera calibration information of the multi-channel RGB camera, and construct a three-dimensional voxel space and a three-dimensional sampling space based on the three-dimensional world coordinate system; A setting module is configured to set a multi-channel virtual rendering camera based on parameters of a three-dimensional light field display device in the three-dimensional world coordinate system, and emit rendering rays for calculating a composite image based on light field coding, wherein each sub-pixel in the composite image corresponds to one of the rendering rays; a sampling module, configured to perform initial sampling of the light emitted by the multi-channel virtual rendering camera in the three-dimensional voxel space based on a preset adaptive light sampling strategy to obtain a plurality of initial sampling points, calculate the point density of the plurality of initial sampling points, and calculate the transparency of the initial sampling points according to the point density; Adaptively adjusting the resampling mode to perform resampling according to a set threshold value for the number of initial sampling points; said adaptively adjusting the resampling mode to perform resampling according to the set threshold value for the number of initial sampling points includes: when the number of initial sampling points reaches the threshold value, performing voxel space resampling according to the cumulative transparency distribution of the plurality of initial sampling points to obtain coordinates and direction vectors of the plurality of resampling points; when the number of initial sampling points does not reach the threshold value, reselecting new sampling points within the range of the three-dimensional sampling space as resampling points to obtain coordinates and direction vectors of the plurality of resampling points; A calculation module is used to calculate the color and point density of all resampled points on each rendering ray based on the coordinates and direction vectors of all resampled points on each rendering ray after position encoding and direction encoding, and to obtain the sub-pixel color corresponding to the rendering ray in the composite image after accumulation, thereby generating three-dimensional light field content; A generating module is configured to generate a target three-dimensional light field based on the three-dimensional light field content and the three-dimensional light field display device.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for generating three-dimensional light field content using adaptive ray sampling according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating three-dimensional light field content using adaptive ray sampling is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating three-dimensional light field content using adaptive ray sampling is implemented.

Citation Information

Patent Citations

  • Three-dimensional light field image generation method and device, display equipment and storage medium

    CN118381888A

  • Light coding method and device for projection three-dimensional display, medium and product

    CN118474333A