A digital sand table system and light field image rendering method using image rendering

By combining image rendering with neural radiation field and light field rendering technologies, the digital sand table system solves the problem of poor interactivity in existing digital sand table systems, and realizes real-time changes in stereoscopic images and improves interactive effects.

CN119228966BActive Publication Date: 2025-11-11NANJING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411390103.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-11-11
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

Existing digital sand table systems suffer from problems such as poor interactivity due to fixed models and inability to change the displayed scene in real time.

Method used

The digital sand table system employs image rendering, utilizing image acquisition units, image rendering units, data transmission units, and directional rendering units, combined with neural radiation field and light field rendering technologies, to reconstruct three-dimensional scenes and render directional light rays, forming stereoscopic images.

Benefits of technology

It improves interactivity and data processing efficiency, can change stereoscopic image content in real time, enhances interactive effects, and adapts to stereoscopic display from different viewing angles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119228966B_ABST
    Figure CN119228966B_ABST
Patent Text Reader

Abstract

This application discloses a digital sand table system and a light field image rendering method using image rendering. The system reconstructs a 3D scene model using neural radiation fields and performs light field rendering on the scene image. It uses eye tracking to locate the viewer's position and performs directional light rendering, significantly improving processing efficiency compared to rendering images from all angles. Furthermore, the system reproduces the rendered directional image using light field display technology, allowing the stereoscopic image to change with the viewer, enhancing interactivity. In addition, the system can change the content of the stereoscopic image by modifying the 3D scene image and provides interactive processing operations, offering advantages over traditional sand tables in terms of freely changeable stereoscopic image content and superior interactive effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and optical imaging, specifically to a digital sand table system and a light field image rendering method that employs image rendering. Background Technology

[0002] Existing sand table systems are primarily implemented using physical sand table models constructed from sand or 3D printed. However, these systems suffer from drawbacks such as fixed content and the inability to update data in real time. For example, in the defense field, this leads to a reduction in the timeliness of battlefield information data, failing to meet the needs for real-time support in the rapidly changing battlefield environment.

[0003] Current 3D image stereoscopic display technology has several mainstream architectures: one is to use optical instruments such as projectors for 3D reconstruction; the other is to use wearable interactive devices (such as VR glasses) based on virtual reality technology for visualization simulation. However, at the hardware level, the limitations of projectors and wearable interactive devices prevent digital sand tables from integrating into a more complete and flexible intelligent system, and they are also heavily reliant on physical objects, failing to fully meet the needs of specific fields for electronic sand table systems. In currently widespread 3D stereoscopic sand tables, most adopt stereoscopic projection combined with physical models to realize the stereoscopic sand table model device.

[0004] In summary, existing technologies have the following drawbacks: First, the technology of combining virtual and real elements is based on physical models, which makes it difficult to change the sand table model once the structure is determined; second, existing sand tables have poor interactivity and cannot be changed in real time according to user needs; third, stereoscopic projection and volumetric display sand tables occupy a large area and are not easy to disassemble, transport, or move; fourth, lenticular stereoscopic displays can only be viewed from a fixed direction and are prone to causing eye fatigue, which does not meet the actual use conditions of sand tables. Summary of the Invention

[0005] The technical problem addressed by this application is that existing digital sand table displays are affected by fixed models, resulting in poor interactivity and difficulty in changing the display scene. To solve these problems, this application provides a digital sand table system using image rendering and a light field image rendering method.

[0006] According to a first aspect of this application, this application provides a digital sand table system employing image rendering, comprising: an image acquisition unit for acquiring images of a three-dimensional scene, obtaining scene images, and transmitting them to a host computer; an image rendering unit, running on the host computer, for acquiring the scene images from the image acquisition unit, and using the scene images to reconstruct the three-dimensional scene and perform light field rendering to obtain a light field model; a data transmission unit, running on the host computer, for converting the light field model into a neural network framework type to obtain a converted model and deploying it to a lower-level computer; a directional rendering unit, running on the lower-level computer, for tracking and obtaining the orientation information of the viewer, and using the orientation information to perform directional light rendering on the converted model to obtain a directional rendered image; and an image display unit, communicatively connected to the lower-level computer, for projecting the directional rendered image in three dimensions to form a digital sand table, the digital sand table including a stereoscopic image of the three-dimensional scene facing the viewer.

[0007] Furthermore, the image rendering unit includes a model training module and a light field image generation module; the model training module is used to reconstruct the three-dimensional scene using a neural radiation field to obtain a scene model; the light field image generation module is used to perform multi-view light field rendering on the scene image using the scene model to generate a light field image of the three-dimensional scene and construct a corresponding light field model.

[0008] Furthermore, the data transmission unit includes a neural network framework transfer module and a distributed Ethernet port module; the neural network framework transfer module is used to convert the light field model from a first neural network framework type to a second neural network framework type to obtain a corresponding converted model; the light field model has a first neural network framework type adapted to the host computer, and the converted model has a second neural network architecture type adapted to the slave computer; the distributed Ethernet port module is used to deploy the converted model to the slave computer.

[0009] Furthermore, the directional rendering unit includes an eye-tracking module, a light field image rendering module, and an HDMI transmission module; the eye-tracking module is used to determine the viewer's position and viewing angle through eye tracking, forming the orientation information; the light field image rendering module receives the conversion model and uses the orientation information as input parameters to perform directional light rendering on the conversion model to obtain a directional rendering image of the three-dimensional scene; the HDMI transmission module is used to transmit the directional rendering image to the image display unit using the HDMI protocol.

[0010] Furthermore, the image display unit includes a collimated backlight module, a screen display module, and a composite lens array module; the collimated backlight module is used to provide collimated backlight for the screen display module; the screen display module is used to receive the directional rendering image transmitted by the HDMI module and to display the directional rendering image; the composite lens array module is used to focus the light rays of each pixel in the image of the screen display module, and the resulting three-dimensional projection forms the digital sand table.

[0011] Furthermore, the collimated backlight module has an extinction structure, which eliminates stray light when light is projected onto the screen display module, thereby reducing crosstalk in the display and improving the contrast and clarity of the image.

[0012] Furthermore, the image acquisition unit is used to acquire video of the three-dimensional scene, and uses video frame extraction technology to extract frame images from the video of the three-dimensional scene to obtain images of the three-dimensional scene from multiple perspectives and use them as the scene images.

[0013] According to a second aspect of this application, an image rendering method is provided, comprising: acquiring an image of a three-dimensional scene to obtain a scene image; reconstructing and rendering the three-dimensional scene using the scene image to obtain a light field model; converting the light field model to a neural network framework type to obtain a conversion model, wherein the conversion model can be deployed to a lower-level machine; tracking and obtaining the orientation information of the viewer; using the orientation information to perform directional light rendering on the conversion model to obtain a directional rendering image, wherein the directional rendering image is used for three-dimensional projection.

[0014] Furthermore, the step of reconstructing the 3D scene and rendering the light field using the scene image to obtain a light field model includes: reconstructing the 3D scene using a neural radiation field to obtain a scene model; the light field image generation module is used to perform multi-view light field rendering on the scene image using the scene model to generate a light field image of the 3D scene and construct a corresponding light field model; wherein, during multi-view light field rendering, the scene model can learn the emissivity and light propagation direction of the 3D scene through the scene image.

[0015] Furthermore, the step of tracking to obtain the viewer's location information and using the location information to perform directional ray rendering on the conversion model to obtain a directional rendering image includes: determining the viewer's position and viewing angle through human eye tracking to form the location information; using the location information as input parameters to perform directional ray rendering on the conversion model to obtain a directional rendering image of the three-dimensional scene; wherein, when performing directional ray rendering on the conversion model, the input parameters of the conversion model also include focal length and resolution.

[0016] The beneficial effects of this application are:

[0017] The digital sand table system and image rendering method using image rendering according to the above embodiments reconstruct a 3D scene model using neural radiation fields and perform light field rendering on the scene image. By tracking the viewer's position with the human eye and performing directional light rendering, the processing efficiency is significantly improved compared to rendering images from all angles. The image acquisition unit, image rendering unit, and data transmission unit in the system rely on a host computer, while the directional rendering unit and image display unit rely on a slave computer. This allows for separate processing of establishing the light field model and generating the directional rendered image, significantly saving time in both stages and improving data processing efficiency. The rendered directional image is reproduced using light field display technology, and the stereoscopic image can change according to the viewer, enhancing the interactive effect. This digital sand table system can change the content of the stereoscopic image by modifying the image of the 3D scene and provides interactive processing operations. Compared to traditional sand tables, it has the advantages of freely changing the content of the stereoscopic image and better interactive effects. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the structure of a digital sand table system using image rendering in one embodiment of this application;

[0019] Figure 2 This is a schematic diagram illustrating the principle of the image rendering process in one embodiment of this application;

[0020] Figure 3 This is a schematic diagram illustrating the principle of creating an example model in one embodiment of this application;

[0021] Figure 4 This is a schematic diagram illustrating the principle of generating a light field image in one embodiment of this application;

[0022] Figure 5 This is a schematic diagram illustrating the principle of training a model in one embodiment of this application;

[0023] Figure 6 This is a schematic diagram of the structure of the image display module in one embodiment of this application;

[0024] Figure 7 This is a flowchart illustrating an image rendering method in one embodiment of this application;

[0025] Figure 8 This is a schematic diagram illustrating the control of image rendering and display in one embodiment of this application. Detailed Implementation

[0026] In the description of this invention, it should be understood that the described embodiments are only a part of the embodiments of this invention, and not all of the embodiments.

[0027] The present application will now be described in further detail with reference to specific embodiments and accompanying drawings.

[0028] This application discloses a digital sand table system using image rendering, the structure of which is as follows: Figure 1 As shown, the digital sand table system mainly includes an image acquisition unit 11, an image rendering unit 120, a data transmission unit 130, a directional rendering unit 140, and an image display unit 150. These will be described in detail below.

[0029] The image acquisition unit 11 can be one or more cameras surrounding the 3D scene, used to acquire images of the 3D scene, obtain scene images, and transmit them to the host computer. Here, the 3D scene refers to a space with 3D objects, and the corresponding scene image refers to an image containing 3D objects; the host computer can be a data processing device such as a computer.

[0030] The image rendering unit 120 runs on the host computer and is used to acquire scene images from the image acquisition unit 11, and to reconstruct the three-dimensional scene and render the light field using the scene images to obtain the light field model.

[0031] The data transmission unit 130 runs on the host computer and is used to convert the light field model generated by the image rendering unit 120 into a neural network framework type, obtain the converted model, and deploy it to the slave computer. The slave computer can be a microprocessor (MCU) or other data processing component, capable of mutual data communication with the host computer. The initial non-real-time data processing can be performed on the host computer, while the subsequent real-time data processing can be performed on the slave computer. This combination improves data processing efficiency.

[0032] The directional rendering unit 140 runs on the lower-level machine and is used to track and obtain the viewer's location information. It then uses the location information to perform directional ray rendering on the transformation model to obtain a directional rendering image.

[0033] The image display unit 150 is connected to the lower-level computer and is used to project the directional rendering image in three dimensions to form a digital sand table. The digital sand table includes a stereoscopic image of the three-dimensional scene facing the viewer.

[0034] It should be noted that a digital sand table is a display method that uses digital technology, combining sound, light, electricity, images, animation, etc., with physical models, or directly displaying them on a screen, to convey information to people in a lifelike and varied visual effect. It can provide viewers with three-dimensional images, aiming to provide viewers with an intuitive and vivid information display experience.

[0035] In one specific embodiment, the image acquisition unit 11 is used to acquire video of a 3D scene. It employs video frame extraction technology to extract frame images from the 3D scene video, obtaining images of the 3D scene from multiple perspectives and using them as scene images. For example, the image acquisition unit 11 can acquire multi-angle 2D images of a 3D object, taking pictures of a 3D object from different angles using a camera and transmitting the images to the image rendering unit 120. In actual operation, to facilitate the acquisition of sparse perspectives of the 3D scene, video frame extraction technology is preferentially used to acquire images of the 3D scene, and the acquired coordinate angle information and image information are transmitted to the image rendering unit 120. It can be understood that image data containing 3D scene information is acquired from different perspectives. These images can be real scene images captured by a camera or other devices, or synthetic images generated by computer graphics technology.

[0036] In one specific embodiment, the image rendering unit 120 includes a model training module 12 and a light field image generation module 13. The model training module 12 is used to reconstruct a 3D scene using a neural radiation field to obtain a scene model; the light field image generation module 13 is used to perform multi-view light field rendering on the scene image using the scene model, generating a light field image of the 3D scene and constructing a corresponding light field model. It can be understood that the image rendering unit 120 mainly uses a neural radiation algorithm to train the 3D model and renders and generates light field models corresponding to multi-angle images based on the training content.

[0037] It is understandable that the model training module 12 uses the training set for deep learning, combining NeRF with memory neural radiation fields to make model generation faster and improve algorithm efficiency; the light field image generation module 13 uses the neural network trained by the model training module 12 to transmit the acquired scene images to the neural network for generation, thus enabling multi-view square rendering and generating light field images of the corresponding 3D scene from multiple perspectives. In other words, the light field image generation module 13 uses the multi-angle 2D images (i.e., scene images) transmitted by the image acquisition unit 11 to perform light field model reconstruction operations on the 3D object, generating the model of the scene. The computational framework under this module is Pytroch, meaning that the light field model has a first neural network framework type adapted to the host computer.

[0038] It's important to note that Neural Radiance Field (NeRF) is a type of neural network that can reconstruct complex 3D scenes from a partial set of 2D images. Various simulation, gaming, media, and Internet of Things (IoT) applications require 3D images to make digital interactions more realistic and accurate. NeRF learns the scene geometry, objects, and angles of a specific scene, then renders realistic 3D views from a new perspective, automatically generating synthetic data to fill in the gaps.

[0039] In one specific embodiment, the data transmission unit 130 includes a neural network framework transfer module 14 and a distributed Ethernet port module 15. The neural network framework transfer module 14 is used to convert the optical field model from a first neural network framework type to a second neural network framework type, obtaining a corresponding converted model. The optical field model has a first neural network framework type adapted to the host computer (e.g., the Pytroch framework), and the converted model has a second neural network architecture type adapted to the slave computer (e.g., a mobile terminal framework). The distributed Ethernet port module 15 is used to deploy the converted model to the slave computer.

[0040] It should be noted that the neural network framework transfer module 14 converts the training framework PyTorch into a mobile framework through the ONNX structure, enabling cross-platform deployment and inference of the model while preserving its accuracy and performance. The distributed Ethernet port module 15 uses network transmission technology to transmit the converted model to the lower-level machine for image rendering.

[0041] It should be noted that the data transmission unit 130 can transfer the light field model of the PyTorch framework in the light field image generation module 13 to a personal mobile device, and transmit the directionally rendered light field image to the display screen through the personal mobile device to realize a stereoscopic image facing the viewer. Specifically, the neural network framework transfer module 14 mainly converts the plaza model under the PyTorch framework in the light field image generation module 13 into other platform frameworks, and the distributed Ethernet port module 15 mainly connects the neural network framework transfer module 14 to the image display unit 140, transmitting the received light field model to the directional rendering unit 140.

[0042] In one specific embodiment, the directional rendering unit 140 includes an eye-tracking module 17, a light field image rendering module 16, and an HDMI transmission module 18.

[0043] The eye-tracking module 17 can be a camera or video camera, capable of capturing and recognizing images of the human eye. It is used to determine the viewer's position and viewing angle through eye tracking, forming orientation information. In essence, the eye-tracking module 17 is used to locate the viewer's position and viewing angle, inputting the viewer's orientation information (position and viewing angle) into the light field image rendering module 16. The eye-tracking module 17 can track the viewer's orientation in real time and transmit information, ensuring that the stereoscopic effect is displayed even when the viewer moves.

[0044] The light field image rendering module 16 runs on the lower-level machine and can receive the conversion model from the upper-level machine where the data transmission unit 130 is located. It uses the orientation information generated by the eye-tracking module 17 as input parameters to perform directional light rendering on the conversion model, obtaining a directional rendered image of the 3D scene. In essence, the light field image rendering module 16 determines the viewer's perspective (orientation) through the eye-tracking module 17, and performs directional light rendering of the image based on the viewer's perspective, using the received conversion model for directional rendering.

[0045] The HDMI transmission module 18 runs on the host computer and is used to transmit the directional rendering image to the image display unit 150 via the HDMI protocol. It can be understood that the HDMI transmission module 18 connects the light field image rendering module 16 and the image display unit 150 through the HDMI interface, transmitting the directional rendering image (which can be understood as the light field image to be displayed) generated by the light field image rendering module 16 to the image display unit 150 for stereoscopic display.

[0046] In one specific embodiment, the image display unit 150 includes a collimated backlight module 110, a screen display module 19, and a compound lens array module 111.

[0047] The collimation backlight module 110, also known as the light source, is used to provide collimation backlight for the screen display module 19.

[0048] The screen display module 19 can be an LCD screen used to receive directional rendering images sent from the HDMI transmission module and to display the directional rendering images; the displayed images are illuminated by collimated backlight.

[0049] The composite lens array module 111 is used to focus the light from each pixel in the image of the screen display module, and the resulting three-dimensional projection forms a digital sand table.

[0050] It can be understood that the image display unit 150 transforms the directional rendering image sent by the HDMI transmission module into a three-dimensional stereoscopic image through the basic principle of light field display. The collimation backlight module 110 can provide the screen display module 19 with light of set direction and intensity, and the composite lens array module 111 can perform stereoscopic display of the light field, thereby presenting a stereoscopic image.

[0051] It is understandable that the light field image rendering module 16 uses the lower-level machine to perform directional rendering of the image using a light field model after the frame type conversion. The human eye tracking module 17 determines the viewer's location through human eye tracking so that the light field image rendering module 16 can perform directional rendering, saving rendering time and improving efficiency. The directionally rendered light field image is transmitted to the screen display module 19 through the HDMI transmission module 18. The screen display module 19 can be composed of a borderless 2K / 4K screen. The borderless screen is easy to splice and combine, making it convenient to build the minimum system into the sand table structure we need. The collimation backlight module 110 uses a light-absorbing material to make an extinction structure. By absorbing the scattered light through the extinction structure, most of the light can propagate out within a certain range. The screen display module 19 receives the directionally rendered image generated by the light field image rendering module 16 through the HDMI transmission module 18 and displays the directionally rendered image on the screen. The compound lens array module 111 is connected to the screen display module 19. The compound lens array module 111 projects the displayed image, so that the viewer sees different images at different positions. The stereoscopic projection effect is achieved by the viewer's movement.

[0052] In one specific embodiment, the collimated backlight module 110 has an extinction structure. When light is projected onto the screen display module 19, the extinction structure eliminates stray light, thereby reducing crosstalk in the image display and improving the contrast and clarity of the image. For example, the collimated backlight module 110 projects light onto the screen display module 19 through a designed extinction structure to illuminate the screen image. Moreover, the extinction structure can be a mesa-shaped structure to absorb some of the reflected light, reducing interference from ambient light and thus improving the visual effect.

[0053] Furthermore, the screen display module 19 can be a borderless 2K / 4K display screen, displaying directional rendered images transmitted by the HDMI transmission module 18. Multiple display screens can be combined to achieve the splicing capability of a minimal system. The composite lens array module 111 uses the principle of light field display to guide light using a composite lens array, thereby presenting the image displayed by the screen display module 19 in a three-dimensional form, resulting in a stereoscopic image 112.

[0054] It should be noted that the digital sand table system uses a composite lens array to display the processed images in three dimensions and can refresh the displayed content in real time. The system renders the multi-view images of a pre-input model or a real-world physical model using neural radiation fields to obtain a meta-image and performs sub-pixel position mapping to improve the macro-pixel resolution of the light field image. The system uses a collimated light source as backlighting, which eliminates gaps between spliced ​​screens. Simultaneously, it incorporates eye-tracking technology to locate the viewer's perspective and viewing position for directional rendering, reducing rendering time. Because the digital sand table system adopts a unit / modular design, it can be infinitely expanded from a minimum system. The image generation method uses a neural radiation algorithm, generating light field images through end-to-end convolutional neural network training.

[0055] In one embodiment, Figure 2 A schematic diagram of the image rendering process is provided, and details can be found in each processing step from step 21 to step 26.

[0056] Step 21: The image acquisition unit 11 acquires data from the 3D scene, that is, it acquires image data containing scene information from different perspectives, thus obtaining scene images. These images can be real scene videos captured by cameras or other devices and processed by frame extraction, or they can be synthetic images generated by computer graphics technology.

[0057] Step 22: The model training module 12 in the image rendering unit 120 performs neural network training on the NeRF model (neural radiation field), uses the training set to train the NeRF neural network, and uses the light field image generation module 13 in the image rendering unit 120 to learn the emissivity and light propagation direction in the scene image, so as to obtain a multi-view light field image and construct a light field model.

[0058] Step 23: The neural network framework conversion module 14 in the data transmission unit 130 first performs the PyTorch framework to ONNX conversion. By converting the PyTorch framework to ONNX format, the model can be seamlessly transferred between different frameworks.

[0059] Step 24: The neural network framework conversion module 14 in the data transmission unit 130 then executes the ONNX to mobile framework conversion to achieve cross-platform deployment and inference of the model. ONNX is an open deep learning model exchange format that allows for model conversion and migration between different deep learning frameworks, thus enabling cross-platform deployment.

[0060] In step 25, the distributed Ethernet port module 14 in the data transmission unit 130 deploys the conversion model to the lower-level machine. The light field image rendering module 16 in the directional rendering unit 140 uses the trained NeRF model, sets the corresponding focal length and resolution, inputs the viewer's position and direction (i.e., the orientation information generated by the human eye tracking module 17) for reordering, and outputs the corresponding directional rendering image through the neural network.

[0061] In step 26, the collimation backlight module 110 in the image display unit 150 generates collimation backlight, the screen display module 19 receives the directional rendering image generated by the light field image rendering module 16 through the HDMI transmission module 18, displays the directional rendering image on the screen, and the composite lens array module 111 projects the displayed image so that the viewer can see different images at different positions. The stereoscopic projection effect is achieved by the viewer's movement, thereby generating a digital sand table.

[0062] In acquiring multi-angle 2D images of objects using a camera, to avoid the influence of factors such as component manufacturing errors and redundant lighting, a virtual camera can be used for rendering in a computer simulation environment. This method allows for rapid adjustment of experimental parameters. Software is used to acquire virtual 3D scenes or 3D reconstructions of real scenes, and a camera acquisition array is established. Image information and camera coordinate information are then transmitted to the vision-driven module.

[0063] In one embodiment, the specific operation is as follows: Figure 3 As shown, a schematic diagram illustrating the principle of creating an example model is disclosed. This involves creating a simple 3D model, adjusting virtual camera parameters, adjusting the virtual camera's position and posture, adjusting the size and orientation of the 3D model, and taking photos from multiple angles. This process can be performed using an image acquisition unit. These steps are explained below.

[0064] Step 31: The image acquisition unit 11 establishes a 3D model (i.e., an instance model corresponding to the three-dimensional scene) according to the user's requirements.

[0065] Step 32: The image acquisition unit 11 adjusts the virtual camera parameters, specifically setting the virtual camera resolution, setting the camera type to perspective, adjusting the camera's focal length, lens unit, and the horizontal width of the camera sensor element, etc.

[0066] Step 33: The image acquisition unit 11 adjusts the position of the virtual camera and sets virtual cameras in different positions.

[0067] Step 34: The image acquisition unit 11 adjusts the virtual camera posture and sets virtual cameras at different angles.

[0068] Step 35: The image acquisition unit 11 adjusts the size and position of the 3D model.

[0069] Step 36: The image acquisition unit 11 adjusts the viewpoint of the 3D model.

[0070] Step 37: The image acquisition unit 11 uses a virtual camera to capture multi-angle images of the 3D model, forming scene images.

[0071] This application uses the NeRF algorithm, the basic principle of which is as follows: Figure 4 This can also be seen as a schematic diagram of the principle of generating a light field image, which is divided into a data reading step 41, a 5D function expression step 42, a deep fully connected neural network optimization step 43, a volume rendering step 44, an image reordering step 45, and an output step 46. These will be explained below.

[0072] Step 41: The image rendering unit 120 traverses the scene through camera rays, generates a set of 3D sampling points, and inputs these points and their corresponding 2D perspectives into the neural network.

[0073] Step 42, the image rendering unit 120 represents a static scene as a continuous 5D function, which outputs the radiance emitted in each direction (θ,φ) at each point (x,y,z) in space, as well as the density at each point, which is expressed as differential opacity, controlling the amount of radiance accumulated when light passes through (x,y,z).

[0074] Step 43: Since the neural network does not have convolutional layers (often called a multilayer perceptron or MLP), the image rendering unit 120 needs to represent the function by regressing from a single 5D coordinate (x,y,z,θ,φ) to a single volume density and view-dependent RGB color, thereby achieving deep fully connected neural network optimization.

[0075] In step 44, the image rendering unit 120 uses classic volumetric rendering techniques to accumulate these colors and densities into a 2D image. This model can be optimized using gradient descent, achieving this by minimizing the error between each viewed image and the corresponding view rendered from the radiation field representation.

[0076] Step 45: The image rendering unit 120 reorders the acquired 2D images using the predicted model to form the required light field image and construct the corresponding light field model.

[0077] In one embodiment, Figure 5 The schematic diagram of the training model is disclosed, which can be divided into the following steps: data input (step 51), NeRF creation (step 52), rendering process (step 54), and loss calculation (step 55). These steps will be explained in detail below.

[0078] Step 51: The image rendering unit 120 reads all transmitted images and camera positions into the imgs and pos arrays of the algorithm to realize data input.

[0079] Step 52 is divided into three steps: setting up (e.g., adjusting focal length), converting pixels to rays (e.g., step 56), and obtaining position codes (e.g., step 58). Step 56 aims to determine the near and far positions of the image by inputting the distance of the image from the camera center. Step 57, converting pixels to rays, inputs the width, height, focal length, and extrinsic parameters (typically a 3x4 matrix used to map coordinates from the camera to world coordinates) of an image and outputs the rays containing all pixels of the image. Step 58, obtaining position codes, aims to acquire the 5D information of the rays.

[0080] Step 53: The image rendering unit 120 performs deep neural learning training on the NeRF model using the training set, making the NeRF algorithm generate data faster and with higher accuracy.

[0081] Step 54 is divided into coordinate transformation step 59 and rendering step 510. Coordinate transformation step 59 is used to convert the image coordinates of the image into real-world coordinates and transmit this information to rendering step 510. Rendering step 510 can sample the points on the rays generated in step 57 to obtain their real coordinates, and then render them using these coordinates, transforming 5D information into realistic model information.

[0082] Step 55: The image rendering unit 120 calculates the difference between the transformed model and the original model, that is, calculates the loss for training.

[0083] In one embodiment, the image display unit 160 requires a collimating backlight module 110, see details below. Figure 6 The structure within the design. The following section will elaborate on the alignment backlight design scheme.

[0084] Light source L1 provides a light source for the collimating backlight module. Extinction structure L2 uses a triangular structure to eliminate surface reflections. Display screen L3 displays directional rendered images. Imaging plane L4 projects images onto this plane. Compound lens array L5 displays stereoscopic effects.

[0085] In one embodiment, this application discloses an image rendering method, such as... Figure 7 As shown, image generation is achieved using neural radiation field rendering, employing techniques such as development-side image rendering and algorithm porting. These will be explained in detail below.

[0086] Step 610: Acquire images of the 3D scene to obtain scene images.

[0087] Step 620: Reconstruct the 3D scene and render the light field using the scene images to obtain the light field model.

[0088] Step 630: Convert the light field model to a neural network framework type to obtain a converted model, which can be deployed to the lower-level machine.

[0089] Step 640: Track and obtain the viewer's location information, use the location information to perform directional light rendering on the transformation model, and obtain a directional rendering image. This directional rendering image is used for three-dimensional projection.

[0090] In one specific embodiment, in step 610 above, image data containing scene information can be collected from different perspectives. These scene images can be real scene images captured by cameras or other devices, or synthetic images generated by computer graphics technology.

[0091] In one specific embodiment, step 620 above, which involves reconstructing the 3D scene and rendering the light field using the scene image to obtain a light field model, includes:

[0092] (1) The three-dimensional scene is reconstructed using the neural radiation field to obtain the scene model.

[0093] (2) Use the scene model to perform multi-view light field rendering on the scene image, generate the light field image of the three-dimensional scene and construct the corresponding light field model; wherein, when performing multi-view light field rendering, the scene model can learn the emissivity of the three-dimensional scene and the direction of light propagation through the scene image.

[0094] Understandably, in step 620, the collected data can be used to train the NeRF neural network, which will learn the radiance and the direction of light propagation in the scene.

[0095] In one specific embodiment, step 630 above involves converting the light field model to a neural network framework type to obtain a converted model, including:

[0096] (1) Obtain the first neural network framework type based on the light field model. The first neural network framework type is configured as a PyTorch framework and is used to enable the light field model to run on the host computer.

[0097] (2) Convert the light field model from the first neural network framework type to the second neural network framework type to obtain the corresponding conversion model. The second neural network framework type is configured as the ONNX framework and is used to enable the conversion model to run on the lower-level machine.

[0098] It can be understood that step 630 can be viewed as a PyTorch framework to ONNX conversion step, and an ONNX to mobile framework conversion step, used to achieve cross-platform deployment and inference of the model. ONNX is an open deep learning model exchange format that can perform model conversion and migration between different deep learning frameworks. By converting the PyTorch model to ONNX format, seamless model migration between different frameworks can be achieved, thereby realizing cross-platform deployment.

[0099] In one specific embodiment, step 640 above, tracking and obtaining the viewer's location information, and using the location information to perform directional ray rendering on the conversion model to obtain a directional rendered image, includes:

[0100] (1) The location and viewing angle of the viewer are determined by human eye tracking, forming location information.

[0101] (2) Using orientation information as input parameters, the transformation model is subjected to directional ray rendering to obtain a directional rendering image of the three-dimensional scene; among which, when performing directional ray rendering on the transformation model, the input parameters of the transformation model also include focal length and resolution.

[0102] It is understandable that in step 640, a trained NeRF model can be used, and the viewer's position and orientation can be input. The neural network will then output the corresponding directional rendering image.

[0103] In one embodiment, the digital sand table system and image generation method disclosed above are combined to control the display of the digital sand table system, such as... Figure 7 As shown below, the control process will be explained in detail. In actual use, the specific operation is as follows.

[0104] Step 71: The digital sand table system begins operation.

[0105] Step 72: Construct the 3D model that will be displayed in stereo.

[0106] Step 73: Acquire images of the 3D model from multiple angles to facilitate the generation of scene images.

[0107] Step 74: Train the NeRF model so that it can render high-quality light field images.

[0108] Step 75: Use the NeRF model to perform light field rendering on the scene image.

[0109] Step 76: Obtain the light field image and construct the corresponding light field model based on the rendering results.

[0110] Step 77: Convert the light field model and deploy it to the lower-level machine via data transmission.

[0111] Step 78: Determine the viewer's location information using eye tracking.

[0112] Step 79: On the lower-level machine, using the conversion model, input the orientation information and perform directional rendering based on the viewer's orientation to obtain a directional rendering image.

[0113] Step 710: Display the directional rendering image on the display screen, and display the image as a stereoscopic image through a compound lens array, thereby forming a digital sand table.

[0114] It should be noted that this digital sand table system is based on cutting-edge technologies such as neural radiation fields and stereoscopic light field display. Addressing the shortcomings of current digital sand table technology, it essentially provides a feasible minimum system for an intelligent digital sand table, which can drive the advancement of third-generation display technology. Furthermore, this digital sand table system allows for dynamic interaction and updates, offers excellent immersive stereoscopic experience, and supports simultaneous use and communication by multiple users. For example, commanders can use a mouse, keyboard, gestures, a baton, or other input devices to simulate the battlefield environment and situation in real time.

[0115] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. For those skilled in the art, several simple deductions or substitutions can be made without departing from the inventive concept of this application.

Claims

1. A digital sand table system employing image rendering, characterized in that, include: The image acquisition unit is used to acquire images of the 3D scene, obtain scene images, and transmit them to the host computer. An image rendering unit, running on the host computer, is used to acquire scene images from the image acquisition unit, and to reconstruct and render the 3D scene using the scene images to obtain a light field model. The image rendering unit includes a model training module and a light field image generation module. The model training module is used to perform deep learning using a training set, combining NeRF with memory neural radiation fields, and reconstructing the 3D scene using the neural radiation fields to obtain a scene model. The light field image generation module is used to transmit the acquired scene images to a neural network using the scene model, generate light field images of the 3D scene through multi-view light field rendering, and construct the corresponding light field model. The computational framework of the light field image generation module is PyTorch, meaning that the light field model has a first neural network framework type adapted to the host computer. A data transmission unit, running on the host computer, is used to convert the light field model into a neural network framework type, obtain a converted model, and deploy it to the lower-level computer. The data transmission unit includes a neural network framework transfer module and a distributed Ethernet port module. The neural network framework transfer module is used to convert the PyTorch computing framework into a mobile framework using an ONNX structure to obtain the corresponding converted model, enabling cross-platform deployment and inference of the light field model while preserving its accuracy and performance. The computing framework of the neural network framework transfer module is a mobile framework, meaning the converted model has a second neural network architecture type adapted to the lower-level computer. The distributed Ethernet port module is used to transmit the converted model to the lower-level computer using network transmission technology for image rendering on the lower-level computer. A directional rendering unit, running on the lower-level machine, is used to track and obtain the viewer's location information, and to use the location information to perform directional ray rendering on the conversion model to obtain a directional rendering image. An image display unit, communicatively connected to the lower-level machine, is used to project the directional rendering image in three dimensions to form a digital sand table, the digital sand table including a stereoscopic image of the three-dimensional scene facing the viewer.

2. The digital sand table system as described in claim 1, characterized in that, The directional rendering unit includes an eye-tracking module, a light field image rendering module, and an HDMI transmission module; The human eye tracking module is used to determine the viewer's position and viewing angle through human eye tracking, forming the orientation information; The light field image rendering module receives the conversion model and uses the orientation information as input parameters to perform directional light rendering on the conversion model to obtain a directional rendering image of the three-dimensional scene. The HDMI transmission module is used to transmit the directional rendering image to the image display unit using the HDMI protocol.

3. The digital sand table system as described in claim 2, characterized in that, The image display unit includes a collimation backlight module, a screen display module, and a compound lens array module; The collimation backlight module is used to provide collimation backlight for the screen display module; The screen display module is used to receive the directional rendering image transmitted by the HDMI transmission module, and to display the directional rendering image. The composite lens array module is used to focus the light from each pixel in the image of the screen display module, and the resulting three-dimensional projection forms the digital sand table.

4. The digital sand table system as described in claim 3, characterized in that, The collimated backlight module has an extinction structure, which eliminates stray light when light is projected onto the screen display module, thereby reducing crosstalk in the display and improving the contrast and clarity of the image.

5. The digital sand table system as described in claim 1, characterized in that, The image acquisition unit is used to acquire video of the three-dimensional scene, and uses video frame extraction technology to extract frame images from the video of the three-dimensional scene to obtain images of the three-dimensional scene from multiple perspectives and use them as the scene images.

6. A method for rendering light field images, characterized in that, include: Acquire images of the 3D scene to obtain scene images; The process involves reconstructing and rendering the 3D scene using the scene image to obtain a light field model. This includes using a training set for deep learning, combining NeRF with a memory neural radiation field, and reconstructing the 3D scene using the neural radiation field to obtain a scene model. The process also involves using the scene model to transmit the acquired scene image to a neural network, generating a 3D scene light field image through multi-view light field rendering, and constructing a corresponding light field model. The light field model has a first neural network framework type adapted to the host computer, namely the PyTorch framework. The light field model is converted to a neural network framework type to obtain a converted model. This converted model can be deployed to a lower-level machine. This includes converting the PyTorch computing framework into a mobile framework using the ONNX structure to obtain the corresponding converted model. This enables cross-platform deployment and inference of the light field model while preserving its accuracy and performance. In other words, the converted model has a second neural network architecture type adapted to the lower-level machine, namely a mobile framework. The converted model is then transmitted to the lower-level machine using network transmission technology for image rendering. The location information of the viewer is tracked and obtained. The location information is used to perform directional light rendering on the transformation model to obtain a directional rendering image. The directional rendering image is used for three-dimensional projection.

7. The light field image rendering method as described in claim 6, characterized in that, The process of tracking to obtain the viewer's location information, and using this location information to perform directional ray rendering on the transformation model to obtain a directional rendered image, includes: The location and viewing angle of the viewer are determined by eye tracking, forming the orientation information; Using the orientation information as input parameters, directional ray rendering is performed on the transformation model to obtain a directional rendering image of the three-dimensional scene; When performing directional ray rendering on the conversion model, the input parameters of the conversion model also include focal length and resolution.

Citation Information

Patent Citations

  • Light field display system based on human eye tracking and image rendering method

    CN117319631A