High-resolution light field rapid rendering method and rendering device

Through color space separation and subpixel light field rendering technology, combined with neural network optimization, the problems of large data demand and high computational complexity of traditional light field rendering are solved, high-resolution light field rendering is achieved, and rendering speed and image quality are improved.

CN120259524APending Publication Date: 2025-07-04NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312898.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional light field image rendering technology has problems such as large data demand, high computational complexity and limited resolution, which leads to certain limitations in practical applications.

Method used

By acquiring the scene images of the target scene, color space separation is performed, grayscale images and RGB signals are generated, light information is calculated and multi-dimensional information is generated, and sub-pixel light field image rendering technology is used to optimize the network model with neural networks to perform high-resolution light field rendering.

Benefits of technology

It improves the speed and resolution of light field rendering, reduces computational complexity, enhances the detail presentation and visual effects of images, and is suitable for scenes such as virtual reality and augmented reality that require real-time rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259524A_ABST
    Figure CN120259524A_ABST
Patent Text Reader

Abstract

The invention discloses a high-resolution light field rapid rendering method and rendering device, and compared with a previous image rendering method, the method and device change an image processing mode and information generation and query dimensions. The method comprises the following steps: firstly, performing color space separation processing on an image, separating a grayscale image from an RGB (Red, Green and Blue) signal, and respectively entering different networks for rendering, so that the rendering speed can be improved on the premise of keeping details and textures and fitting the visual characteristics of human eyes; secondly, wavelength dimension information is introduced into image reconstruction, the network structure is simplified, a network model can share the weight more effectively to optimize the memory and the processing efficiency, and meanwhile, the expression ability of the model in different spectral ranges is enhanced; and finally, sub-pixel light field image rendering is introduced, so that the system can capture information at more angles, the resolution and details of the image are enhanced, and a high-resolution image which can be observed at different visual angles is generated in a visual space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer image processing, and in particular to a high-resolution light field fast rendering method and a rendering device. Background Art

[0002] Light field image rendering is a cutting-edge computer graphics technology that achieves realistic image synthesis and multi-perspective presentation by capturing and reconstructing the information of every light in the scene. Light field technology can directly reproduce the propagation of light in space, thereby providing more realistic visual effects and free perspective experience. Traditional light field technology captures and stores the direction and position of all light in the scene, combines stereo reconstruction and training convolutional neural networks, and relies on voxels, point clouds and other solutions to model three-dimensional scenes. When the resolution requirement is high, the storage cost and computing power requirements of voxels and point clouds will show a cubic growth, so the application scenarios are greatly limited, and the actual model precision often does not achieve the ideal effect.

[0003] Currently, more efficient means of rendering light field images are to perform model prediction or image-based 3D expression of the scene. The most famous method, NeRF (Neural Radiance Fields), combines implicit representation with volume rendering. It calculates the true position of the light corresponding to each pixel in the image in space based on the input camera pose, samples the coordinates of the query point on the light, and uses the coordinates as input parameters to query the MLP model. In this way, NeRF can reconstruct a dense light field of a scene and generate images from different perspectives. This method inherits the basic principles of light field rendering, and also improves rendering quality through deep learning models, reduces the storage requirements for raw light field data, and improves the ability to handle complex scenes.

[0004] At the same time, in order to solve the problem that traditional light field rendering is limited by resolution, sub-pixel technology has been introduced into light field image rendering. This technology further subdivides each pixel into multiple sub-pixels to achieve more precise light calculation and presentation. It can not only improve the spatial resolution of the image, but also generate smoother transition effects under different viewing angles, making the images seen in all directions within the viewing cone more natural and delicate.

[0005] In general, traditional light field image rendering technology has defects such as large data requirements, high computational complexity and resolution limitation, which makes traditional light field technology have certain limitations in practical applications. In recent years, neural network technologies such as NeRF and sub-pixel technology have gradually attracted attention, trying to solve the shortcomings of these traditional light field rendering methods. Summary of the invention

[0006] The technical problem solved by this application is: how to perform light field rendering on an image and improve the rendering speed and resolution of light field rendering. To solve this problem, this application provides a high-resolution light field fast rendering method and a rendering device.

[0007] According to the first aspect, this application provides a high-resolution light field fast rendering method, including: obtaining a scene image of a target scene and performing color space separation processing to obtain a grayscale image and an RGB signal of the target scene; obtaining detection data of a target object in the target scene according to the grayscale image, and calculating light information by using the detection data; generating multi-dimensional information of the target object according to the light information, where the multi-dimensional information includes a spatial position, a detection direction, and a light wavelength; calculating light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively based on the multi-dimensional information; and performing sub-pixel light field image rendering by using the light intensity information corresponding to the target object in the grayscale channel and the light intensity information in the RGB channel to obtain a high-resolution three-dimensional image of the target object at multiple angles.

[0008] Further, the performing color space separation processing on the target scene image includes: separating the scene image according to the channel dimension to obtain three independent channel output data of red, green, and blue, and forming an RGB signal; and generating the grayscale image of the target scene by calculating a weighted average value of the three independent channel output data of red, green, and blue.

[0009] Further, obtaining detection data of a target object in the target scene according to the grayscale image, and calculating light information by using the detection data includes: identifying the target object in the target scene from the grayscale image, obtaining the detection data of the target object, and determining each imaging light point corresponding to the detection data in the camera coordinate system; calculating the light direction and color corresponding to each imaging light point respectively, and forming the light information based on the light direction and color corresponding to each imaging light point respectively.

[0010] Further, generating multi-dimensional information of the target scene according to the light information includes: analyzing, by using the light direction and color corresponding to each imaging light point in the light information, to obtain the coordinate values of each feature point on the target object in three-dimensional space, represented by x, y, and z, and further obtaining the angle values of each feature point relative to the camera, represented by θ and φ, and obtaining the light wavelength value of each feature point, represented by λ; forming the spatial position of each feature point on the target object by the coordinate values x, y, and z, forming the detection direction of each feature point by the angle values θ and φ, and forming the light wavelength of each feature point by the light wavelength value λ; and forming the multi-dimensional information based on the spatial position, detection direction, and light wavelength of each feature point on the target object.

[0011] Further, calculating the light intensity information corresponding to the target object in the grayscale channel and the RGB channel based on the multi-dimensional information includes: performing high-dimensional encoding on the multi-dimensional information to obtain encoding information corresponding to the spatial position, detection direction, and light wavelength of each feature point on the target object; calculating the light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively using the encoding information; wherein, the high-dimensional encoding refers to performing a high-dimensional mapping of the Fourier characteristics on the spatial position, detection direction, and light wavelength of each feature point, and obtaining the corresponding encoding information using the high-dimensional mapping result.

[0012] Further, calculating the light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively using the encoding information includes: inputting the encoding information of the spatial position, detection direction, and light wavelength of each feature point into a preset light field chromatic aberration network to generate the intensity information and color information of the light, and forming the light intensity information in the RGB channel; inputting the encoding information of the spatial position and detection direction corresponding to each feature point into a preset light field intensity network to generate the intensity information of the light, and forming the light intensity information in the grayscale channel.

[0013] Further, after generating the light intensity information of each feature point using the multi-dimensional information, it further includes: comparing to obtain the error information between the light intensity information of each feature point and the corresponding true light intensity value, where the error information includes PSMR value and / or mean square error value; adjusting the network parameters of the neural network according to the error information to optimize the network model corresponding to the neural network through negative feedback, and the neural network is the light field chromatic aberration network or the light field intensity network.

[0014] Further, performing sub-pixel light field image rendering using the light intensity information corresponding to the target object in the grayscale channel and the light intensity information in the RGB channel to obtain a high-resolution three-dimensional image of the target object at multiple angles, includes: irradiating a liquid crystal panel that has undergone sub-pixel processing through backlight, calculating the light intensity value using the light intensity information corresponding to the target object in the grayscale channel and the light intensity information in the RGB channel, and driving the sub-pixel points on the liquid crystal panel to emit light according to the calculated light intensity value, and the light converges into points after passing through the microstructure array filled with lenses, forming a high-resolution three-dimensional image of the target object at multiple angles in the visual space.

[0015] According to a second aspect, the present application provides a high-resolution light field fast rendering device, including: a color space separation module, configured to obtain a scene image of a target scene and perform color space separation processing on the scene image according to the channel dimension to obtain a grayscale image and an RGB signal of the target scene; a light ray generation module, connected to the color space separation module, configured to obtain detection data of a target object in the target scene according to the grayscale image, and obtain light ray information by using the detection data; an information generation module, connected to the light ray generation module, configured to generate multi-dimensional information of the target object according to the light ray information, where the multi-dimensional information includes a spatial position, a detection direction, and a light ray wavelength; an information query module, connected to the information generation module, configured to calculate light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively according to the multi-dimensional information; a light field image rendering module, connected to the information query module, configured to perform sub-pixel light field image rendering by using the light ray intensity information corresponding to the target object in the grayscale channel and the light intensity information in the RGB channel to obtain a high-resolution three-dimensional image of the target object at multiple angles.

[0016] Further, the 6D light field image high-resolution fast rendering device further includes an error calculation module and a model optimization module; the error calculation module is signal-connected to the light ray rendering module, configured to compare and obtain error information between the calculated light intensity information and the true light intensity value, where the error information includes a PSMR value and / or a mean square error value; the model optimization module is signal-connected to the error calculation module and the information query module, configured to adjust network parameters of a neural network according to the error information to optimize a network model corresponding to the neural network through negative feedback; the neural network is a light field color difference network for generating light intensity information in the RGB channel, or a light field intensity network for generating light intensity information in the grayscale channel.

[0017] The beneficial effects of the present application are as follows:

[0018] According to the above high-resolution light field fast rendering method and rendering device, since the dimensions of information generation and query are changed during the rendering process, by introducing the light ray wavelength parameter, the network can not only uniformly process the information of the three RGB channels, reduce the computational complexity, but also more accurately capture the characteristics of light rays at different wavelengths to achieve hyperspectral image rendering; secondly, the strategy of high-resolution reconstruction of the grayscale image and low-resolution reconstruction of the RGB signal is adopted, which not only ensures the accurate presentation of the scene contour and details, but also reduces the burden of processing multi-channel data and optimizes the rendering efficiency.

[0019] The technical solution of this application optimizes the network performance by designing a hierarchical network for the neural radiance field, reasonably allocating the processing tasks of high-resolution grayscale images and low-resolution RGB signals to different network structures, thereby greatly reducing the computational burden while maintaining the image quality. The technical solution of this application performs deep learning by extracting information from images, generates the corresponding rays in space for each pixel in the image according to the camera parameters, extracts features of the rays through training the MLP model and refines the light intensity distribution in the scene, then generates the light intensity values ray by ray according to the query information, and finally performs light field image rendering. It can be understood that the technical solution of this application not only calculates the light intensity in grayscale and RGB images by adding the light wavelength variable to the input information and combining with the NeRF network, but also realizes end-to-end training, directly training from the input data to the final output without manually designing intermediate steps, designing feature extraction or adding additional components. Furthermore, it compresses the number of generated rays, speeds up the training speed, and implicitly stores the light intensity information in the form of neural network weights. Due to the continuity characteristics of storing weights and functions, a large amount of information can be implanted with a very small storage space, greatly improving the rendering speed.

[0020] The technical solution of this application introduces sub-pixels for light field image rendering, enabling the system to capture the details of the image more precisely, which has obvious positive significance for scenes with complex textures or obvious light and shadow changes, and enables the image to maintain consistent high-quality display from different perspectives in the visual space. At the same time, sub-pixel light field rendering provides high flexibility for viewing angle changes because the rendering process considers the light propagation from different perspectives, enabling the image seen by the observer in the visual space to change naturally according to the movement of the viewing angle, enhancing the sense of three-dimensionality and immersion, and is suitable for scenarios such as virtual reality and augmented reality that require real-time rendering. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic flowchart of a high-resolution light field fast rendering method in an embodiment of this application;

[0022] Figure 2 It is a schematic flowchart of a high-resolution light field fast rendering method in another embodiment of this application;

[0023] Figure 3 It is a schematic diagram of the principle of high-dimensional mapping of multi-dimensional information in an embodiment of this application;

[0024] Figure 4 It is a schematic diagram of the network architecture of the light field chromatic aberration network in an embodiment of this application;

[0025] Figure 5 It is a schematic diagram of the network architecture of the light field intensity network in an embodiment of this application;

[0026] Figure 6 Schematic diagram of the principle of sub-pixel light field rendering in an embodiment of the present application;

[0027] Figure 7 Schematic diagram of light field rendering without sub-pixel processing in an embodiment of the present application;

[0028] Figure 8 Schematic diagram of light field rendering with sub-pixel processing in an embodiment of the present application;

[0029] Figure 9 Flowchart of model training of a preset neural network in an embodiment of the present application;

[0030] Figure 10 Schematic diagram of the structure of a high-resolution light field fast rendering device in an embodiment of the present application;

[0031] Figure 11 Schematic diagram of the structure of a high-resolution light field fast rendering device in another embodiment of the present application. Detailed implementation manners

[0032] The present application will be further described in detail below with reference to the accompanying drawings through specific implementation manners.

[0033] In an embodiment of the present application, a high-resolution light field fast rendering method is provided. As Figure 1 shown, it mainly includes steps S11 to S16.

[0034] Step S11: Obtain the scene image of the target scene and perform color space separation processing to obtain the grayscale image and RGB signals of the target scene. It can be understood here that the image is separated according to the channel dimension to achieve the separation of the grayscale image and RGB signals of the target scene.

[0035] In a specific implementation, the color space separation processing of the target scene image includes the following process:

[0036] (1) Load the target scene image in the form of a digital matrix, usually with a size of (height, width, 3), where each layer of the third dimension corresponds to a different color channel. By extracting this layer, the values of each pixel in the R, G, and B channels can be obtained respectively.

[0037] (2) Weight and combine the information of the R, G, and B channels into a single grayscale value. The added weights reflect the sensitivity of the human eye to different colors, where it is most sensitive to green and least sensitive to blue. A common grayscale weighting formula is:

[0038] I gray= 0.299×R + 0.587×G + 0.114×B。

[0039] Step S12: Obtain the detection data of the target object in the target scene from the grayscale image, and calculate the light information using the detection data. It can be understood that the light direction and color of the target object imaging in the target scene are respectively generated. Here, the target object can be a physical object or an animal in the environment, and the specific object is not limited.

[0040] In a specific implementation, obtaining the detection data of the target object in the target scene from the grayscale image and calculating the light information using the detection data includes the following process:

[0041] (1) Identify the target object in the target scene from the grayscale image, obtain the detection data of the target object, and determine each imaging light point corresponding to the detection data in the camera coordinate system. It can be understood that for the detection camera, each imaging light point is the mapping of the light rays radiated by a feature point on the target object on the imaging plane of the camera. In the camera coordinate system of the detection camera, the detection data reflects the imaging results of each corresponding imaging light point on the imaging plane.

[0042] (2) Calculate the light direction and color corresponding to each imaging light point respectively, and form the light information based on the light direction and color corresponding to each imaging light point respectively. For example, the light direction and color of the imaging light point can be calculated according to the internal parameter K and external parameter c2w of the detection camera, and even the light center coordinate information can be obtained. The light direction, color, and light center coordinate information can be obtained through matrix and coordinate operations based on the internal and external parameters of the camera.

[0043] Step S13: Generate multi-dimensional information of the target object in different color spaces. The multi-dimensional information mentioned here includes spatial position, detection direction, and light wavelength. This step can be understood as the generation step of the query direction, coordinates, and wavelength. Based on the obtained light direction, color, and even the light center coordinate, the scene boundary is inferred, and then the query direction, coordinates, and light wavelength are generated.

[0044] In a specific embodiment, generating the multi-dimensional information of the target object according to the light information includes the following process:

[0045] (1) Use the light direction and color corresponding to each imaging light point in the light information to analyze and obtain the coordinate values of each feature point on the target object in the three-dimensional space, represented by x, y, and z, and also obtain the angular values of each feature point relative to the camera, represented by θ and φ, and obtain the light wavelength value of each feature point, represented by λ. Since the grayscale channel only includes the brightness information of the target object and no additional processing of the wavelength information is required, the light wavelength λ of each feature point is not formed in this color space.

[0046] (2) The coordinate values x, y, and z form the spatial positions of each feature point on the target object, the angular values θ and φ form the detection directions of each feature point, and the light wavelength λ forms the light wavelength of each feature point. θ and φ in the detection direction can be understood as the deflection angles of the target point relative to the detection point in the horizontal and vertical directions.

[0047] (3) Based on the spatial positions (x, y, z), detection directions (θ, φ), and light wavelengths λ of each feature point on the target object, multi-dimensional information in different color spaces is formed.

[0048] It can be understood that the multi-dimensional information includes 6 data volumes, so it can also be called 6D information.

[0049] Step S14, calculate the light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively based on the multi-dimensional information. It can be understood as the sampling information generation step. Specifically, the multi-dimensional information in different color spaces is input into the neural network model, and the light intensity information at the sampling points is obtained through neural network calculation.

[0050] In a specific embodiment, calculating the light intensity information of the target object based on the dimensional information includes the following process:

[0051] (1) Perform high-dimensional encoding on the multi-dimensional information to obtain the encoded information corresponding to the spatial positions, detection directions, and light wavelengths of each feature point on the target object. Specifically, refer to Figure 3 , perform high-dimensional mapping 34 with Fourier characteristics on the spatial position 31, detection direction 32, and light wavelength 33 of each feature point in the multi-dimensional information, and use the high-dimensional mapping results (i.e., the distinguishable high-dimensional information 35) to obtain the corresponding encoded information. It can be understood that for the reason of making the preset neural network better extract information, the high-dimensional information is encoded by the encoder, and an L value is set for the coordinates, directions, and wavelengths respectively, and the following encoding method is constructed:

[0052] γ(p) = (sin(2 0 πp), cos(2 0 πp), …, sin(2 L-1 πp), cos(2 L-1 πp))

[0053] According to the above encoding method, the high-dimensional mapping of the Fourier characteristics of the information can be completed, where p represents the data to be processed.

[0054] (2) Input the encoding information corresponding to the spatial position of each feature point into the front network layer of the preset light field chromatic aberration network, and input the encoding information corresponding to the detection direction and the encoding information corresponding to the light wavelength into the rear network layer of the preset light field chromatic aberration network. Use the output results of the network layer to obtain the intensity information and color information of the light of any pixel in the RGB channels, and form the light intensity information in the RGB channels. The preset light field chromatic aberration network can refer to Figure 4 , the front network layer (Backbone, marked 42) is used to process the input spatial position encoding information (Input, marked 41), extract geometric features from it, and the rear network layer (Direction&Wavelength, marked 43) is used to receive and fuse the encoding information of the detection direction and the light wavelength. Finally, through the output layer of the network (marked 44), output the intensity information and color information of the corresponding light for any pixel, and form the light intensity information of this pixel in the RGB channels.

[0055] (3) Input the encoding information corresponding to the spatial position of each feature point into the front network layer of the preset light field intensity network, and input the encoding information corresponding to the detection direction into the rear network layer of the preset light field intensity network. Use the output results of the network layer to obtain the intensity information of the light of any pixel, and form the light intensity information in the grayscale channel. The preset light field intensity network can refer to Figure 5 , the front network layer (Backbone, marked 52) is used to process the input spatial position encoding information (Input, marked 51), extract geometric features from it, and the rear network layer (Direction, marked 53) is used to receive and fuse the encoding information of the detection direction. Finally, through the output layer of the network (marked 54), output the intensity information of the corresponding light for any pixel, and form the light intensity value of this pixel in the grayscale channel.

[0056] It can be understood that in Figure 4 the schematic light field chromatic aberration network and Figure 5 the schematic light field intensity network shown, the coordinate information corresponding to the spatial position is sent into the backbone network shown by Backbone (marked 42, 52), and the coordinate information is spliced here to obtain intermediate data, that is, the geometric features of the target object; in Figure 4 , the intermediate data is sent into the single-layer network shown by Direction&Wavelength (marked 43), and the light intensity value I R in the RGB channels is calculated by splicing the direction information corresponding to the detection direction and the wavelength information corresponding to the light wavelength. G and I B . And in Figure 5Among them, the intermediate data is sent into the single-layer network shown by Direction (mark 53), and after splicing the direction information corresponding to the detection direction, the light intensity value I under the same gray-scale combination is calculated. g

[0057] It can be understood that for Figure 4 the light field color difference network, the depth of its front network layer is 4 layers, while for Figure 5 the light field intensity network in it, the depth of its front network layer is 6 layers. This is designed according to the human eye visual characteristics: the human eye is very sensitive to the change of brightness (i.e., gray scale). The deep design of the light field intensity network helps to capture more details; while the human eye is not sensitive to the change of color (i.e., RGB). Therefore, for the light field color difference network, using a shallower depth simplifies the processing process. Such a design optimizes the calculation efficiency and at the same time maintains the required level of detail in the gray-scale and RGB outputs.

[0058] It can be understood that for Figure 4 the light field color difference network in it and Figure 5 the light field intensity network in it, the network width is 128 in both cases. In this way, the complexity of the neural network structure can be significantly reduced, so that the model rendering time can be greatly reduced while ensuring the image accuracy, and the improvement of the running speed is also extremely significant.

[0059] Step S15, using the light intensity information corresponding to the target object in the gray-scale channel and the light intensity information in the RGB channel to perform sub-pixel light field image rendering, and obtaining high-resolution three-dimensional images of the target object at multiple angles. Specifically, refer to Figure 6 , irradiate the liquid crystal panel (mark 62) after sub-pixel processing through backlight (mark 61), calculate the light intensity value using the light intensity information corresponding to the target object in the gray-scale channel and the light intensity information in the RGB channel, and drive the liquid crystal panel to be composed of RGB sub-pixels and gray-scale sub-pixels. The sub-pixels emit light according to the calculated light intensity. After the light passes through the microarray structure filled with lenses (mark 63), it converges into multiple visual points (such as A and B in the figure) in the image space (mark 64). Finally, the light enters the visual space (mark 65), and thus high-resolution three-dimensional images of the target object at multiple angles can be formed in the visual space.

[0060] It can be understood that using a liquid crystal panel after sub-pixel processing can improve the imaging resolution. Due to the capacity limitation of the photosensitive element itself, each pixel only represents the nearby color, so there is usually a spacing of 4-5 microns between pixels. Specifically, refer to Figure 7 , the light emitted by the pixels p1, p2, and p3 without sub-pixel processing forms visual cones S1, S2, and S3 respectively after passing through the lens L. Refer to Figure 8, software approximate calculation processes pixels into several sub-pixels (i.e., RGB sub-pixels and grayscale sub-pixels). The light emitted by R1 and R2 forms visual cones S R1 and S R2 , respectively, after passing through lens L. The light emitted by G1 and G2 forms visual cones S G1 and S G2 , respectively, after passing through lens L. The light emitted by B1 and B2 forms visual cones S B1 and S B2 , respectively, after passing through lens L. The light emitted by g1 and g2 forms visual cones S g1 and S g2 . Figure 8 The visual cones generated after sub-pixel processing in Figure 7 are denser than the visual cones generated without sub-pixel processing. For details, refer to Figure 8 . The visual cones S R1 , S g2 and S B2 generated after sub-pixel processing are denser than the visual cone S2 generated without sub-pixel processing. The density of the visual cones can reflect the resolution of reconstructing the target object in the visual space. The denser the imaging area, the more continuous the imaging of the target object during reconstruction in the visual space, more scene details can be restored, and blurring or jaggedness can be reduced, resulting in a higher resolution of the reconstructed three-dimensional image.

[0061] It can be understood that during sub-pixel processing, it is not limited to the traditional division of pixels into RGB three channels, but combines the RGB channels with grayscale values for processing, which can more accurately restore the lighting changes and details in the scene. This processing method enables color information to be processed separately without affecting the light intensity, while grayscale information enhances the ability to capture scene details and depth, greatly improving the clarity and visual realism of the rendering results. This processing method not only improves the resolution of the image but also enhances the adaptability of light field rendering to complex lighting scenes, and can better present rich colors and details.

[0062] It can be understood that dense visual cones can form a high-resolution three-dimensional image in the visual space for observers to view from various angles. When collecting the light of the target object, the light is reflected from the object surface and reaches the camera's sensor. According to the reversibility of the light path, light is "emitted" from the pixel point and "returns" to the object surface along the original path to achieve reverse light field reconstruction. In Figure 8 , multiple visual cones (labeled as S R1etc.) respectively represent the process of light propagating in different directions after entering the lens L from different angles. Each cone represents a unique viewing angle, and each cone in the image corresponds to a sub-pixel (labeled as R1, etc.) at its bottom. This indicates that the light emitted from sub-pixels at different positions will converge after passing through the lens and be projected to different positions in the visual space at different angles. Since each cone covers a different direction, when looking at the same point from different viewing angles, the received light is different. This difference in the direction of light directly leads to different image contents when observed from different angles, and this characteristic is precisely the manifestation of the anisotropy of the light field. The light field image rendering technology records the direction and angle information of these lights, allowing the observer to view the lights emitted from different directions through the change of the viewing angle in the later stage, so as to form high-resolution three-dimensional images of the target object from multiple angles in the visual space.

[0063] It can be understood that in the process of sub-pixel processing, it is necessary to flexibly adjust the separation ratio of sub-pixels according to the specific scene and rendering effect. According to the rendered image effect, dynamically adjust the weight distribution of color (RGB) and light intensity (gray scale), and find the most suitable ratio scheme to achieve the best visual presentation. For example, in some scenes, it may be necessary to increase the weight of gray scale to enhance the performance of details, while in other scenes, the color information may be emphasized preferentially to improve the realism of visual colors. This strategy of flexibly adjusting the ratio according to the effect enables the sub-pixel separation scheme to adapt to various complex rendering environments, thereby maximizing the accuracy and realism of the light field image.

[0064] In another embodiment, the three-dimensional rendering method may further include, in addition to the steps S11 - S15 mentioned above, Figure 2 the steps S16 and S17 shown in

[0065] Step S16, located between step S14 and step S15, after generating the light intensity information of each feature point using multi-dimensional information, it is also necessary to compare the error information between the light intensity information of each pixel and the corresponding real light intensity value. The error information includes the PSMR value and / or the mean square error value.

[0066] It can be understood that step S16 can be understood as calculating the difference between the predicted light intensity value output by the network and the target light intensity value. This error measures the deviation degree of the current network's prediction result from the real scene light intensity. For example, using the mean square error (MSE) for calculation, that is:

[0067]

[0068] where N is the number of lights or pixels; I pred,i is the predicted light intensity of the i-th light; Itarget,i is the target light intensity of the i-th ray.

[0069] Step S17: Obtain the error information from step S16, and adjust the network parameters of the neural network according to the error information to optimize the network model corresponding to the neural network through negative feedback. The neural network is a light field chromatic aberration network or a light field intensity network.

[0070] It can be understood that step S17 can be regarded as a process of optimizing and updating the model weights. If the error information exceeds a certain threshold, the weight information in the network model corresponding to the neural network is adjusted, and then the calculation of the ray intensity information is restarted from step S14.

[0071] It should be noted that in the Figure 2 process of the process flow, since step S16 is added to calculate the error information, a specific judgment can be made in step S15. For example, when the error information is within a certain range, the light field image is rendered according to the ray intensity value.

[0072] In the 3D rendering method, it is necessary to calculate the light intensity of the target object based on multi-dimensional information. This calculation process uses the Figure 4 light field chromatic aberration network in Figure 5 and the Figure 9 light field intensity network in

[0073] Before using the preset neural network, it is also necessary to perform network training on it to obtain a trained network model. Refer to

[0074] Step S91: Preparation stage. It is necessary to configure the running environment of the network model, and prepare sample images and camera parameters, etc.

[0075] Step S92: Set all training parameters, such as the number of training epochs, expected threshold, noise, training starting point, etc.

[0076] Step S93: The program checks whether the training data exists, is complete, and matches according to the specified path, and then reads the file into memory and converts it into a suitable data type.

[0077] Step S94: The program generates a model class according to the specified parameters to implement the initialization of the network model.

[0078] Step S95: Check whether the pre-trained parameters are read in step S92, specifically whether there is weight data. If so, go to step S96; otherwise, go to step S97.

[0079] Step S97, in the case of no weight data, randomly initialize the weight data and load it into the network model.

[0080] Step S98, obtain the thermal infrared detection data of the target object and process it to obtain light information. For details, refer to step S11 above.

[0081] Step S99, generate multi-dimensional information of the target object based on the light information, and calculate the light intensity information of the target object based on the multi-dimensional information. For details, refer to steps S12, S14, and S15 above.

[0082] Step S910, compare to obtain the error information between the true light value and the corresponding rendered light intensity value of any pixel. The error information includes PSMR value and / or mean square error value. For details, refer to step S16 above.

[0083] Step S911, when the error information is less than the corresponding set value, or the number of training rounds is greater than the corresponding set value, then enter step S514; otherwise, enter step S513.

[0084] Step S912, adjust the network parameters according to the error information to optimize the network model corresponding to the preset neural network through negative feedback. For details, refer to step S17 above.

[0085] Step S913, at this time, it is considered that the training of the network model is basically qualified, and all parameters of the network model during training can be packaged into a compressed file and stored according to the specified path, mainly saving the weight data.

[0086] Step S914, release the occupied memory, clear the temporarily stored training data, and end the training program, that is, the training of the network model ends.

[0087] According to a high-resolution light field fast rendering method of the above embodiment, since the color space of the image is separated, and wavelength parameters are added to change the dimensions of information generation and query, the rendering speed is improved on the premise of conforming to the visual characteristics of the human eye, the memory and processing efficiency are optimized, and at the same time, the expression ability of the network model in different spectral ranges is enhanced. At the same time, through sub-pixel sampling, the distribution of the cones is densified, so that the light directions that can be recorded are more refined, and the details of the image are sampled and recorded at more angles, and finally a clearer and more delicate image can be obtained during reconstruction. And, through the optimization of the network model training method, the rendering time is greatly reduced, and the running speed is significantly accelerated. The following is a specific analysis.

[0088] For the set of target object images obtained by shooting and subjected to color space separation, first obtain the camera parameters of each image through methods such as colmap, and input the images and corresponding parameters into the trained network model. First, the detection data of the target object generates the light information (such as the optical center coordinates, light direction, color) of the light corresponding to each pixel point in the actual space; divide the blocks according to the image color characteristics, assign importance based on whether the block contains the target object, and generate a small number of lights for the blocks with low importance to compress the number of lights. In each round of training, randomly extract a fixed number of lights to generate multi-dimensional information of the target object, randomly sample a certain number of query points on the light, and calculate their coordinates, query directions, and wavelengths to the light (i.e., spatial positions, detection directions, and wavelengths). For the reason of enabling the network model to better extract information, encode the data through the encoder to obtain the encoded information with high-dimensional mapping. When using the network model to calculate the pixel information of the target object, send the coordinate information into the front network layer, and splice the coordinate information again at Add_Layer. After obtaining the intermediate data, fuse and splice the direction information (for the light field chromatic aberration network, it is also necessary to fuse and splice the wavelength information) and then calculate the light intensity value I i . When optimizing the network model, automatically backpropagate according to the calculated error information to modify each weight of the network, complete one round of training, and repeat the training until satisfied and then stop training. A trained model is required for sub-pixel light field image rendering. Import the model weights into the network model, first generate the imaging position and shape (camera parameters) of the image to be rendered, generate the light information corresponding to each pixel, repeat the above training steps for each light to obtain the light intensity, and then input the light intensity value into the sub-pixel light field image rendering device for high-precision reconstruction of three-dimensional multi-view images.

[0089] In one embodiment, based on the high-resolution light field fast rendering method mentioned above, a high-resolution light field fast rendering device is also disclosed. Please refer to Figure 10 , this rendering device mainly includes a color space separation module 101, a light generation module 102, an information generation module 103, an information query module 104, and a light field image rendering module 105, which will be described separately below.

[0090] The color space separation module 101 is used to obtain the scene image of the target scene and perform color space separation processing on the scene image according to the channel dimension to obtain the grayscale image and RGB signal of the target scene. Specifically, load the target scene image in the form of a digital matrix, usually with a size of (height, width, 3), where each layer of the third dimension corresponds to a different color channel. By extracting this layer, the values of each pixel in the R, G, and B channels can be obtained respectively; weight and combine the information of the R, G, and B channels into a single grayscale value, thereby realizing image color space separation.

[0091] The light generation module 102 is connected to the color space separation module 101 and is used to obtain the detection data of the target object in the target scene according to the grayscale image, and obtain the light information by using the detection data. Specifically, the light generation module 102 can obtain the detection data of the target object from the detection camera and determine each imaging light point corresponding to the detection data in the camera coordinate system of the detection camera. It can be understood that for the detection camera, each imaging light point is the mapping of a feature point light on the target object on the imaging plane of the detection camera, and the detection data in the camera coordinate system of the detection camera reflects the imaging results of each corresponding imaging light point on the imaging plane. In addition, the light generation module 102 can calculate the light direction and color corresponding to each imaging light point respectively and form them into light information. For example, the light direction and color of the imaging light point can be calculated according to the internal parameter K and external parameter c2w of the detection camera, and even the light center coordinate information can be obtained. The light direction, color and light center coordinate information can be obtained through matrix and coordinate operations based on the internal and external parameters of the camera.

[0092] The information generation module 103 is connected to the light generation module 102 and is used to generate multi-dimensional information of the target object according to the light information. The multi-dimensional information includes spatial position, detection direction and light wavelength. Specifically, the information generation module 103 can use the light direction and color corresponding to each imaging light point in the light information to process and obtain the coordinate values x, y, z of each feature point on the target object in the three-dimensional space, the angle values θ, φ of each feature point relative to the thermal infrared camera, and the light wavelength λ of each feature point. Then, the information generation module 103 forms the coordinate values x, y, z into the spatial position of each feature point on the target object, forms the angle values θ, φ into the detection direction of each feature point, and forms the light wavelength λ into the light wavelength of each feature point. The θ and φ in the detection direction can be understood as the deflection angles of the target point relative to the detection point in the horizontal and vertical directions. Then, the information generation module 103 can form multi-dimensional information in different color spaces based on the spatial position (x, y, z), detection direction (θ, φ), and light wavelength λ of each feature point on the target object.

[0093] The information query module 104 is connected to the information generation module 103 and is used to calculate the light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively according to the multi-dimensional information. Specifically, the information query module 104 can encode the multi-dimensional information in a high dimension to obtain the encoded information corresponding to the spatial position, detection direction and light wavelength respectively. For details, please refer to Figure 3, perform a high-dimensional mapping 34 of the Fourier characteristics on the spatial position 31, detection direction 32, and light wavelength 33 of each feature point in the multi-dimensional information, and obtain the encoded information corresponding to the spatial position, detection direction, and light wavelength respectively using the high-dimensional mapping results of each feature point (i.e., the easily distinguishable high-dimensional information 35). Subsequently, the information query module 104 inputs the encoded information corresponding to the spatial position into the light field chromatic aberration network and the light field intensity network respectively to output the light intensity information of the light. The light field chromatic aberration network can refer to Figure 4 , the front network layer (Backbone, labeled 42) is used to process the input encoded information of the spatial position (Input, labeled 41), extract geometric features from it, and the rear network layer (Direction&Wavelength, labeled 43) is used to receive and fuse the encoded information of the detection direction and the light wavelength. Finally, through the output layer of the network (labeled 44), the corresponding light intensity value is output for any pixel, representing the light intensity value of the pixel in the RGB channel. The light field intensity network can refer to Figure 5 , the front network layer (Backbone, labeled 52) is used to process the input encoded information of the spatial position (Input, labeled 51), extract geometric features from it, and the rear network layer (Direction, labeled 53) is used to receive and fuse the encoded information of the detection direction. Finally, through the output layer of the network (labeled 54), the corresponding light intensity value is output for any pixel, representing the light intensity value of the pixel in the grayscale channel.

[0094] The light field image rendering module 105 is connected to the information query module 104, and is used to perform sub-pixel light field image rendering using the light intensity information corresponding to the target object in the grayscale channel and the light intensity information in the RGB channel, and obtain a high-resolution three-dimensional image of the target object at multiple angles. Specifically, it can refer to Figure 6 , the backlight (labeled 61) irradiates the liquid crystal panel (labeled 62) after sub-pixel processing. The liquid crystal panel is composed of RGB sub-pixels and grayscale sub-pixels. The sub-pixels emit light with the calculated light intensity. After passing through the microarray structure filled with lenses (labeled 63), the light converges into multiple visual points (such as A and B in the figure) in the image space (labeled 64). Finally, the light enters the visual space (labeled 65), and thus a high-resolution three-dimensional image that can be observed from all angles of the target object in the visual space can be formed.

[0095] In another embodiment, please refer to Figure 7 , the above-mentioned three-dimensional rendering device for thermal imaging, in addition to including the color space separation module 101, the light generation module 102, the information generation module 103, the information query module 104, and the light field image rendering module 105, may further include an error calculation module 106 and a model optimization module 107. The following will be described separately.

[0096] The error calculation module 106 is signal-connected to the information query module 104 and is used to compare and obtain the error information between the calculated light intensity information and the true light intensity value. The error information includes the PSMR value and / or the mean square error value. Specifically, the error calculation module 106 can input the light intensity information corresponding to the rendered light intensity value into the mean square error function to measure the deviation degree between the prediction result of the current network and the true scene light intensity.

[0097] The model optimization module 107 is signal-connected to the error calculation module 106 and the information query module 104 and is used to adjust the network parameters of the neural network according to the error information so as to optimize the network model corresponding to the neural network through negative feedback. Specifically, the model optimization module 107 adjusts the network parameters of the neural network in the information query module 104 so as to optimize the network model corresponding to the neural network through negative feedback. The neural network is a light field color difference network used to generate light intensity information in the RGB channel or a light field intensity network used to generate light intensity information in the grayscale channel. It can be understood that by adjusting the parameters of the network model, the rendering effect can be optimized. For pixels with large errors, the relevant parameters in the preset neural network can be adjusted to improve the rendering accuracy.

[0098] It can be understood that updating network parameters and optimizing network models are common ways for neural networks to train and learn, and specific descriptions will not be given here.

[0099] The light field image rendering module 105 here can be connected to the error calculation module 106. When the error information calculated by the error calculation module 106 is within the preset range, the light intensity of different color spaces is integrated to perform light field image rendering.

[0100] According to the above content, the high-resolution light field fast rendering device specifically includes a color space separation module 101, a light ray generation module 102, an information generation module 103, an information query module 104, a light field image rendering module 105, an error calculation module 106, and a model optimization module 107. The functions of each module can be referred to the corresponding content in the introduction of the high-resolution light field fast rendering method, and no detailed description will be given here.

[0101] The above uses specific examples to elaborate on the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the technical field to which the present invention belongs, according to the idea of the present invention, several simple deductions, deformations or substitutions can also be made.

Claims

1. A high-resolution light field fast rendering method, characterized in that Including: Obtain the scene image of the target scene and perform color space separation processing to obtain the grayscale image and RGB signals of the target scene; Obtain the detection data of the target object in the target scene according to the grayscale image, and calculate the light information by using the detection data; Generate the multi-dimensional information of the target object according to the light information, where the multi-dimensional information includes spatial position, detection direction, and light wavelength; Calculate the light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively based on the multi-dimensional information; Use the light intensity information corresponding to the target object in the grayscale channel and the light intensity information in the RGB channel to perform sub-pixel light field image rendering to obtain the high-resolution three-dimensional image of the target object at multiple angles.

2. The light field fast rendering method according to claim 1, wherein The obtaining the scene image of the target scene and performing color space separation processing to obtain the grayscale image and RGB signals of the target scene includes: Separate the scene image according to the channel dimension to obtain the output data of three independent channels of red, green, and blue, and form them into RGB signals; Generate the grayscale image of the target scene by calculating the weighted average value of the output data of three independent channels of red, green, and blue.

3. The light field fast rendering method according to claim 2, wherein, The obtaining the detection data of the target object in the target scene according to the grayscale image and calculating the light information by using the detection data includes: Identify the target object in the target scene from the grayscale image, obtain the detection data of the target object, and determine each imaging light point corresponding to the detection data in the camera coordinate system; Calculate the light direction and color corresponding to each imaging light point respectively, and form the light information based on the light direction and color corresponding to each imaging light point respectively.

4. The light field fast rendering method according to claim 3, wherein, The generating the multi-dimensional information of the target object according to the light information includes: Use the light direction and color corresponding to each imaging light point in the light information to analyze and obtain the coordinate values of each feature point on the target object in three-dimensional space, represented by x, y, and z, and also obtain the angle values of each feature point relative to the camera, represented by θ and φ, and obtain the light wavelength value of each feature point, represented by λ; Form the spatial position of each feature point on the target object with the coordinate values x, y, and z, form the detection direction of each feature point with the angle values θ and φ, and form the light wavelength of each feature point with the light wavelength value λ; Generate the multi-dimensional information based on the spatial position, detection direction, and light wavelength of each feature point on the target object.

5. The light field fast rendering method according to claim 4, wherein, The calculating the light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively based on the multi-dimensional information includes: Perform high-dimensional encoding on the multi-dimensional information to obtain the encoding information corresponding to the spatial position, detection direction, and light wavelength of each feature point on the target object respectively; Calculate the light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively by using the encoding information; Wherein, the high-dimensional encoding refers to performing a high-dimensional mapping of the Fourier characteristics on the spatial position, detection direction, and light wavelength of each feature point, and obtaining the corresponding encoding information by using the high-dimensional mapping result.

6. The light field fast rendering method according to claim 5, wherein Calculating the light intensity information corresponding to the target object in the grayscale channel and the RGB channel by using the encoding information includes: Inputting the encoding information of the spatial position, detection direction, and light wavelength of each feature point into a preset light field chromatic aberration network to generate the intensity information and color information of the light, and forming the light intensity information in the RGB channel; Inputting the encoding information of the spatial position and detection direction corresponding to each feature point into a preset light field intensity network to generate the intensity information of the light, and forming the light intensity information in the grayscale channel.

7. The light field fast rendering method according to claim 6, wherein After generating the light intensity information of each feature point by using the multi-dimensional information, it further includes: Comparing to obtain the error information between the light intensity information of each feature point and the corresponding true light intensity value, where the error information includes the PSMR value and / or the mean square error value; Adjusting the network parameters of the neural network according to the error information to optimize the network model corresponding to the neural network through negative feedback, where the neural network is the light field chromatic aberration network or the light field intensity network.

8. The light field fast rendering method according to claim 7, characterized in that Performing sub-pixel light field image rendering by using the light intensity information corresponding to the target object in the grayscale channel and the light intensity information in the RGB channel to obtain a high-resolution three-dimensional image of the target object at multiple angles, including: Illuminating the liquid crystal panel after sub-pixel processing by backlight, calculating the light intensity value by using the light intensity information corresponding to the target object in the grayscale channel and the light intensity information in the RGB channel, and driving the sub-pixel points on the liquid crystal panel to emit light according to the calculated light intensity value. The light converges into points after passing through the microstructure array filled with lenses, and a high-resolution three-dimensional image of the target object at multiple angles is formed in the visual space.

9. A rendering device, characterized in that, It includes: A color space separation module for acquiring a scene image of the target scene and performing color space separation processing on the scene image according to the channel dimension to obtain the grayscale image and RGB signal of the target scene; A light generation module for acquiring the detection data of the target object in the target scene according to the grayscale image and obtaining the light information by using the detection data; An information generation module for generating the multi-dimensional information of the target object according to the light information, where the multi-dimensional information includes the spatial position, detection direction, and light wavelength; An information query module for calculating the light intensity information corresponding to the target object in the grayscale channel and the RGB channel respectively by using the multi-dimensional information; A light field image rendering module for performing sub-pixel light field image rendering by using the light intensity information corresponding to the target object in the grayscale channel and the light intensity information in the RGB channel to obtain a high-resolution three-dimensional image of the target object at multiple angles.

10. The rendering device according to claim 9, wherein It further includes an error calculation module and a model optimization module; The error calculation module is signal-connected to the light rendering module and is used to compare and obtain the error information between the calculated light intensity information and the true light intensity value, where the error information includes the PSMR value and / or the mean square error value; The model optimization module is signal-connected to the error calculation module and the information query module, and is used to adjust the network parameters of the neural network according to the error information, so as to optimize the network model corresponding to the neural network through negative feedback; the neural network is a light field chromatic aberration network for generating light intensity information in the RGB channel, or a light field intensity network for generating light intensity information in the grayscale channel.