Efficient rendering method for complex scenes based on visually perceived radiance field

By constructing an initial visual perception radiation field and using a visual sampling rate map and a preset loss function to train and generate a visual perception radiation field, the problems of the neural radiation field method being time-consuming and ignoring significant features are solved, and efficient and high-quality new perspective image rendering is achieved.

CN118608667BActive Publication Date: 2025-09-23BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410772911.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2025-09-23
Estimated Expiration
2044-06-14

AI Technical Summary

Technical Problem

Existing neural radiance field methods and their variants take a long time to render images and tend to ignore significant features in the periphery of the central visual area, resulting in low-quality rendering results and difficulty in generating high-quality new perspective images in a timely manner.

Method used

Construct an initial visual perception radiation field, including an initial density grid, an initial color grid, and an initial visual saliency grid. Generate the visual perception radiation field through training using a visual sampling rate map and a preset loss function. Use visual sensitivity and gaze point information for image rendering to generate high-quality new perspective images.

Benefits of technology

By constructing a visual perception radiation field based on a grid structure, the image rendering time is shortened, the rendering quality is improved, and high-quality new perspective images can be generated in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118608667B_ABST
    Figure CN118608667B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose an efficient rendering method for complex scenes based on a visual perception radiation field. A specific implementation of the method includes: constructing an initial visual perception radiation field; selecting a scene image as a sample image from a scene image set, and performing the following steps: generating a visual sampling rate map based on user gaze point information and the initial visual perception radiation field; determining an image rendering result based on the visual sampling rate map and the initial visual perception radiation field; determining a target difference value between the image rendering result and the sample image rendering data; in response to determining that the target difference value is less than a preset difference threshold, determining the trained initial visual perception radiation field as a visual perception radiation field; inputting rendering perspective information into the visual perception radiation field to output a target rendered image; and controlling a display device to display the target rendered image. This implementation can improve the efficiency and quality of image rendering and generate high-quality new perspective images in a timely manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the fields of computer graphics and virtual reality, and more particularly to a method for efficiently rendering complex scenes based on visually perceived radiation fields. Background Art

[0002] In virtual reality (VR) scenarios, image rendering performance significantly impacts the efficiency of synthesizing new perspective images. Currently, the most common approach to image rendering is to use neural radiance fields (NRFs) or their variants based on multi-layer perceptrons (MLPs) to synthesize new perspective images based on the VR scene's geometric information (e.g., depth, opacity) and the characteristics of the central visual area.

[0003] However, in practice, it is found that when using the above method for image rendering, the following technical problems often occur:

[0004] First, neural radiance field and its variants typically require long network inference times during training and runtime, and tend to overlook significant features outside the central visual area, resulting in a long image rendering process and reduced rendering quality. This makes it difficult to generate high-quality new perspective images in a timely manner.

[0005] Second, the neural radiation field and its variant methods usually use uniform sampling or coarse sampling followed by fine sampling to sample points on light rays. Since the uniform sampling method easily leads to undersampling of visually salient areas, and the coarse sampling followed by fine sampling method easily leads to oversampling in some visually salient areas, it is easy to cause the new perspective image synthesized based on the sampling points to have low quality.

[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention

[0007] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0008] Some embodiments of the present disclosure propose an efficient rendering method for complex scenes based on visually perceived radiation fields to solve one or more of the technical problems mentioned in the above background technology section.

[0009] Some embodiments of the present disclosure provide a method for efficiently rendering complex scenes based on a visual perception radiation field, the method comprising: constructing an initial visual perception radiation field based on a pre-acquired scene image set corresponding to a target three-dimensional scene, wherein each scene image in the above scene image set corresponds to a visual sensitivity image in a visual sensitivity image set, and the above initial visual perception radiation field includes an initial density grid, an initial color grid, and an initial visual saliency grid; selecting a scene image from the above scene image set as a sample image, and performing the following initial visual perception radiation field training steps based on the selected sample image: generating a visual sampling grid based on the user gaze point information corresponding to the selected sample image, the initial density grid, and the initial visual saliency grid included in the initial visual perception radiation field. sampling rate map; based on the above-mentioned visual sampling rate map, the initial density grid and the initial color grid included in the initial visual perception radiation field, determine the image rendering result corresponding to the above-mentioned sample image; based on the preset loss function group, determine the target difference value between the image rendering result corresponding to the above-mentioned sample image and the sample image rendering data, wherein the above-mentioned sample image rendering data includes the sample image and the visual sensitivity image corresponding to the above-mentioned sample image; in response to determining that the above-mentioned target difference value is less than the preset difference threshold, determine the trained initial visual perception radiation field as the visual perception radiation field; input the preset rendering perspective information into the above-mentioned visual perception radiation field to output the target rendered image corresponding to the above-mentioned target three-dimensional scene; control the associated display device to display the above-mentioned target rendered image.

[0010] Embodiments of the present disclosure have the following beneficial effects: Through the efficient rendering methods for complex scenes based on visually perceived radiance fields, as described in some embodiments of the present disclosure, the efficiency and quality of image rendering can be improved, allowing for the timely generation of high-quality new-perspective images. Specifically, the difficulty in generating high-quality new-perspective images in a timely manner arises from the fact that neural radiance fields and their variants typically require lengthy network inference during training and runtime, and tend to overlook salient features outside the central visual area, resulting in a lengthy image rendering process and reduced quality. Based on this, the efficient rendering methods for complex scenes based on visually perceived radiance fields, as described in some embodiments of the present disclosure, first construct an initial visually perceived radiance field based on a pre-acquired set of scene images corresponding to a target three-dimensional scene. Each scene image in the scene image set corresponds to a visual sensitivity image in a visual sensitivity image set. The initial visually perceived radiance field includes an initial density grid, an initial color grid, and an initial visual saliency grid. Thus, an untrained initial visually perceived radiance field corresponding to the target three-dimensional scene can be constructed based on a grid structure. Then, a scene image is selected from the above scene image set as a sample image, and based on the selected sample image, the following initial visual perception radiation field training steps are performed: based on the user gaze point information corresponding to the selected sample image, the initial density grid and the initial visual saliency grid included in the initial visual perception radiation field, a visual sampling rate map is generated; based on the above visual sampling rate map, the initial density grid and the initial color grid included in the initial visual perception radiation field, an image rendering result corresponding to the above sample image is determined; based on a preset loss function group, a target difference value between the image rendering result corresponding to the above sample image and the sample image rendering data is determined, wherein the above sample image rendering data includes the sample image and the visual sensitivity image corresponding to the above sample image; in response to determining that the above target difference value is less than a preset difference threshold, the trained initial visual perception radiation field is determined as the visual perception radiation field. In this way, the initial visual perception radiation field can be trained from sample images of multiple perspectives to obtain a visual perception radiation field with higher image rendering quality. When training the initial visual perception radiation field based on each sample image, a visual sampling rate corresponding to each pixel in the image to be rendered can be generated based on the human eye's sensitivity to scene content and gaze point information, and the initial density grid and initial color grid can be sampled based on the visual sampling rate corresponding to each pixel. Furthermore, a predicted image rendering result is generated based on the sampling results. Subsequently, the preset rendering perspective information is input into the visual perception radiation field to output a target rendered image corresponding to the target three-dimensional scene. Thus, a new perspective image corresponding to the target three-dimensional scene can be generated using the visual perception radiation field. Finally, the associated display device is controlled to display the target rendered image.Therefore, in some embodiments of the present disclosure, the efficient rendering method for complex scenes based on visual perception radiation fields, by pre-constructing a visual perception radiation field based on a grid structure, can sample a relatively limited number of grids for image rendering during training and operation without spending a long time on network inference, thereby shortening the image rendering time. In addition, because the visual sampling rate used for ray sampling is determined based on visual sensitivity and gaze point information, the possibility of ignoring high visual sensitivity areas outside the gaze point can be reduced, thereby improving the image rendering quality. As a result, new perspective images with higher quality can be generated in a timely manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0012] Figure 1 is a flow chart of some embodiments of a method for efficient rendering of complex scenes based on visually perceived radiance fields according to the present disclosure;

[0013] Figure 2 It is a schematic diagram consisting of a visual acuity map, a contrast sensitivity map, and a visual sampling rate map generated by mixing the two at a specific viewing angle according to the efficient rendering method for complex scenes based on visual perception radiation field disclosed in the present invention. DETAILED DESCRIPTION

[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0015] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0017] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0019] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0020] Figure 1 The process 100 of some embodiments of the method for efficiently rendering complex scenes based on visually perceived radiation fields according to the present disclosure is shown. The method for efficiently rendering complex scenes based on visually perceived radiation fields comprises the following steps:

[0021] Step 101: construct an initial visual perception radiation field based on a pre-acquired scene image set corresponding to a target three-dimensional scene.

[0022] In some embodiments, an entity (e.g., a computing device) executing a method for efficiently rendering complex scenes based on a visually perceived radiance field can construct an initial visually perceived radiance field using various methods based on a pre-acquired set of scene images corresponding to a target three-dimensional scene. The complex scene can be a scene containing a large number of objects with complex features. These complex features can include, but are not limited to, complex textures or complex lighting. The target three-dimensional scene can be a complex three-dimensional scene for which a new perspective image is to be synthesized. The new perspective scene image can be an image at a new perspective position different from the original perspective. The scene image set can be a collection of scene images from different perspectives. The scene image can be an RGB (Red, Green, Blue) image corresponding to the target three-dimensional scene. Each scene image in the scene image set can correspond one-to-one to a visual sensitivity image in a pre-generated visual sensitivity image set. The visual sensitivity images in the visual sensitivity image set can be pre-generated grayscale images that characterize the visual salience and contour features of the corresponding scene image. The visual salience can characterize the degree to which an observer visually notices an object in the image. The initial visually perceived radiance field can include an initial density grid, an initial color grid, and an initial visual salience grid. The above-mentioned initial density grid, initial color grid and initial visual saliency grid can be constructed based on a pre-generated voxel grid at different feature dimensions. The above-mentioned voxel grid can be composed of various sub-voxel grids. The above-mentioned sub-voxel grids can be grids obtained by evenly dividing the entire three-dimensional scene according to voxels. Each grid can contain a 3D (Three-Dimensional) space point at the location of the corresponding voxel. For example, when the above-mentioned voxel grid is a grid of 128*128*128 size, the sub-voxel grid is a grid of 1*1*1 size. Each sub-voxel grid is associated with a sub-grid identifier. The sub-grid identifier can be a unique identifier for the corresponding sub-voxel grid. Each sub-voxel grid can store various characteristic values ​​of the corresponding voxel at the corresponding position in the scene. The above-mentioned various characteristic values ​​can include but are not limited to density values, color characteristic values ​​and visual sensitivity characteristic values.

[0023] Furthermore, the density value can represent the probability that an object in a three-dimensional scene exists within the corresponding sub-voxel grid. The density value can be obtained by trilinear interpolation of the eight nearest sub-voxel grids of the corresponding sub-voxel grid. The color eigenvalue can represent the color of the corresponding voxel at the corresponding position in the scene. The color eigenvalue can be represented by a color vector group. Each color vector in the color vector group corresponds one-to-one to a color channel of an RGB image. Each color vector can represent the color information of the corresponding voxel at the corresponding position in the scene stored in the corresponding color channel. Each color vector can be composed of nine spherical harmonic coefficients. Color information can be modeled using second-order spherical harmonic functions, and nine spherical harmonic coefficients can be determined for each color channel. The visual sensitivity eigenvalue can represent the visual sensitivity of the corresponding voxel at the corresponding position in the scene. The visual sensitivity can represent the degree of visual attention of the corresponding voxel by the observer at the corresponding viewing angle. The visual sensitivity eigenvalue can be represented by a visual sensitivity vector. The visual sensitivity vector can be composed of four spherical harmonic coefficients. The visual sensitivity information can be modeled by the first-order spherical harmonic function, and four spherical harmonic coefficients can be obtained by solving. The above-mentioned initial density grid may include a sub-initial density grid set. Each sub-initial density grid corresponds one-to-one to a sub-voxel grid in each of the above-mentioned sub-voxel grids. Each sub-initial density grid may be a grid that stores the density value of the corresponding sub-voxel grid. The above-mentioned initial color grid may include a sub-initial color grid set. Each sub-initial color grid corresponds one-to-one to a sub-voxel grid in each of the above-mentioned sub-voxel grids. Each sub-initial color grid may be a grid that stores the color feature value of the corresponding sub-voxel grid. The above-mentioned initial visual saliency grid may include a sub-initial visual saliency grid set. Each sub-initial visual saliency grid corresponds one-to-one to a sub-voxel grid in each of the above-mentioned sub-voxel grids. Each sub-initial visual saliency grid may be a grid that stores the visual sensitivity feature value of the corresponding sub-voxel grid.

[0024] In some optional implementations of some embodiments, the execution entity may construct an initial visually perceived radiation field based on a pre-acquired scene image set corresponding to the target three-dimensional scene through the following steps:

[0025] The first step is to construct a 3D voxel model based on the scene image set. The 3D voxel model can be a model that represents the target 3D scene using a voxel grid. The 3D voxel model can be constructed based on the scene image set using a pre-defined voxel model construction method. For example, the voxel model construction method can be, but is not limited to, at least one of the following: a voxel model construction method based on feature representation learning, or a voxel model generation method based on graph convolution.

[0026] The second step is to determine the first feature grid, the second feature grid, and the third feature grid corresponding to the above-mentioned three-dimensional voxel model. The first feature grid, the second feature grid, and the third feature grid may have the same structure as the above-mentioned three-dimensional voxel model. The above-mentioned first feature grid may be a grid used only to store the density values ​​corresponding to each sub-voxel grid. The above-mentioned second feature grid may be a grid used only to store the color feature values ​​corresponding to each sub-voxel grid. The above-mentioned third feature grid may be a grid used only to store the visual sensitivity feature values ​​corresponding to each sub-voxel grid. The voxel grid corresponding to the above-mentioned three-dimensional voxel model may be dimensionally split according to the feature values ​​to obtain the first feature grid, the second feature grid, and the third feature grid.

[0027] In the third step, the first feature grid is initialized with eigenvalues ​​to obtain an initial density grid. Each density value in the first feature grid can be initialized to a random density value. The random density value can be a random number between 0 and 1 generated by a random generator. The initialized first feature grid is then determined as the initial density grid.

[0028] In the fourth step, the second feature grid is subjected to eigenvalue initialization processing to obtain an initial color grid. Each color eigenvalue in the second feature grid can be initialized to a random color eigenvalue. The random color eigenvalue can be a randomly generated color vector group. For each color vector in the color vector group to be generated, a random generator can be used to generate nine random numbers between 0 and 1 to form a color vector. The initialized second feature grid is then determined as the initial color grid.

[0029] Step 5: Initialize the eigenvalues ​​of the third feature grid to obtain an initial visual saliency grid. Each visual sensitivity eigenvalue in the third feature grid can be initialized to a random visual sensitivity value. The random visual sensitivity value can be a visual sensitivity vector consisting of four random numbers between 0 and 1. The initialized third feature grid is then determined as the initial visual saliency grid.

[0030] In the sixth step, the model represented by the initial density grid, the initial color grid, and the initial visual saliency grid is determined as the initial visual perception radiation field.

[0031] Optionally, the above-mentioned visual sensitivity image set may be pre-generated by the following steps:

[0032] For each scene image in the scene image set, perform the following steps to generate a visual sensitivity image in the visual sensitivity image set:

[0033] In the first step, edge detection is performed on the scene image using a preset edge detection algorithm to obtain a first grayscale image. The first grayscale image may be a grayscale image of the same size as the scene image. For example, the edge detection algorithm may be, but is not limited to, one of the following: a Sobel operator or a Canny operator.

[0034] In the second step, a preset visual saliency detection algorithm is used to perform visual saliency detection on the scene image to obtain a second grayscale image. The second grayscale image may be a grayscale image of the same size as the scene image. For example, the visual saliency detection algorithm may be, but is not limited to, one of the following: a frequency-tuned (FT) algorithm or a residual spectrum algorithm.

[0035] In the third step, the first grayscale image and the second grayscale image are fused to obtain a visual sensitivity image. The pixel values ​​at the same position in the two grayscale images can be added together, and the result obtained is used as the pixel value at the corresponding position in the visual sensitivity image.

[0036] Step 102: Select a scene image from the scene image set as a sample image, and perform the following initial visual perception radiance field training steps based on the selected sample image:

[0037] Step 1021 : Generate a visual sampling rate map based on the user gaze point information corresponding to the selected sample image, the initial density grid and the initial visual saliency grid included in the initial visual perception radiation field.

[0038] In some embodiments, the execution entity may generate a visual sampling rate map in various ways based on the user gaze point information corresponding to the selected sample image, the initial density grid and the initial visual saliency grid included in the initial visual perception radiation field. The user gaze point information may be pre-acquired information about the position of the user gaze point on the corresponding sample image. The user gaze point may be acquired through an eye tracking device. The eye tracking device may be a device for measuring and recording eye position and movement. For example, the eye tracking device may be an eye tracker. The visual sampling rate map may be a grayscale image with the visual sampling rate as the pixel value. The visual sampling rate may be the sampling rate when sampling along a ray. The sampling rate may be represented by a numerical value between 0 and 1. 1 represents the highest sampling rate and 0 represents the lowest sampling rate. When the visual sampling rate is high, dense sampling is performed along the ray; when the visual sampling rate is low, sparse sampling is performed along the ray.

[0039] In some optional implementations of some embodiments, the execution entity may generate a visual sampling rate map based on the user gaze point information corresponding to the selected sample image, the initial density grid, and the initial visual saliency grid included in the initial visual perception radiance field through the following steps:

[0040] The first step is to construct a visual acuity map based on the user gaze point information corresponding to the sample image, the image resolution information corresponding to the preset camera image, and the corresponding camera field of view angle information. The camera image may be a predicted RGB image to be captured by the camera. The image resolution information may include information about the horizontal and vertical resolutions of the camera image. The image resolution information may be represented by a one-dimensional vector consisting of the horizontal and vertical resolutions of the camera image. It should be noted that the image resolution information also includes information about the resolution of the sample image. The camera field of view angle information may include information about the horizontal and vertical field of view angles when the camera captured the sample image. The camera field of view angle information may be represented by a one-dimensional vector consisting of the horizontal and vertical field of view angles of the camera. Each pixel in the camera image may correspond one-to-one to each ray in a preset ray set. Each ray in the ray set may represent a ray of light passing through the target three-dimensional scene. Each ray is associated with a unique number. The visual acuity map has the same resolution as the camera image. Each pixel in the visual acuity map stores a numerical value between 0 and 1 representing visual acuity. Each visual acuity in the visual acuity map can represent the importance of the corresponding pixel in the camera image. The visual acuity in the visual acuity map can be generated by the following formula group:

[0041]

[0042] Wherein, V represents visual acuity. ω0 represents the lower limit of visual acuity. m represents the visual acuity slope. The visual acuity slope represents the magnitude of change in visual acuity as the eccentricity changes. F represents the coordinates of the pixel point corresponding to the ray. G represents the coordinates of the user's gaze point. e represents the eccentricity of the pixel point corresponding to the ray relative to the user's gaze point. s represents the distance from the user's viewpoint to the imaging screen. The above-mentioned imaging screen is the screen of the display where the camera image is located. a / 2 represents a one-dimensional vector consisting of 1 / 2 of the horizontal field of view angle and 1 / 2 of the vertical field of view angle. a in a / 2 represents the vector corresponding to the above-mentioned camera field of view information. d / 2 represents the center point position of the imaging screen. d in d / 2 represents the vector corresponding to the above-mentioned image resolution information. tan(·) represents the tangent function. atan(·) represents the inverse tangent function. f represents the eccentricity of the pixel point corresponding to the ray relative to the center point of the imaging screen. g represents the eccentricity of the user's gaze point relative to the center point of the imaging screen. x represents the horizontal direction. y represents the vertical direction. fx represents the component of eccentricity f in the x direction. y represents the component of eccentricity f in the y direction. g x It represents the component of eccentricity g in the x direction. y represents the component of the eccentricity g in the y direction.

[0043] In the second step, an initial contrast sensitivity map is constructed based on the above-mentioned ray set, the initial density grid and the initial visual saliency grid included in the initial visual perception radiation field. The above-mentioned initial contrast sensitivity map can be a grayscale image with a resolution lower than that of the above-mentioned camera image. The above-mentioned initial contrast sensitivity map can represent the distribution of initial contrast sensitivities corresponding to each pixel in the scene image. Each pixel value in the above-mentioned initial contrast sensitivity map can represent the initial contrast sensitivity. Each pixel value in the above-mentioned initial contrast sensitivity map can be a value between 0 and 1. The initial contrast sensitivity can represent the sensitivity of human vision to pixels in the scene image. For example, there is a simple scene: in the middle of a white room, there is an apple. Then human vision is more sensitive to this apple. Therefore, the initial contrast sensitivity of the pixel position where the apple is located can be close to 1, and the initial contrast sensitivity of the pixel corresponding to the background white room can be close to 0.

[0044] In some optional implementations of some embodiments, the execution entity may construct an initial contrast sensitivity map based on the ray set, the initial density grid, and the initial visual saliency grid included in the initial visual perception radiation field through the following steps:

[0045] Step 1: For each ray in the above ray set, perform the following steps:

[0046] Sub-step 1: determining the pixel corresponding to the ray in the camera image as the target pixel.

[0047] Sub-step 2: uniformly sample the points on the ray to obtain a first sampling point sequence. The points on the ray may be 3D spatial points in the scene. Each spatial point is associated with a spatial point identifier. The spatial point identifier may be a unique identifier for the spatial point. The first sampling point sequence may be an ordered set of sampling points obtained by uniformly sampling the points on the ray for the first time along the direction of travel of the ray. The first sampling point sequence may be obtained by uniformly sampling the points on the ray according to a preset sampling interval. The preset sampling interval may be a preset interval between two adjacent sampling points.

[0048] Sub-step three: Based on the initial density grid included in the initial visually perceived radiation field, determine the sampling point density value corresponding to each first sampling point in the first sampling point sequence to obtain a sampling point density value sequence. The sampling point density values ​​in the sampling point density value sequence may correspond to the first sampling point with the same sequence number in the first sampling point sequence. The sampling point density values ​​in the sampling point density value sequence may represent the probability of an object existing at the location of the corresponding sampling point in the scene. For each first sampling point in the first sampling point sequence, perform the following steps:

[0049] In the first sub-step, a sub-initial density grid that matches the first sampling point is selected from the sub-initial density grid set included in the initial density grid as a sampling sub-density grid, wherein matching the first sampling point may be that the sub-initial density grid includes the first sampling point.

[0050] In the second sub-step, trilinear interpolation is performed on the density values ​​of the above-mentioned sampling sub-density grid to obtain an updated density value.

[0051] The third sub-step is to determine the updated density value as the sampling point density value.

[0052] Sub-step 4: Based on the initial visual saliency grid included in the initial visual perception radiation field, determine the sampling point visual saliency value corresponding to each first sampling point in the above-mentioned first sampling point sequence to obtain a sampling point visual saliency value sequence. The sampling point visual saliency value in the above-mentioned sampling point visual saliency value sequence may correspond to the first sampling point with the same sequence number in the above-mentioned first sampling point sequence. The sampling point visual saliency value in the above-mentioned sampling point visual saliency value sequence may represent the visual sensitivity of the corresponding sampling point in the scene. For each first sampling point in the above-mentioned first sampling point sequence, perform the following steps:

[0053] In a first sub-step, a sub-initial visual saliency grid that matches the first sampling point is selected from the sub-initial visual saliency grid set included in the initial visual saliency grid as a sampling sub-visual saliency grid. Matching the first sampling point may mean that the sub-initial visual saliency grid includes the first sampling point.

[0054] In the second sub-step, trilinear interpolation is performed on the visual sensitivity feature values ​​of the sampled sub-visual saliency grid to obtain updated visual sensitivity values.

[0055] In the third sub-step, the updated visual sensitivity value is determined as the visual saliency value of the sampling point.

[0056] Sub-step 5: Generate a sampling point feature information sequence based on the sampling point density value sequence and the sampling point visual saliency value sequence. The sampling point feature information in the sampling point feature information sequence may include a sampling point identifier, a sampling point density value, and a sampling point visual sensitivity value. For each first sampling point in the first sampling point sequence, perform the following steps to generate the sampling point feature information in the sampling point feature information sequence:

[0057] The first sub-step is to select the sampling point density value corresponding to the first sampling point from the sampling point density value sequence.

[0058] In the second sub-step, the visual saliency value of the sampling point corresponding to the first sampling point is selected from the visual saliency value sequence of the sampling point as the visual sensitivity value of the sampling point.

[0059] In a third sub-step, the spatial point identifier corresponding to the first sampling point is used as the sampling point identifier, and the sampling point identifier, the selected sampling point density value, and the corresponding sampling point visual sensitivity value are determined as sampling point feature information.

[0060] Sub-step 6: Based on the sampling point feature information sequence, determine the initial contrast sensitivity corresponding to the target pixel. The initial contrast sensitivity corresponding to the target pixel can be generated by the following formula group:

[0061]

[0062] Where r represents the number of the ray. S(·) represents the initial contrast sensitivity. S(r) represents the initial contrast sensitivity of the target pixel corresponding to the ray r. i and j both represent serial numbers, and j is less than i. N1 represents the maximum serial number in the first sampling point sequence corresponding to the ray r. T represents the cumulative transmittance, which reflects the occlusion of the light when it travels to a sampling point. T i It represents the cumulative transmittance of the above ray r when it travels to the i-th first sampling point, that is, the probability that the light ray propagates from the first first sampling point to the i-th first sampling point without being intercepted. σ represents the sampling point density value. i Indicates the sampling point density value of the i-th first sampling point. δ indicates the interval between sampling points. i Represents the interval between the i-th first sampling point and the i+1-th first sampling point. s represents the visual sensitivity value of the sampling point. i Represents the visual sensitivity value of the sampling point of the i-th first sampling point. σ j Indicates the sampling point density value of the jth first sampling point. j represents the interval between the jth first sampling point and the j+1th first sampling point. exp(·) represents the natural exponential function.

[0063] Step 2: Based on the determined initial contrast sensitivities, an initial contrast sensitivity map corresponding to the camera image is generated. First, each initial contrast sensitivity is stored as a pixel value at a corresponding pixel point in the grayscale image, and the grayscale image containing the stored initial contrast sensitivities is determined as the initial contrast sensitivity map.

[0064] In a third step, the initial contrast sensitivity map is upsampled to obtain a contrast sensitivity map. The contrast sensitivity map may include individual contrast sensitivities. The contrast sensitivity map may be a grayscale image having the same resolution as the camera image. The initial contrast sensitivity map may be upsampled using bilinear interpolation to obtain the contrast sensitivity map.

[0065] In practice, after the execution entity performs upsampling processing on the initial contrast sensitivity map, the initial contrast sensitivity map can be restored to the same image resolution as the camera image to be generated, ensuring a one-to-one correspondence between the contrast sensitivity map and each pixel in the camera image.

[0066] The fourth step is to fuse the visual acuity map and the contrast sensitivity map to obtain a visual sampling rate map. Specifically, the following steps can be performed:

[0067] In the first sub-step, for each pixel in the visual acuity map, the following steps may be performed to generate a visual sampling rate in the visual sampling rate map:

[0068] Sub-step 1: Selecting a contrast sensitivity matching the pixel from the contrast sensitivity map as the associated contrast sensitivity. Matching the pixel may mean that the pixel coordinates corresponding to the contrast sensitivity in the contrast sensitivity map are the same as the pixel coordinates.

[0069] Sub-step 2: determining the visual acuity corresponding to the above-mentioned pixel as the associated visual acuity.

[0070] Sub-step three: in response to determining that the associated contrast sensitivity is greater than or equal to the associated visual acuity, determining the associated contrast sensitivity as a visual sampling rate.

[0071] Sub-step four: in response to determining that the associated contrast sensitivity is less than the associated visual acuity, determining the associated visual acuity as the visual sampling rate.

[0072] The second sub-step is to store each obtained visual sampling rate as a pixel value at a corresponding pixel point in the grayscale image, and determine the grayscale image storing the visual sampling rate as a visual sampling rate map.

[0073] Figure 2This is a schematic diagram of a visual acuity map, a contrast sensitivity map, and a visual sampling rate map generated by mixing the two generated at a specific viewing angle according to the method for efficiently rendering complex scenes based on visual perception radiation fields disclosed in the present invention. Figure 2 Includes visual acuity map (upper left), contrast sensitivity map (lower left), and visual sampling rate map generated by mixing the two (right).

[0074] Step 1022 : Determine an image rendering result corresponding to the sample image based on the visual sampling rate map, the initial density grid and the initial color grid included in the initial visually perceived radiance field.

[0075] In some embodiments, the execution entity may determine an image rendering result corresponding to the sample image in various ways based on the visual sampling rate map, the initial density grid, and the initial color grid included in the initial visually perceived radiance field. The image rendering result may represent the visual effect of the predicted camera image.

[0076] In some optional implementations of some embodiments, the execution entity may determine an image rendering result corresponding to the sample image based on the visual sampling rate map and the initial density grid and the initial color grid included in the initial visually perceived radiance field through the following steps:

[0077] In the first step, for each ray in the above ray set, perform the following steps:

[0078] In a first sub-step, a visual sampling rate that matches the target pixel is selected from the visual sampling rate map as a target visual sampling rate. Matching the target pixel may mean that the pixel coordinates corresponding to the visual sampling rate in the visual sampling rate map are the same as the coordinates of the target pixel.

[0079] The second sub-step is to determine the second channel value corresponding to the target pixel based on the importance weight threshold corresponding to the ray and the contrast sensitivity corresponding to the target pixel, and store the second channel value in the contrast sensitivity map. The importance weight threshold may be an upper limit of the importance weight. The importance weight may characterize the importance of the sampling point on the ray in the imaging process. The more significant the pixel feature corresponding to the ray, the greater the importance weight corresponding to the ray. The second channel value may be the result of upsampling the importance weight threshold and the visual sensitivity of the corresponding target pixel. The second channel value may be stored in the second channel of the target pixel in the contrast sensitivity map.

[0080] Optionally, the importance weight threshold corresponding to the above ray may be pre-generated by the following steps:

[0081] Step 1: Based on the sampling point feature information sequence corresponding to the ray, determine the sampling point weight value corresponding to each first sampling point corresponding to the ray to obtain a sampling point weight value sequence. The sampling point weight value represents the importance of the sampling point in synthesizing the corresponding target pixel. Each sampling point weight value can represent the importance of the corresponding sampling point in synthesizing the corresponding target pixel. The sampling point weight value corresponding to each first sampling point can be determined by the following formula:

[0082] w f =T i ×(1-exp(-σ i ×δ i )).

[0083] Where w represents the weight value. i represents the sampling point weight value corresponding to the i-th first sampling point in the above first sampling point sequence.

[0084] Step 2: Select a sampling point weight value that meets a preset weight condition from the sampling point weight value sequence as the target importance weight value. The preset weight condition may be that the sampling point weight value is the maximum value in the sampling point weight value sequence.

[0085] Step 3: Based on the target visual sampling rate and the target importance weight value, generate the importance weight threshold corresponding to the ray. The importance weight threshold corresponding to the ray can be generated by the following formula:

[0086]

[0087] Where τ represents the importance weight threshold. r Indicates the importance weight threshold corresponding to ray r. Represents the target importance weight value corresponding to ray r. P represents the visual sampling rate. r Indicates the visual sampling rate corresponding to ray r.

[0088] The third sub-step is to generate a sampling limit based on the preset sampling upper limit, sampling lower limit, and the above-mentioned target visual sampling rate. The sampling upper limit may be the maximum value of the number of samples allowed. The sampling lower limit may be the minimum value of the number of samples allowed. The sampling limit may be the number of points on the ray that are sampled. The sampling limit may be generated using the following formula:

[0089]

[0090] Where N represents the number of samples. r Indicates the number of sampling limits corresponding to ray r. max indicates the maximum value. min indicates the minimum value. N maxIndicates the upper limit of sampling. N min Indicates the lower limit of sampling. Indicates a round-up operation.

[0091] The fourth sub-step is to perform non-uniform sampling on the points on the ray based on the target visual sampling rate, the sampling limit, and the importance weight threshold to obtain a secondary sampling point sequence. The secondary sampling point sequence may be obtained by sampling the points on the ray a second time. Specifically, the following steps may be performed:

[0092] Sub-step 1: Sample points along the ray along its travel direction according to a preset upper limit for subsampling and a preset sampling strategy. After each sampling session, perform a test on each cumulatively sampled point based on the sampling limit and the importance weight threshold to determine whether to terminate the sampling process prematurely. The upper limit for subsampling may be a preset upper limit on the number of sampling points. The preset sampling strategy may be a preset sampling strategy. The preset sampling strategy may include: if the sampling point weight corresponding to the current sampling point is greater than the importance weight threshold, continue sampling using a smaller ray travel step; otherwise, continue sampling using a larger ray travel step. The sampling process may be terminated prematurely when the number of target sampling points exceeds the sampling limit. The target sampling point may be a sampling point that matches the importance weight threshold among the cumulatively sampled points during subsampling. Matching the importance weight threshold may mean that the sampling point weight corresponding to the sampling point is greater than the importance weight threshold. It should be noted that when the number of samples reaches the upper limit for subsampling, sampling is terminated even if the number of target sampling points is still below the sampling limit.

[0093] Optionally, the ray marching length corresponding to the above preset sampling strategy can be defined by the following formula:

[0094]

[0095] Among them, STEP represents the progress of the ray. i→i+1 Indicates the ray travel step length from the current sampling point to the next sampling point. base represents the preset initial step size. β represents a hyperparameter.

[0096] Sub-step three: in response to determining that the sampling process is terminated, determining each sampling point obtained by sampling along the ray travel direction as a secondary sampling point sequence.

[0097] As an example, when the upper limit of subsampling is 50 and the sampling limit is 30, if the number of samplings has reached 40 (less than 50) and the number of target sampling points has reached 31 (greater than 30), the sampling process is terminated early, and the total number of subsampling points obtained by subsampling is 40. If the number of samplings has reached 50, the sampling process is terminated, and the total number of subsampling points obtained by subsampling is 50.

[0098] The above-described secondary sampling point sequence generation steps and related content, as an inventive feature of an embodiment of the present disclosure, address the second technical issue mentioned in the background art: low quality of new-perspective images synthesized based on sampling points. The low quality of new-perspective images synthesized based on sampling points is often caused by the following: Neural Radiance Field and its variants typically sample points on light rays using uniform sampling or coarse-first-then-fine sampling. Uniform sampling easily leads to undersampling of visually salient areas, while coarse-first-then-fine sampling easily leads to oversampling in some visually salient areas. Addressing these issues can improve the quality of synthesized new-perspective images. To achieve this, for each pixel in the camera image to be generated, the target visual sampling rate and importance weight threshold corresponding to the pixel are first determined. Then, sampling is performed based on the target visual sampling rate. During the sampling process, if the sampling point weight of a sampling point is greater than the importance weight threshold, a smaller ray marching step length can be used to perform the next sampling step near the object surface. Otherwise, a larger ray marching step length is used to skip blank voxels to quickly reach the object surface or boundary. Finally, the sampling points obtained along the ray marching direction are determined as a secondary sampling point sequence. Therefore, when using the visual sampling rate to guide the sampling process, we can allocate fewer samples to areas of low visual sensitivity and more samples to areas of high visual sensitivity, based on the visual sensitivity of each pixel in different regions of the scene. Furthermore, we can filter out invalid samples that contribute little to the pixel color based on an importance weight threshold, while allocating more samples to visible scene objects. This improves the quality of the new perspective image when synthesized based on the sampling points.

[0099] The fifth sub-step involves determining, based on the initial density grid included in the initial visually perceived radiation field, a sub-sampling point density value corresponding to each sub-sampling point in the sub-sampling point sequence, thereby obtaining a sub-sampling point density value sequence. The sub-sampling point density values ​​in the sub-sampling point density value sequence may correspond to sub-sampling points with the same sequence number in the sub-sampling point sequence. The sub-sampling point density values ​​in the sub-sampling point density value sequence may represent the probability of an object existing at the location of the corresponding sub-sampling point in the scene. For each sub-sampling point in the sub-sampling point sequence, the following steps are performed:

[0100] Sub-step 1: Selecting a sub-initial density grid that matches the secondary sampling point from the sub-initial density grid set included in the initial density grid as the secondary sampling sub-density grid, wherein matching the first sampling point may be that the sub-initial density grid includes the secondary sampling point.

[0101] Sub-step 2: performing trilinear interpolation processing on the density value of the above-mentioned sampling sub-density grid to obtain an optimized density value.

[0102] Sub-step three: determining the above optimized density value as the secondary sampling point density value.

[0103] In a sixth sub-step, sub-sampling points that match the importance weight threshold are sequentially selected from the sub-sampling point sequence as target sub-sampling points, thereby obtaining a target sub-sampling point sequence. Matching the importance weight threshold may mean that the sampling point weight corresponding to the sub-sampling point in the sub-sampling point sequence is greater than the importance weight threshold.

[0104] The seventh sub-step involves determining, based on the initial color grid included in the initial visually perceived radiance field, the sub-sampling point color value corresponding to each target sub-sampling point in the target sub-sampling point sequence, thereby obtaining a sub-sampling point color value sequence. The sub-sampling point color values ​​in the sub-sampling point color value sequence may correspond to sub-sampling points with the same sequence number in the sub-sampling point sequence. The sub-sampling point color values ​​in the sub-sampling point color value sequence may represent the color of the light in the scene at the corresponding sampling point. For each sub-sampling point in the sub-sampling point sequence, the following steps are performed:

[0105] Sub-step 1: Selecting a sub-initial color grid that matches the secondary sampling point from the sub-initial color grid set included in the initial color grid as a sampling sub-color grid, wherein matching the secondary sampling point may be that the sub-initial color grid includes the secondary sampling point.

[0106] Sub-step 2: performing trilinear interpolation processing on the color feature values ​​of the above-mentioned sampled sub-color grid to obtain optimized color values.

[0107] Sub-step three: determining the optimized color value as the secondary sampling point color value.

[0108] In practice, only when the sampling point weight corresponding to the secondary sampling point is greater than the aforementioned importance weight threshold is the sub-initial color grid sampled and its color value calculated. This can reduce unnecessary color synthesis and improve image rendering performance.

[0109] The eighth sub-step is to generate a pixel color value corresponding to the target pixel based on the sub-sampling point density value sequence and the sub-sampling point color value sequence. The pixel color value may be the color value of the target pixel. Each sub-sampling point density value in the sub-sampling point density value sequence may be used as the sampling point density value of the corresponding sub-sampling point, and each sub-sampling point color value in the sub-sampling point color value sequence may be used as the sampling point color value of the corresponding sub-sampling point. The pixel color value corresponding to the target pixel may be generated using the following formula:

[0110]

[0111] Where C(·) represents the pixel color value. h represents the sequence number. H represents the maximum sequence number of the secondary sampling point sequence corresponding to the above ray r. C(r) represents the pixel color value of the target pixel corresponding to the above ray r. C represents the color value of the sampling point. c h Indicates the color value of the sampling point of the hth secondary sampling point. h It represents the cumulative transmittance when the above ray r travels to the i-th secondary sampling point. h Indicates the sampling point density value of the hth secondary sampling point. h Represents the interval between the hth subsampling point and the h+1th subsampling point.

[0112] In a ninth sub-step, the contrast sensitivity corresponding to the target pixel is determined as the pixel visual sensitivity.

[0113] In a tenth sub-step, the pixel color value and the pixel visual sensitivity are determined as a pixel rendering result corresponding to the target pixel.

[0114] In the second step, each determined pixel rendering result is determined as an image rendering result corresponding to the sample image.

[0115] Step 1023 : Determine a target difference value between the image rendering result corresponding to the sample image and the sample image rendering data based on a preset loss function group.

[0116] In some embodiments, the execution entity may determine a target difference value between the image rendering result corresponding to the sample image and the sample image rendering data based on a preset loss function set. The sample image rendering data may include the sample image and a visual sensitivity image corresponding to the sample image. The loss functions in the loss function set may be functions used to evaluate the difference between the predicted rendered image and the actual rendered image from different dimensions. The target difference value may represent the difference between the predicted rendered image and the actual rendered image.

[0117] Optionally, the above-mentioned loss function group may include a photometric loss function, a visual perception loss function and an importance weight constraint loss function. Among them, the above-mentioned photometric loss function can be used to determine the color difference between the image rendering result corresponding to the above-mentioned sample image and the above-mentioned sample image rendering data. For example, the above-mentioned photometric loss function can be a mean square error loss function. The above-mentioned visual perception loss function can be used to determine the difference in visual sensitivity between the image rendering result corresponding to the above-mentioned sample image and the above-mentioned sample image rendering data. For example, the above-mentioned visual perception loss function can be a mean square error loss function. The above-mentioned importance weight constraint loss function can be used to constrain the density value of the secondary sampling point. Among them, the above-mentioned importance weight constraint loss function can be expressed by the following formula:

[0118]

[0119] Among them, L weight represents the loss value of the importance weight constraint loss function. R represents the ray set. ||·|| represents the first norm. N2 represents the number of secondary sampling points corresponding to the ray. f1(·) is the sigmoid function, which simulates the process of filtering sampling points in a differentiable form so that the gradient can be back-propagated. k represents the hyperparameter used to control the function in w i =τ. When w i When <τ, f1(w i ,τ) takes the value of 0, when w i >τ, f1(w i , τ) takes the value of 1. The above importance weight constraint loss function can make the weight distribution of all sampling points meet the following conditions: for each ray, the ratio of the number of all sampling points on the ray to the number of sampling points satisfying w>τ is 1:P. Among them, P in 1:P is the visual sampling rate of the target pixel corresponding to the ray. It should be noted that in the above importance weight constraint loss function, for each ray r in R, there are corresponding N2, τ, w i .

[0120] In addition, the execution subject may determine a target difference value between the image rendering result corresponding to the sample image and the sample image rendering data by the following steps:

[0121] In the first step, a color loss value is generated based on the color values ​​of each pixel in the image rendering result and the sample image in the sample image rendering data using the photometric loss function. The color loss value can represent the color difference between the predicted rendered image and the actual rendered image.

[0122] The second step is to generate a visual sensitivity loss value based on the visual sensitivity of each pixel in the image rendering result and the visual sensitivity image in the sample image rendering data using the photometric loss function. The visual sensitivity loss value can represent the difference in visual sensitivity between the predicted rendered image and the actual rendered image.

[0123] The third step is to determine the density loss value corresponding to the image rendering result by constraining the loss function with the importance weight, wherein the density loss value can represent the difference between the predicted density value and the constrained density value.

[0124] In the fourth step, the sum of the color loss value, the visual sensitivity loss value, and the weight loss value is determined as the target difference value.

[0125] In practice, the execution subject adopts a training method based on importance weight constraints, which can better constrain the solution space of scene geometry and ensure the effectiveness of the sampling strategy in the above step 1022.

[0126] Step 1024 : In response to determining that the target difference value is less than the preset difference threshold, the trained initial visual perception radiation field is determined as the visual perception radiation field.

[0127] In some embodiments, the execution entity may, in response to determining that the target difference value is less than a preset difference threshold, determine the trained initial visual perception radiation field as the visual perception radiation field. The preset difference threshold may be a pre-set upper limit of the difference value. The visual perception radiation field may include a density grid, a color grid, and a visual saliency grid. The density grid may be the trained initial density grid. The color grid may be the trained initial color grid. The visual saliency grid may be the trained initial visual saliency grid.

[0128] Optionally, the execution subject may further adjust the parameters in the initial visual perception radiation field in response to determining that the target difference value is not less than the preset difference threshold, select an unused scene image from the scene image set as a sample image, use the adjusted initial visual perception radiation field as the initial visual perception radiation field, and perform the initial visual perception radiation field training step again. The parameters in the initial visual perception radiation field may be the individual feature values ​​stored in the initial density grid, the initial color grid, and the initial visual saliency grid. The parameters in the initial visual perception radiation field may be continuously updated through a back-propagation learning method to gradually reduce the target difference value.

[0129] Step 103: input the preset rendering perspective information into the visual perception radiation field to output a target rendered image corresponding to the target three-dimensional scene.

[0130] In some embodiments, the execution entity may input preset rendering perspective information into the visually perceived radiation field to output a target rendered image corresponding to the target three-dimensional scene. The rendering perspective information may be information about the camera pose matrix corresponding to the image to be synthesized. The camera pose matrix may be a matrix representing the position and orientation of the camera. A corresponding set of rays exists for the rendering perspective information. Based on the set of rays corresponding to the rendering perspective information, the visually perceived radiation field may first generate a visual sampling rate map by uniformly sampling each ray based on a density grid and a visual saliency grid. Then, based on the visual sampling rate map, each ray is non-uniformly sampled to determine the eigenvalues ​​of each sampling point, including density and color values. Subsequently, based on the eigenvalues ​​of each sampling point, the color value corresponding to each pixel in the image to be synthesized is determined using the pixel color value generation formula to perform pixel rendering. Finally, the rendered image to be synthesized is output as the target rendered image.

[0131] Step 104: Control the associated display device to display the target rendered image.

[0132] In some embodiments, the execution entity may control an associated display device to display the target rendered image, wherein the display device may be a device with a display screen.

[0133] Embodiments of the present disclosure have the following beneficial effects: Through the efficient rendering methods for complex scenes based on visually perceived radiance fields, as described in some embodiments of the present disclosure, the efficiency and quality of image rendering can be improved, allowing for the timely generation of high-quality new-perspective images. Specifically, the difficulty in generating high-quality new-perspective images in a timely manner arises from the fact that neural radiance fields and their variants typically require lengthy network inference during training and runtime, and tend to overlook salient features outside the central visual area, resulting in a lengthy image rendering process and reduced quality. Based on this, the efficient rendering methods for complex scenes based on visually perceived radiance fields, as described in some embodiments of the present disclosure, first construct an initial visually perceived radiance field based on a pre-acquired set of scene images corresponding to a target three-dimensional scene. Each scene image in the scene image set corresponds to a visual sensitivity image in a visual sensitivity image set. The initial visually perceived radiance field includes an initial density grid, an initial color grid, and an initial visual saliency grid. Thus, an untrained initial visually perceived radiance field corresponding to the target three-dimensional scene can be constructed based on a grid structure. Then, a scene image is selected from the above scene image set as a sample image, and based on the selected sample image, the following initial visual perception radiation field training steps are performed: based on the user gaze point information corresponding to the selected sample image, the initial density grid and the initial visual saliency grid included in the initial visual perception radiation field, a visual sampling rate map is generated; based on the above visual sampling rate map, the initial density grid and the initial color grid included in the initial visual perception radiation field, an image rendering result corresponding to the above sample image is determined; based on a preset loss function group, a target difference value between the image rendering result corresponding to the above sample image and the sample image rendering data is determined, wherein the above sample image rendering data includes the sample image and the visual sensitivity image corresponding to the above sample image; in response to determining that the above target difference value is less than a preset difference threshold, the trained initial visual perception radiation field is determined as the visual perception radiation field. In this way, the initial visual perception radiation field can be trained from sample images of multiple perspectives to obtain a visual perception radiation field with higher image rendering quality. When training the initial visual perception radiation field based on each sample image, a visual sampling rate corresponding to each pixel in the image to be rendered can be generated based on the human eye's sensitivity to scene content and gaze point information, and the initial density grid and initial color grid can be sampled based on the visual sampling rate corresponding to each pixel. Furthermore, a predicted image rendering result is generated based on the sampling results. Subsequently, the preset rendering perspective information is input into the visual perception radiation field to output a target rendered image corresponding to the target three-dimensional scene. Thus, a new perspective image corresponding to the target three-dimensional scene can be generated using the visual perception radiation field. Finally, the associated display device is controlled to display the target rendered image.Therefore, in some embodiments of the present disclosure, the efficient rendering method for complex scenes based on visual perception radiation fields, by pre-constructing a visual perception radiation field based on a grid structure, can sample a relatively limited number of grids for image rendering during training and operation without spending a long time on network inference, thereby shortening the image rendering time. In addition, because the visual sampling rate used for ray sampling is determined based on visual sensitivity and gaze point information, the possibility of ignoring high visual sensitivity areas outside the gaze point can be reduced, thereby improving the image rendering quality. As a result, new perspective images with higher quality can be generated in a timely manner.

[0134] The technical contents not elaborated in detail in the present invention belong to the common knowledge of those skilled in the art.

[0135] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. An efficient rendering method for complex scenes based on visually perceived radiation fields, comprising: constructing an initial visual perception radiance field based on a pre-acquired scene image set corresponding to a target three-dimensional scene, wherein each scene image in the scene image set corresponds to a visual sensitivity image in a pre-generated visual sensitivity image set, and the initial visual perception radiance field includes an initial density grid, an initial color grid, and an initial visual saliency grid; A scene image is selected from the scene image set as a sample image, and the following initial visual perception radiance field training steps are performed based on the selected sample image: Generate a visual sampling rate map based on the user gaze point information corresponding to the selected sample image, the initial density grid and the initial visual saliency grid included in the initial visual perception radiation field; Determining an image rendering result corresponding to the sample image based on the visual sampling rate map, an initial density grid and an initial color grid included in the initial visually perceived radiance field; Determining a target difference value between an image rendering result corresponding to the sample image and sample image rendering data based on a preset loss function group, wherein the sample image rendering data includes a sample image and a visual sensitivity image corresponding to the sample image; In response to determining that the target difference value is less than a preset difference threshold, determining the trained initial visual perception radiation field as the visual perception radiation field; Inputting preset rendering perspective information into the visual perception radiation field to output a target rendered image corresponding to the target three-dimensional scene; controlling an associated display device to display the target rendered image; The step of constructing an initial visual perception radiation field based on a pre-acquired scene image set corresponding to the target three-dimensional scene includes: constructing a three-dimensional voxel model based on the scene image set; Determining a first feature grid, a second feature grid, and a third feature grid corresponding to the three-dimensional voxel model, wherein the first feature grid is a grid used only to store density values ​​corresponding to each sub-voxel grid, where the density value represents the probability of an object in the three-dimensional scene existing within the corresponding sub-voxel grid; the second feature grid is a grid used only to store color feature values ​​corresponding to each sub-voxel grid; and the third feature grid is a grid used only to store visual sensitivity feature values ​​corresponding to each sub-voxel grid, where the visual sensitivity represents the degree to which the corresponding voxel receives visual attention of the observer at a corresponding viewing angle; Performing eigenvalue initialization processing on the first feature grid to obtain an initial density grid; performing eigenvalue initialization processing on the second feature grid to obtain an initial color grid; performing eigenvalue initialization processing on the third feature grid to obtain an initial visual saliency grid; A model represented by the initial density grid, the initial color grid, and the initial visual saliency grid is determined as an initial visually perceived radiance field.

2. The method according to claim 1, wherein The method further comprises: In response to determining that the target difference value is not less than the preset difference threshold, the parameters in the initial visual perception radiation field are adjusted, and an unused scene image is selected from the scene image set as a sample image, and the adjusted initial visual perception radiation field is used as the initial visual perception radiation field, and the initial visual perception radiation field training step is performed again.

3. The method according to claim 2, wherein: The generating of a visual sampling rate map based on the user gaze point information corresponding to the selected sample image, the initial density grid and the initial visual saliency grid included in the initial visual perception radiation field includes: Constructing a visual acuity map based on user gaze point information corresponding to the sample image, image resolution information corresponding to a preset camera image, and corresponding camera field of view angle information, wherein each pixel in the camera image corresponds one-to-one to each ray in a preset ray set; constructing an initial contrast sensitivity map based on the ray set, an initial density grid and an initial visual saliency grid included in the initial visual perception radiation field; performing upsampling processing on the initial contrast sensitivity map to obtain a contrast sensitivity map, wherein the contrast sensitivity map includes various contrast sensitivities; The visual acuity map and the contrast sensitivity map are fused to obtain a visual sampling rate map.

4. The method according to claim 3, wherein: The constructing of an initial contrast sensitivity map based on the ray set, the initial density grid and the initial visual saliency grid included in the initial visual perception radiation field comprises: For each ray in the ray set, perform the following steps: determining a pixel in the camera image corresponding to the ray as a target pixel; uniformly sampling the points on the ray to obtain a first sampling point sequence; Determining, based on an initial density grid included in the initial visually perceived radiation field, a sampling point density value corresponding to each first sampling point in the first sampling point sequence to obtain a sampling point density value sequence; Determining, based on an initial visual saliency grid included in the initial visual perception radiation field, a sampling point visual saliency value corresponding to each first sampling point in the first sampling point sequence to obtain a sampling point visual saliency value sequence; generating a sampling point feature information sequence based on the sampling point density value sequence and the sampling point visual saliency value sequence; Determining an initial contrast sensitivity corresponding to the target pixel based on the sampling point feature information sequence; Based on the determined initial contrast sensitivities, an initial contrast sensitivity map corresponding to the camera image is generated.

5. The method according to claim 4, wherein The determining of an image rendering result corresponding to the sample image based on the visual sampling rate map and the initial density grid and the initial color grid included in the initial visually perceived radiance field comprises: For each ray in the ray set, perform the following steps: Selecting a visual sampling rate that matches the target pixel from the visual sampling rate map as a target visual sampling rate; determining a second channel value corresponding to the target pixel based on an importance weight threshold corresponding to the ray and a contrast sensitivity corresponding to the target pixel, and storing the second channel value in the contrast sensitivity map; Generate a sampling limit number based on a preset sampling upper limit number, a preset sampling lower limit number and the target visual sampling rate; Based on the target visual sampling rate, the sampling limit number and the importance weight threshold, performing non-uniform sampling processing on the points on the ray to obtain a secondary sampling point sequence; Determining, based on an initial density grid included in the initial visually perceived radiation field, a secondary sampling point density value corresponding to each secondary sampling point in the secondary sampling point sequence to obtain a secondary sampling point density value sequence; Selecting, in sequence, secondary sampling points that match the importance weight threshold from the secondary sampling point sequence as target secondary sampling points to obtain a target secondary sampling point sequence; Determining, based on an initial color grid included in the initial visually perceived radiation field, a subsampling point color value corresponding to each target subsampling point in the target subsampling point sequence to obtain a subsampling point color value sequence; generating a pixel color value corresponding to the target pixel based on the sub-sampling point density value sequence and the sub-sampling point color value sequence; Determining the contrast sensitivity corresponding to the target pixel as the pixel visual sensitivity; Determining the pixel color value and the pixel visual sensitivity as a pixel rendering result corresponding to the target pixel; The determined rendering results of each pixel are determined as image rendering results corresponding to the sample image.

6. The method according to claim 5, wherein: The importance weight threshold corresponding to the ray is generated by the following steps: Determining, based on a sampling point feature information sequence corresponding to the ray, a sampling point weight value corresponding to each first sampling point corresponding to the ray, to obtain a sampling point weight value sequence, wherein the sampling point weight value represents the importance of the sampling point in synthesizing the corresponding target pixel; Selecting a sampling point weight value that meets a preset weight condition from the sampling point weight value sequence as a target importance weight value; An importance weight threshold corresponding to the ray is generated based on the target visual sampling rate and the target importance weight value.

7. The method according to claim 6, wherein: The loss function group includes a photometric loss function, a visual perception loss function and an importance weight constraint loss function. The photometric loss function is used to determine the color difference between the image rendering result corresponding to the sample image and the sample image rendering data. The visual perception loss function is used to determine the difference in visual sensitivity between the image rendering result corresponding to the sample image and the sample image rendering data. The importance weight constraint loss function is used to constrain the density value of the secondary sampling point.

Citation Information

Patent Citations

  • Image rendering model training method and device and image rendering method and device

    CN114493995A

  • Scene-adaptive fixation point neural radiation field rendering method and system

    CN117058293A