Semantic segmentation method based on HDR dynamic nerve radiation field
By introducing a semantic segmentation method based on HDR dynamic neural radiation field in image segmentation technology, using tone mapping module and mask deformation field network, the problem of exposure unevenness caused by changes in lighting conditions in real-world data and dynamic object segmentation problems are solved, and efficient four-dimensional segmentation of objects in HDR dynamic scenes is achieved.
Patent Information
- Application Number
- CN202510535173.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The prior art is difficult to deal with the problem of uneven exposure caused by changes in lighting conditions in real-world data in image segmentation tasks, and it is difficult to effectively segment dynamic objects, resulting in inconsistency across views and poor segmentation effects.
The semantic segmentation method based on HDR dynamic neural radiation field is adopted, and the tone mapping module and mask deformation field network is constructed to realize the four-dimensional segmentation of any object in the HDR dynamic scene, and the segmentation ability of the two-dimensional visual model is used to expand to four-dimensional space.
It realizes efficient semantic segmentation of objects in HDR dynamic scenes, improves segmentation effect and consistency, and can accurately segment target objects in dynamic scenes without expensive four-dimensional scene annotations.
Smart Images

Figure CN120070898A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular, to a semantic segmentation method based on HDR dynamic neural radiance fields. Background Art
[0002] Neural Radiance Field (NeRF) is a deep learning model for three-dimensional implicit space modeling. Based on the scene images captured under known views and the internal and external camera parameters, it predicts the color value and volume density of each sampling point through a neural network, and synthesizes images under new views based on the volume rendering method, showing great success in representing 3D scenes and synthesizing new view images. NeRF represents the three-dimensional space as a set of learnable and continuous radiance fields. Given the position of the input sampling point and the ray direction, it outputs the color and volume density of the sampling point. HDR (High Dynamic Range) technology can capture images with different brightness levels and present images with a larger exposure range than traditional images. Using images with different exposures as input, the HDR neural radiance field can reconstruct an HDR scene with uniform exposure.
[0003] For the image segmentation task, SA3D attempts to combine an arbitrary segmentation model (SAM) with a neural radiance field to segment 3D objects. It provides manual segmentation hint points for the target object in a single view, and then SA3D alternately performs mask inverse rendering and cross-view self-hinting between various views to iteratively complete the 3D mask of the target object constructed with a voxel grid. Under the volume density distribution of the neural radiance field, SA3D projects the 2D mask obtained by SAM in the current view into 3D space and automatically extracts reliable hints from the 2D view rendered by the neural radiance field in another view as the input of SAM.
[0004] However, the input of SAM needs to maintain uniform illumination and contain sufficient scene information. In real-world datasets, affected by lighting conditions, the input images always contain overexposed or underexposed areas, which may significantly lead to cross-view inconsistencies and have an adverse impact on the learning processes of the neural radiance field and SAM. In addition, previous methods mainly targeted static scenes, which are limited in segmenting 4D dynamic objects. A serious problem in dynamic interactive segmentation is that the hint points belonging to one object at timestamp t0 may belong to other objects at timestamp t1, making it difficult to propagate the hint points to other timestamps. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a semantic segmentation method based on HDR dynamic neural radiance fields in view of the deficiencies of the above-mentioned prior art. By utilizing the segmentation ability of two-dimensional vision large models on pictures and videos, the segmentation ability of two-dimensional models is extended to four dimensions. Based on the constructed tone mapping module and mask deformation field network, mask inverse rendering is realized, and thus semantic segmentation of any object in the HDR dynamic neural radiance field is achieved.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows: The present invention provides a semantic segmentation method based on HDR dynamic neural radiance fields, including the following steps: S1: Construct a scene dataset, including scene picture files and scene label files, and divide the scene dataset into a training set and a test set according to a ratio; The scene label file includes the file path, rotation matrix, translation vector, and timestamp information corresponding to each scene picture file. The training set is used for scene reconstruction based on the HDR dynamic neural radiance field, and the test set is used to evaluate the quality of scene reconstruction and the quality of segmentation; S2: Use the HDR dynamic neural radiance field model to reconstruct the pictures with different exposure rates in the training set into a training set HDR dynamic scene with uniform exposure, and output a training set HDR picture sequence F1 based on the training set HDR dynamic scene, including n training set HDR pictures; Input the pictures with different exposure rates in the training set into the HDR dynamic neural radiance field model to obtain a training set HDR dynamic scene with uniform exposure; in the training set HDR dynamic scene, any ray is: ; wherein, is the camera origin, is the ray direction, is the distance from a point on the ray to the origin along the ray direction; For the sampling points and the ray direction on the ray in the training set HDR dynamic scene, based on the mapping function calculate the spatial color value and the volume density of the ray , and use the volume rendering method to obtain the integral of the ray on the imaging plane and the integral of the opacity of the sampling points on the ray in the ray direction, which is used to represent the ray Upsampling points The weights are as shown in the following formula: ; ; Among them, and respectively represent the near point and the far point of the sampling ray ; Based on the training set HDR dynamic scene, output the training set HDR picture sequence F1, including n training set HDR pictures; S3: Construct a tone mapping module, input the training set HDR picture sequence F1 into the tone mapping module, and obtain the training set LDR low dynamic range picture sequence F2 with uniform exposure; The tone mapping module includes a normalization layer, an adjustment layer, and an output layer; the normalization layer obtains the training set HDR picture sequence F1, and normalizes the n training set HDR pictures included in the training set HDR picture sequence F1, and uses the sigmoid function to map the pixel value of each pixel point in the training set HDR picture to between 0 and 1 to obtain the normalized training set HDR picture; the adjustment layer adjusts the brightness and contrast of the normalized training set HDR picture, where the adjustment of brightness and contrast is controlled by hyperparameters to reduce the model training parameters; the output layer maps the pixel value of each pixel point in the adjusted training set HDR picture to the LDR color value range of 0-255 to obtain the training set LDR low dynamic range picture sequence F2 with uniform exposure; S31: Obtain the training set HDR picture sequence F1, and normalize the n training set HDR pictures included therein, and use the sigmoid function to map the pixel value of each pixel point in the training set HDR picture to between 0 and 1 to obtain the normalized training set HDR picture, as shown in the following formula: ; Among them, is the pixel value of the pixel point in the training set HDR picture; S32: Adjust the brightness and contrast of the normalized training set HDR picture to obtain the adjusted training set HDR picture; Use the adjustment layer of the tone mapping module controlled by hyperparameters for brightness and contrast to adjust the brightness of the training set HDR picture normalized in S31, as shown in the following formula: ; Among them, is the brightness hyperparameter, is thei Zhang training set HDR images, For the i th training set HDR image after brightness adjustment; Perform contrast adjustment on the training set HDR images after brightness adjustment, as shown in the following formula: ; ; Where, is the contrast hyperparameter, For the i th training set HDR image after contrast adjustment, is the contrast enhancement function, as shown in the following formula: ; ; ; Where, is the image brightness function defined based on the human eye's sensitivity to the three primary colors, is the brightness enhancement function, is the training set HDR image after brightness adjustment The pixel value of the pixel point on the red channel, is the training set HDR image after brightness adjustment The pixel value of the pixel point on the green channel, is the training set HDR image after brightness adjustment The pixel value of the pixel point on the blue channel; S33: Map the pixel value of each pixel point in the adjusted training set HDR image to the LDR color value range of 0 - 255 to obtain a sequence F2 of training set LDR low - dynamic - range images with uniform exposure; S4: Use a two - dimensional vision segmentation auxiliary model to segment the sequence F2 of training set LDR low - dynamic - range images with uniform exposure to obtain a two - dimensional segmentation mask of the target object in the training set HDR dynamic scene; S41: Select a two - dimensional vision large model as the two - dimensional vision segmentation auxiliary model and input the sequence F2 of training set LDR low - dynamic - range images into the two - dimensional vision large model; S42: Set a number of segmentation hint points and background points on the target object in the training set HDR dynamic scene, and use the two - dimensional vision segmentation auxiliary model to obtain a two - dimensional segmentation mask of the target object in the training set HDR dynamic scene; For the RGB three - channel image with size H×W×3 in the sequence F2 of training set LDR low - dynamic - range images, where H is the height of the training set LDR low - dynamic - range image and W is the width of the training set LDR low - dynamic - range image, obtain its two - dimensional segmentation mask It is a two-dimensional segmentation mask with dimensions H×W×1. Stores a single-channel boolean value for each pixel in H×W×1. The area with a mask is True, and the boolean value of the pixels in this area is 1. The area without a mask is False, and the boolean value of the pixels in this area is 0. S43: If the two-dimensional visual segmentation auxiliary model cannot achieve the desired segmentation effect, then adjust the brightness hyperparameter in S32 and the contrast hyperparameter to optimize the segmentation effect, and repeat S41 to S42 to obtain the two-dimensional segmentation mask sequence with the best segmentation effect for the target object in the HDR dynamic scene. S5: Construct a mask deformation field network, input the two-dimensional segmentation mask sequence with the best segmentation effect for the target object in the HDR dynamic scene obtained by the two-dimensional segmentation model auxiliary model into the mask deformation field network for mask inverse rendering, project the two-dimensional segmentation mask into the four-dimensional space, and obtain the four-dimensional segmentation mask of the target object in the training set HDR dynamic scene, thereby realizing the segmentation of the target object in the training set HDR dynamic scene. The mask deformation field network includes a spatio-temporal interpolation encoder and a mask deformation decoder. The spatio-temporal interpolation encoder network includes a HexPlane sub-module and an MLP sub-module. The mask deformation decoder includes a lightweight multi-layer perceptron network. The input of the mask deformation field network is the position coordinates of the up-sampling points on the rays in the training set HDR dynamic scene, the ray directions and the timestamp information, and the output is the mask confidence score of the up-sampling points on the rays in the training set HDR dynamic scene, indicating the confidence score that the up-sampling points on the rays belong to the target object. S51: Input the position coordinates, ray directions, and timestamp information of each sampling point on the rays in the training set HDR dynamic scene into the HexPlane sub-module. Based on the HexPlane plane interpolation method, obtain the features of each sampling point on the rays in the training set HDR dynamic scene. Through the MLP sub-module, fuse the features of all sampling points on the rays in the training set HDR dynamic scene to obtain the fused features . Based on the HexPlane plane interpolation method, decompose the 4D voxel space into six learnable parameter planes, as shown in the following formula: ; Among them, , , , represent the plane bilinear interpolation sequence numbers of the sampling point position coordinates and time stamps, , , , , , are all learnable parameter planes, , , are respectively the vectors orthogonal to the two planes of the bilinear interpolation, , , are respectively the dimensions of the plane for performing bilinear interpolation, represents element-wise multiplication; Using six parameter planes to represent a four-dimensional scene to reduce memory and computational overhead, the ray upsampled points on the , ray direction and time stamp information of the training set HDR dynamic scene are input into the HexPlane sub-module to obtain the feature output by the HexPlane sub-module, as shown in the following formula: ; Among them, is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate Bilinear interpolation of , ";" is the concatenation operation of the tensor; The MLP submodule uses a multi-layer perceptron network Rays in the training set HDR dynamic scene The characteristics of each sampling point Fusion is performed to obtain the fused features , as shown in the following formula: ; S52: Mask Deformable Decoder uses a lightweight multi-layer perceptron network to predict rays in the training set HDR dynamic scene The confidence score of each sampling point is obtained by mask inverse rendering, and the two-dimensional prediction mask of the target object in the HDR dynamic scene of the training set is obtained; Mapping function in HDR dynamic neural radiance field model based on S2 , the fused features obtained based on the spatiotemporal interpolation encoder and the ray direction Mapping to a periodic function , as shown in the following formula: ; in, To control the sampling point The highest frequency hyperparameters; Use a lightweight multi-layer perceptron network Predicting rays in the training set HDR dynamic scene Upsampling point Confidence score , as shown in the following formula: ; in, Rays in the training set HDR dynamic scene The set of confidence scores of all sampling points on , is the number of sampling points; S53: Use the pre-trained HDR dynamic neural radiance field to query the rays in the training set HDR dynamic scene The volume density of each sampling point , to obtain the opacity of the target object in the training set HDR dynamic scene, and obtain the rays in the pre-trained HDR dynamic neural radiation field The weights of the sampling points on , combined with the mask deformation field network prediction of the training set HDR dynamic scene rays Upsampling point Confidence score , through mask inverse rendering, the two-dimensional prediction mask of the target object in the training set HDR dynamic scene is obtained , as shown in the following formula: ; Use the two-dimensional segmentation mask of the target object in the training set HDR dynamic scene as the ground truth and the two-dimensional prediction mask to calculate the loss and achieve dynamic mask inverse rendering. The loss function is as shown in the following formula: ; ; ; where is the mask projection loss, is the total variance loss, is the ray set of image I, , , are all hyperparameters, is a hyperparameter, is the plane with coordinates in it, is the plane c with coordinates in it, is the plane c with coordinates in it, is the plane c with coordinates in it; Update the mask deformation field network parameters by the gradient descent method, where the model parameters of the neural radiance field are frozen during the segmentation process; S6: Use the training set to train the mask deformation field network. After reaching the preset number of training rounds, obtain the four-dimensional mask of the HDR dynamic scene of the segmented training set. Obtain the segmentation result of the target object in the HDR dynamic scene of the training set by querying. Use the test set to evaluate the quality of scene reconstruction and the quality of segmentation. Set the part greater than 0 in the queried two-dimensional mask to True and the part less than 0 to False to obtain the segmentation result of the target object in the HDR dynamic scene of the test set. Based on the segmentation result of the test set, the trained mask deformation field network is used to achieve semantic segmentation of any object in the HDR dynamic neural radiance field.
[0007] The beneficial effects of adopting the above technical solutions are as follows: A semantic segmentation method based on HDR dynamic neural radiance fields provided by the present invention makes full use of the advantages of two-dimensional vision large models in segmenting image sequences, expands the segmentation ability of two-dimensional models to four-dimensional space; proposes a tone mapping module controlled by hyperparameters to control the mapping from HDR views to LDR, greatly improving the segmentation ability in HDR scenarios; proposes a mask deformation field network and a dynamic mask inverse rendering method, which do not require expensive annotation of four-dimensional scenes and can segment any object in four-dimensional dynamic scenes using two-dimensional image masks. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 It is a flowchart of a semantic segmentation method based on HDR dynamic neural radiance fields provided by an embodiment of the present invention; Figure 2 It is a structural diagram of a mask deformation field network provided by an embodiment of the present invention; Figure 3 It is a schematic diagram of the mask inverse rendering process provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0009] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0010] A semantic segmentation method based on HDR dynamic neural radiance fields in this embodiment aims to solve the problem that the geometric information of objects in HDR dynamic neural radiance fields is implicitly stored in model parameters and it is difficult to achieve semantic segmentation. As Figure 1 shown, taking an HDR dynamic neural radiance field as a prior, using the HDR dynamic neural radiance field, reconstructing the HDR scene according to input images with different exposure rates, constructing a tone mapping module, mapping the multi-view reconstructed images of the HDR scene to LDR, using a two-dimensional vision segmentation auxiliary model to segment the target object masks of the LDR image set, constructing a mask deformation field network, and inverse rendering the masks segmented by the two-dimensional vision model to the four-dimensional dynamic neural radiance field to achieve the segmentation of any object in the four-dimensional dynamic neural radiance field, including the following steps: S1: Construct a scene data set, including scene picture files and scene label files, and divide the scene data set into a training set and a test set according to a ratio; The scene label file includes the file path, rotation matrix, translation vector, and timestamp information corresponding to each scene picture file. The training set is used for scene reconstruction based on the HDR dynamic neural radiance field, and the test set is used to evaluate the quality of scene reconstruction and the quality of segmentation; The methods for obtaining image files include, but are not limited to, downloading the HDR dynamic neural radiance field dataset from the Internet, constructing a virtual scene dataset based on 3D mapping software such as Blender, and using pose estimation methods such as COLMAP to estimate the poses of pictures taken in the real world to generate data in the input format of the neural radiance field; In this embodiment, the scene dataset is divided into a training set and a test set in a ratio of 8:2; S2: Using the HDR dynamic neural radiance field model, reconstruct the pictures with different exposure rates in the training set into a training set HDR dynamic scene with uniform exposure, and output a training set HDR picture sequence F1 based on the training set HDR dynamic scene, including n training set HDR pictures; Input the pictures with different exposure rates in the training set into the HDR dynamic neural radiance field model to obtain a training set HDR dynamic scene with uniform exposure; in the training set HDR dynamic scene, any ray is: ; Among them, is the camera origin, is the ray direction, is the distance from a point on the ray to the origin along the ray direction; For the sampling points on the ray in the training set HDR dynamic scene and the ray direction , calculate the spatial color value and the volume density of the ray based on the mapping function , and use the volume rendering method to obtain the integral of the ray in the training set HDR dynamic scene on the imaging plane and the integral of the opacity of the sampling points on the ray in the ray direction, which is used to represent the weight of the sampling points on the ray , as shown in the following formula: ; ; Among them, and respectively represent the near point and the far point of the sampling ray ; Output a training set HDR picture sequence F1 based on the training set HDR dynamic scene, including n training set HDR pictures; The exposure rate of the HDR dynamic neural radiance field in the single learning scenario is learned, and the HDR dynamic scene of the training set with uniform exposure is reconstructed. In this embodiment, the HDR-HexPlane dynamic neural radiance field is used to reconstruct the HDR dynamic scene of the training set with uniform exposure. Based on the reconstructed HDR dynamic scene of the training set, the training set HDR image sequence F1 is output, including n training set HDR pictures; S3: Construct a tone mapping module, input the training set HDR image sequence F1 into the tone mapping module, and obtain the training set LDR (low dynamic range) image sequence F2 with uniform exposure; Since scene capture in the real world usually includes inputs captured at different exposure rates, the HDR neural radiance field can contain more information about the scene. Since the segmentation ability of two-dimensional vision large models can only accept the input of LDR image sequences, this embodiment proposes a tone mapping module controlled by hyperparameters to map the HDR scene view to LDR, and the scene exposure and contrast mapping are controlled by hyperparameters to solve the problem of local over-dark or over-bright that may be learned in the HDR scene.
[0011] The tone mapping module includes a normalization layer, an adjustment layer, and an output layer; the normalization layer obtains the training set HDR image sequence F1 and normalizes the n training set HDR pictures included in the training set HDR image sequence F1, and uses the sigmoid function to map the pixel value of each pixel point in the training set HDR picture to between 0 and 1 to obtain the normalized training set HDR picture; the adjustment layer adjusts the brightness and contrast of the normalized training set HDR picture, where the adjustment of brightness and contrast is controlled by hyperparameters to reduce the model training parameters; the output layer maps the pixel value of each pixel point in the adjusted training set HDR picture to the LDR color value range of 0-255 to obtain the training set LDR low dynamic range image sequence F2 with uniform exposure; S31: Obtain the training set HDR image sequence F1 and normalize the n training set HDR pictures included in it, and use the sigmoid function to map the pixel value of each pixel point in the training set HDR picture to between 0 and 1 to obtain the normalized training set HDR picture, as shown in the following formula: ; where, is the pixel value of the pixel point in the training set HDR picture; S32: Adjust the brightness and contrast of the normalized training set HDR picture to obtain the adjusted training set HDR picture; Since there will be overexposed or underexposed areas when the reconstructed training set HDR dynamic scenes are compressed to the LDR low dynamic range, a tone mapping module that controls brightness and contrast with hyperparameters is used to adjust the brightness of the S31-normalized training set HDR images through an adjustment layer, as shown in the following formula: ; Among them, is the brightness hyperparameter, is the i th training set HDR image normalized before brightness adjustment, is the i th training set HDR image after brightness adjustment; Perform contrast adjustment on the training set HDR images after brightness adjustment, as shown in the following formula: ; Among them, is the contrast hyperparameter, is the i th training set HDR image after contrast adjustment, is the contrast enhancement function, as shown in the following formula: ; ; ; Among them, is the picture brightness function defined based on the human eye's sensitivity to the three primary colors, is the brightness enhancement function, is the pixel value of the pixel point on the red channel of the training set HDR image after brightness adjustment, is the pixel value of the pixel point on the green channel of the training set HDR image after brightness adjustment, is the pixel value of the pixel point on the blue channel of the training set HDR image after brightness adjustment; S33: Map the pixel value of each pixel point in the adjusted training set HDR image to the LDR color value range of 0 - 255 to obtain a sequence F2 of training set LDR low dynamic range images with uniform exposure; S4: Use a two-dimensional vision segmentation auxiliary model to segment the sequence F2 of training set LDR low dynamic range images with uniform exposure to obtain a two-dimensional segmentation mask of the target object in the training set HDR dynamic scene; Since the cue points belonging to one object at timestamp t0 in dynamic interactive segmentation may belong to other objects at timestamp t1, this embodiment proposes a two-dimensional vision model assistance module and a mask deformation field network to implement dynamic mask inverse rendering, so as to solve the problem of geometric inconsistency in the three-dimensional segmentation process.
[0012] S41: Select a two-dimensional vision large model as the two-dimensional vision segmentation assistance model, and input the training set LDR low-dynamic-range image sequence F2 into the two-dimensional vision large model; Since the two-dimensional vision large model SAM2 can segment any object in the image sequence, in this embodiment, the SAM2 model is used as the two-dimensional vision segmentation assistance model, and the assistance segmentation is realized based on the segmentation ability of the two-dimensional vision large model SAM2 for video or image sequences. The training set LDR low-dynamic-range image sequence F2 in.jpg format is input into the two-dimensional vision large model SAM2, and the model weight is sam2_hiera_large.pt; S42: Set a number of segmentation cue points and background points on the target object in the training set HDR dynamic scene, and use the two-dimensional vision segmentation assistance model to obtain the two-dimensional segmentation mask of the target object in the training set HDR dynamic scene; In this embodiment, when selecting a number of cue points and background points on the target object in the HDR dynamic scene, based on the two-dimensional vision large model SAM2, the number of segmentation cue points is set to be no less than 4. The more background points, the more accurate the segmentation mask can be obtained; For the RGB three-channel image with size H×W×3 in the training set LDR low-dynamic-range image sequence F2, where H is the height of the training set LDR low-dynamic-range image and W is the width of the training set LDR low-dynamic-range image, obtain its two-dimensional segmentation mask is H×W×1, and the two-dimensional segmentation mask stores a single-channel boolean value for each pixel point in H×W×1. The masked area is True, and the boolean value of the pixel points in this area is 1. The unmasked area is False, and the boolean value of the pixel points in this area is 0. The segmentation mask of each image in the training set LDR low-dynamic-range image sequence F2 is stored in a numpy file with a.npy format; S43: If the two-dimensional vision segmentation assistance model cannot achieve the desired segmentation effect, then adjust the brightness hyperparameter in S32 and the contrast hyperparameter to optimize the segmentation effect, and repeat S41 to S42 to obtain the two-dimensional segmentation mask sequence with the best segmentation effect for the target object in the HDR dynamic scene; S5: Construct a mask deformation field network, input the sequence of 2D segmentation masks with the best segmentation effect of the target object in the HDR dynamic scene obtained by the 2D segmentation model auxiliary model into the mask deformation field network for mask inverse rendering, project the 2D segmentation masks into the four-dimensional space, obtain the four-dimensional segmentation masks of the target object in the training set HDR dynamic scene, and realize the segmentation of the target object in the training set HDR dynamic scene; The mask deformation field network is as Figure 2 shown, including a spatio-temporal interpolation encoder and a mask deformation decoder. The spatio-temporal interpolation encoder network includes a HexPlane sub-module and an MLP sub-module. The mask deformation decoder includes a lightweight multi-layer perceptron network. The input of the mask deformation field network is the position coordinates of the up-sampling points on the rays in the training set HDR dynamic scene, the ray directions and the timestamp information, and the output is the mask confidence score of the up-sampling points on the rays in the training set HDR dynamic scene, indicating the confidence score that the up-sampling points on the rays belong to the target object. The mask inverse rendering process is as Figure 3 shown.
[0013] S51: Input the position coordinates, ray directions and timestamp information of each sampling point on the rays in the training set HDR dynamic scene into the HexPlane sub-module. Based on the HexPlane plane interpolation method, obtain the features of each sampling point on the rays in the training set HDR dynamic scene. Through the MLP sub-module, fuse the features of all sampling points on the rays in the training set HDR dynamic scene to obtain the fused features ; Decompose the 4D voxel space into six learnable parameter planes based on the HexPlane plane interpolation method, as shown in the following formula: ; Among them, , , , represent the plane bilinear interpolation sequence numbers of the sampling point position coordinates and the timestamp, , , , , , are all learnable parameter planes, , , are vectors orthogonal to two planes of bilinear interpolation respectively, , , are the dimensions of the plane for bilinear interpolation respectively, represents element-wise multiplication; Using six parameter planes to represent a four-dimensional scene to reduce memory and computational overhead, the position coordinates of the upsampling points on the ray in the training set HDR dynamic scene, the ray direction and the timestamp information are input into the HexPlane sub-module to obtain the feature output by the HexPlane sub-module, as shown in the following formula: ; where, is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , and ";" is the operation of concatenating tensors; The MLP sub-module uses a multi-layer perceptron network to fuse the features of each sampling point on the ray in the training set HDR dynamic scene to obtain the fused feature , as shown in the following formula: ; S52: The mask deformation decoder uses a lightweight multi-layer perceptron network to predict the confidence scores of each sampling point on the rays in the training set HDR dynamic scene, and obtains the two-dimensional predicted mask of the target object in the training set HDR dynamic scene through mask inverse rendering; Using only the position coordinates of the sampling points on the rays in the training set HDR dynamic scene and the ray directions is not sufficient to fully represent the geometric information in the HDR dynamic scene. Therefore, based on the mapping function in the HDR dynamic neural radiance field model used in S2, the fused features obtained based on the spatio-temporal interpolation encoder and the ray directions are mapped to a periodic function as shown in the following formula: ; where is the hyperparameter that controls the highest frequency of the sampling point ; A lightweight multi-layer perceptron network is used to predict the confidence scores of the sampling points on the rays in the training set HDR dynamic scene as shown in the following formula: ; where is the set of confidence scores of all sampling points on the ray in the training set HDR dynamic scene, and is the number of sampling points; S53: Using the pre-trained HDR dynamic neural radiance field, query the volume density of each sampling point on the rays in the training set HDR dynamic scene to obtain the opacity of the target object in the training set HDR dynamic scene, and obtain the weight values of the sampling points on the ray in the pre-trained HDR dynamic neural radiance field. Combining the confidence scores of the sampling points on the ray in the training set HDR dynamic scene predicted by the mask deformation field network, the two-dimensional predicted mask of the target object in the training set HDR dynamic scene is obtained through mask inverse rendering as shown in the following formula: ; The two-dimensional segmentation mask of the target object in the training set HDR dynamic scene is obtained by querying the volume density of each sampling point on the ray in the pre-trained HDR dynamic neural radiance field to obtain the opacity of the target object in the training set HDR dynamic scene, and obtaining the weight values of the sampling points on the ray in the pre-trained HDR dynamic neural radiance field. Combining the confidence scores of the sampling points on the ray in the training set HDR dynamic scene predicted by the mask deformation field network, the two-dimensional predicted mask of the target object in the training set HDR dynamic scene is obtained through mask inverse rendering as shown in the following formula: ; The two-dimensional segmentation mask of the target object in the training set HDR dynamic scene As the ground truth and the two-dimensional prediction mask Calculate the loss to achieve dynamic mask inverse rendering. The loss function is as shown in the following formula: ; ; ; where, is the mask projection loss, is the total variance loss, is the ray set of image I, , , are all hyperparameters, is a hyperparameter, is the plane with coordinates in it, is the plane c with coordinates in it, is the plane c with coordinates in it, is the plane c with coordinates in it; Update the mask deformation field network parameters by the gradient descent method, where the model parameters of the neural radiance field are frozen during the segmentation process.
[0014] S6: Use the training set to train the mask deformation field network. After reaching the preset number of training rounds, obtain the four-dimensional mask of the segmented HDR dynamic scene of the training set. Obtain the segmentation result of the target object in the HDR dynamic scene of the training set by querying. Use the test set to evaluate the quality of scene reconstruction and the quality of segmentation. Set the part greater than 0 in the queried two-dimensional mask to True and the part less than 0 to False to obtain the segmentation result of the target object in the HDR dynamic scene of the test set. Based on the segmentation result of the test set, the trained mask deformation field network realizes semantic segmentation of any object in the HDR dynamic neural radiance field; After 5000 rounds of iterative training, a segmented four-dimensional mask is obtained. The projection under each query view is a single-channel confidence score map. Set the part with a score greater than 0 to True and vice versa to False. Thus, the mask of the target object can be obtained under any view.
[0015] In this embodiment, the HDR-HexPlane dataset is used to verify the semantic segmentation method based on the HDR dynamic neural radiance field, and the average IoU reaches 86.63% and the average precision reaches 99.49%.
[0016] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A semantic segmentation method based on HDR dynamic neural radiation field, characterized by: The following steps are involved: S1: Build a scene dataset, including scene image files and scene label files, and divide the scene dataset into a training set and a test set in proportion; S2: Use the HDR dynamic neural radiation field model to reconstruct the images with different exposure rates in the training set into the training set HDR dynamic scenes with uniform exposure, and output the training set HDR image sequence F1 based on the training set HDR dynamic scenes, including n HDR images of the training set; S3: construct a tone mapping module, input the training set HDR image sequence F1 into the tone mapping module, and obtain a training set LDR low dynamic range image sequence F2 with uniform exposure; S4: Use the two-dimensional visual segmentation auxiliary model to segment the evenly exposed training set LDR low dynamic range image sequence F2 to obtain the two-dimensional segmentation mask of the target object in the training set HDR dynamic scene; S5: construct a mask deformation field network, input the best 2D segmentation mask sequence obtained by the 2D segmentation model auxiliary model into the mask deformation field network for mask inverse rendering, project the 2D segmentation mask into the four-dimensional space, obtain the four-dimensional segmentation mask of the target object in the HDR dynamic scene of the training set, and achieve the segmentation of the target object in the HDR dynamic scene of the training set; S6: Use the training set to train the mask deformation field network, and obtain the four-dimensional mask of the HDR dynamic scene of the segmented training set after reaching the preset training rounds. Obtain the segmentation result of the target object in the HDR dynamic scene of the training set by query, and use the test set to evaluate the quality of scene reconstruction and segmentation. Set the part of the queried two-dimensional mask greater than 0 to True and the part less than 0 to False, and obtain the segmentation result of the target object in the HDR dynamic scene of the test set. Based on the segmentation result of the test set, obtain the trained mask deformation field network to realize the semantic segmentation of any object in the HDR dynamic neural radiation field.
2. The semantic segmentation method based on HDR dynamic neural radiation field according to claim 1, characterized in that: The scene label file S1 includes the file path, rotation matrix, translation vector and timestamp information corresponding to each scene image file. The training set is used for scene reconstruction based on HDR dynamic neural radiation field, and the test set is used to evaluate the quality of scene reconstruction and segmentation.
3. The semantic segmentation method based on HDR dynamic neural radiation field according to claim 2, characterized in that: The specific method of S2 is: The images with different exposure rates in the training set are input into the HDR dynamic neural radiation field model to obtain the training set HDR dynamic scene with uniform exposure; in the training set HDR dynamic scene, any ray for: ; in, is the camera origin, is the ray direction, is the distance from a point on the ray to the origin along the ray direction; Rays in the training set HDR dynamic scene The sampling points on and the ray direction , based on the mapping function Calculate Rays The spatial color value With body density , using the volume rendering method to obtain the rays in the training set HDR dynamic scene Integration on the imaging plane and rays The sampling points on Integral of opacity in the direction of the ray , used to represent rays The sampling points on The weight is as shown in the following formula: ; ; in, and Represents the sampling rays the near and far points; Based on the training set HDR dynamic scene, the training set HDR picture sequence F1 is output, including n Training set HDR images.
4. The semantic segmentation method based on HDR dynamic neural radiation field according to claim 3, characterized in that: S3 The tone mapping module includes a normalization layer, an adjustment layer, and an output layer; the normalization layer obtains the training set HDR picture sequence F1, and performs a tone mapping operation on the HDR picture sequence F1. n The training set HDR pictures are normalized, and the sigmoid function is used to map the pixel value of each pixel in the training set HDR pictures to between 0 and 1 to obtain the normalized training set HDR pictures; the adjustment layer adjusts the brightness and contrast of the normalized training set HDR pictures, where the adjustment of brightness and contrast is controlled by hyperparameters to reduce model training parameters; the output layer maps the pixel value of each pixel in the adjusted training set HDR pictures to the LDR color value domain ranging from 0 to 255 to obtain a training set LDR low dynamic range picture sequence F2 with uniform exposure.
5. The semantic segmentation method based on HDR dynamic neural radiation field according to claim 4, characterized in that: The S3 includes: S31: Get the training set HDR image sequence F1 and analyze the n The training set HDR images are normalized, and the sigmoid function is used to map the pixel value of each pixel in the training set HDR image to between 0 and 1 to obtain the normalized training set HDR image, as shown in the following formula: ; in, is the pixel value of the pixel in the training set HDR image; S32: adjusting the brightness and contrast of the normalized training set HDR image to obtain an adjusted training set HDR image; The brightness of the S31 normalized training set HDR image is adjusted using the tone mapping module adjustment layer with brightness and contrast controlled by hyperparameters, as shown in the following formula: ; in, is the brightness hyperparameter, is the normalized value before brightness adjustment i training set HDR images, After brightness adjustment i HDR images of the training set; For the training set HDR images after adjusting the brightness Adjust the contrast as shown in the following formula: ; in, is the contrast hyperparameter, After contrast adjustment i training set HDR images, To enhance the contrast function, the formula is as follows: ; ; ; in, is the image brightness function defined based on the sensitivity of the human eye to the three primary colors. To enhance the brightness function, The training set HDR images after brightness adjustment The pixel value of the pixel on the red channel, The training set HDR images after brightness adjustment The pixel value of the pixel on the green channel, The training set HDR images after brightness adjustment The pixel value of the pixel in the blue channel; S33: Map the pixel value of each pixel in the adjusted training set HDR image to an LDR color value domain ranging from 0 to 255 to obtain a training set LDR low dynamic range image sequence F2 with uniform exposure.
6. The semantic segmentation method based on HDR dynamic neural radiation field according to claim 5, characterized in that: The S4 includes: S41: Selecting a large two-dimensional visual model as a two-dimensional visual segmentation auxiliary model, and inputting the training set LDR low dynamic range image sequence F2 into the large two-dimensional visual model; S42: setting a number of segmentation cue points and background points on the target object in the HDR dynamic scene of the training set, and obtaining a two-dimensional segmentation mask of the target object in the HDR dynamic scene of the training set using a two-dimensional visual segmentation auxiliary model; S43: If the two-dimensional visual segmentation auxiliary model cannot achieve the desired segmentation effect, adjust the brightness hyperparameter in S32 and the contrast hyperparameter To optimize the segmentation effect, S41 to S42 are repeated to obtain a two-dimensional segmentation mask sequence with the best segmentation effect of the target object in the HDR dynamic scene.
7. The semantic segmentation method based on HDR dynamic neural radiation field according to claim 6, characterized in that: S42 specific method is: For the RGB three-channel image of size H×W×3 in the training set LDR low dynamic range image sequence F2, where H is the height of the training set LDR low dynamic range image and W is the width of the training set LDR low dynamic range image, its two-dimensional segmentation mask is obtained H×W×1, two-dimensional segmentation mask A single-channel Boolean value is stored for each pixel in H×W×1. The masked area is True, and the Boolean value of the pixel in this area is 1. The unmasked area is False, and the Boolean value of the pixel in this area is 0.
8. The semantic segmentation method based on HDR dynamic neural radiation field according to claim 7, characterized in that: S5 The mask deformation field network includes a spatiotemporal interpolation encoder and a mask deformation decoder, the spatiotemporal interpolation encoder network includes a HexPlane submodule and an MLP submodule, and the mask deformation decoder includes a lightweight multi-layer perceptron network; the input of the mask deformation field network is the ray in the training set HDR dynamic scene Upsampling point Location coordinates , ray direction And timestamp information, output as rays in the training set HDR dynamic scene Upsampling point The mask confidence score of , which means ray Upsampling point Confidence score that the object belongs to the target.
9. The semantic segmentation method based on HDR dynamic neural radiation field according to claim 8, characterized in that: The S5 includes: S51: Rays in the training set HDR dynamic scene The position coordinates, ray direction and timestamp information of each sampling point are input into the HexPlane submodule, and the ray in the HDR dynamic scene of the training set is obtained based on the HexPlane plane interpolation method. The characteristics of each sampling point , through the MLP submodule, the rays in the HDR dynamic scene of the training set The features of all sampling points are fused to obtain the fused features ; S52: Mask Deformable Decoder uses a lightweight multi-layer perceptron network to predict rays in the training set HDR dynamic scene The confidence score of each sampling point is obtained by mask inverse rendering, and the two-dimensional prediction mask of the target object in the HDR dynamic scene of the training set is obtained; Mapping function in HDR dynamic neural radiance field model based on S2 , the fused features obtained based on the spatiotemporal interpolation encoder and the ray direction Mapping to a periodic function , as shown in the following formula: ; in, To control the sampling point The highest frequency hyperparameters; Use a lightweight multi-layer perceptron network Predicting rays in the training set HDR dynamic scene Upsampling point Confidence score , as shown in the following formula: ; in, Rays in the training set HDR dynamic scene The set of confidence scores of all sampling points on , is the number of sampling points; S53: Use the pre-trained HDR dynamic neural radiance field to query the rays in the training set HDR dynamic scene The volume density of each sampling point , to obtain the opacity of the target object in the training set HDR dynamic scene, and obtain the rays in the pre-trained HDR dynamic neural radiation field The weights of the sampling points on , combined with the mask deformation field network prediction of the training set HDR dynamic scene rays Upsampling point Confidence score , through mask inverse rendering, the two-dimensional prediction mask of the target object in the training set HDR dynamic scene is obtained , as shown in the following formula: ; The 2D segmentation mask of the target object in the training set HDR dynamic scene As the true value and the 2D prediction mask Calculate loss, implement dynamic mask inverse rendering, loss function As shown in the following formula: ; ; ; in, is the mask projection loss, is the total variance loss, is the ray set of image I, , , are all hyperparameters, is a hyperparameter, For plane The median coordinate is point, For plane c The median coordinate is point, For plane c The median coordinate is point, For plane c The median coordinate is point; The mask deformation field network parameters are updated by gradient descent, where the model parameters of the neural radiance field are frozen during the segmentation process.
10. The semantic segmentation method based on HDR dynamic neural radiation field according to claim 9, characterized in that: The specific method of S51 is: Based on the HexPlane interpolation method, the 4D voxel space Decomposed into six learnable parameter planes, as shown in the following formula: ; in, , , , The plane bilinear interpolation sequence number representing the sampling point position coordinates and timestamp, , , , , , are all learnable parameter planes, , , are vectors orthogonal to the two planes of bilinear interpolation, , , are the dimensions of the plane for bilinear interpolation, Represents element-wise multiplication; Six parameter planes are used to represent the four-dimensional scene to reduce memory and computational overhead. Upsampling point Location coordinates , ray direction And the timestamp information is input into the HexPlane submodule to obtain the features output by the HexPlane submodule , as shown in the following formula: ; in, For sampling point exist Coordinate-based Bilinear interpolation of For sampling point In plane Based on coordinates Bilinear interpolation of For sampling point exist Coordinate-based Bilinear interpolation of For sampling point In plane Based on coordinates Bilinear interpolation of For sampling point In plane Based on coordinates Bilinear interpolation of For sampling point exist Coordinate-based Bilinear interpolation of , ";" is the concatenation operation of tensors; The MLP submodule uses a multi-layer perceptron network Rays in the training set HDR dynamic scene The characteristics of each sampling point Fusion is performed to obtain the fused features , as shown in the following formula: 。
Citation Information
Patent Citations
Image high dynamic range reconstruction method based on deep learning
CN111292264A
High-dynamic three-dimensional measurement method based on overexposure connected domain intensity adaptive distribution
CN116310101A
Image generation method and device, storage medium and chip
CN117392030A
Large-scale three-dimensional scene semantics and building segmentation method based on neural radiation representation
CN117726810A
High Dynamic Range View Synthesis from Noisy Raw Images
US20250037244A1
Cited By
Bubble graph generation method based on visual segmentation and multi-modal large model
CN120912711A
Screen rendering processing system and method based on visual technology
CN121685806A