A semantic segmentation method based on HDR dynamic neural radiance fields
Through the tone mapping module and mask deformation network of HDR dynamic neural radiation field, combined with the two-dimensional visual model, the problem of lighting influence and cross-view inconsistency in dynamic scenes is solved, and efficient semantic segmentation of dynamic scenes is achieved.
Patent Information
- Application Number
- CN202510535173.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When processing image segmentation of dynamic scenes, the prior art is greatly affected by lighting conditions and it is difficult to effectively perform semantic segmentation across views, especially in the segmentation of dynamic objects, there is a problem of inconsistent propagation of cue points.
Using the method based on HDR dynamic neural radiation field, the tone mapping module and mask deformation field network are constructed, and the two-dimensional visual model is used for segmentation to achieve expansion from two-dimensional to four-dimensional, and the semantic segmentation of dynamic scenes is combined with mask inverse rendering technology.
It improves the segmentation ability in HDR scenes, can accurately segment any object in dynamic scenes, reduces the expensive annotation requirement for four-dimensional scenes, and uses two-dimensional image masks to segment four-dimensional dynamic scenes.
Smart Images

Figure CN120070898B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular, to a semantic segmentation method based on HDR dynamic neural radiance fields. Background Art
[0002] Neural Radiance Field (NeRF) is a deep learning model for three-dimensional implicit space modeling. Based on the scene images captured under known views and the internal and external camera parameters, it predicts the color value and volume density of each sampling point through a neural network, and synthesizes images under new views based on the volume rendering method, showing great success in representing 3D scenes and synthesizing new view images. NeRF represents the three-dimensional space as a set of learnable and continuous radiance fields. Given the position of the input sampling point and the ray direction, it outputs the color and volume density of the sampling point. HDR (High Dynamic Range) technology can capture images with different brightness levels and present images with a larger exposure range than traditional images. Using images with different exposures as input, the HDR neural radiance field can reconstruct a uniformly exposed HDR scene.
[0003] For the image segmentation task, SA3D attempts to combine an arbitrary segmentation model (SAM) with a neural radiance field to segment 3D objects. It provides manual segmentation hint points for the target object in a single view, and then SA3D alternately performs masked inverse rendering and cross-view self-hinting between various views to iteratively complete the 3D mask of the target object constructed with a voxel grid. Under the volume density distribution of the neural radiance field, SA3D projects the 2D mask obtained by SAM in the current view into 3D space and automatically extracts reliable hints from the 2D view rendered by the neural radiance field in another view as the input of SAM.
[0004] However, the input of SAM needs to maintain uniform illumination and contain sufficient scene information. In real-world datasets, affected by lighting conditions, the input images always contain overexposed or underexposed areas, which may significantly lead to cross-view inconsistencies and have an adverse impact on the learning processes of the neural radiance field and SAM. In addition, previous methods mainly target static scenes, which are limited in segmenting 4D dynamic objects. A serious problem in dynamic interactive segmentation is that the hint points belonging to one object at timestamp t0 may belong to other objects at timestamp t1, making it difficult to propagate the hint points to other timestamps. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a semantic segmentation method based on HDR dynamic neural radiance fields in view of the deficiencies of the above-mentioned prior art. By utilizing the segmentation ability of a two-dimensional vision large model on pictures and videos, the segmentation ability of the two-dimensional model is extended to four dimensions. Based on the constructed tone mapping module and mask deformation field network, mask inverse rendering is realized, and then semantic segmentation of any object in the HDR dynamic neural radiance field is achieved.
[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0007] The present invention provides a semantic segmentation method based on HDR dynamic neural radiance fields, including the following steps:
[0008] S1: Construct a scene dataset, including scene picture files and scene label files, and divide the scene dataset into a training set and a test set according to a certain proportion;
[0009] The scene label file includes the file path, rotation matrix, translation vector, and timestamp information corresponding to each scene picture file. The training set is used for scene reconstruction based on the HDR dynamic neural radiance field, and the test set is used to evaluate the quality of scene reconstruction and the quality of segmentation;
[0010] S2: Use the HDR dynamic neural radiance field model to reconstruct the pictures with different exposure rates in the training set into a training set HDR dynamic scene with uniform exposure, and output a training set HDR picture sequence F1 based on the training set HDR dynamic scene, including n training set HDR pictures;
[0011] Input the pictures with different exposure rates in the training set into the HDR dynamic neural radiance field model to obtain a training set HDR dynamic scene with uniform exposure; in the training set HDR dynamic scene, any ray is:
[0012] ;
[0013] wherein, is the camera origin, is the ray direction, is the distance from a point on the ray to the origin along the ray direction;
[0014] For the sampling points and the ray direction on the ray in the training set HDR dynamic scene, based on the mapping function calculate the spatial color value and the volume density of the ray , and use the volume rendering method to obtain the ray Integration on the imaging plane and rays sampling points on the integration of opacity in the ray direction , used to represent the ray sampling points on weights, as shown in the following formula:
[0015] ;
[0016] ;
[0017] wherein and respectively represent the near point and the far point of the sampling ray ;
[0018] Based on the training set HDR dynamic scene, output the training set HDR picture sequence F1, including n training set HDR pictures;
[0019] S3: Construct a tone mapping module, input the training set HDR picture sequence F1 into the tone mapping module, and obtain the training set LDR low dynamic range picture sequence F2 with uniform exposure;
[0020] The tone mapping module includes a normalization layer, an adjustment layer, and an output layer; the normalization layer obtains the training set HDR picture sequence F1 and normalizes the n training set HDR pictures included in the training set HDR picture sequence F1, uses the sigmoid function to map the pixel value of each pixel point in the training set HDR picture to between 0 and 1, and obtains the normalized training set HDR picture; the adjustment layer adjusts the brightness and contrast of the normalized training set HDR picture, where the adjustment of brightness and contrast is controlled by hyperparameters to reduce the model training parameters; the output layer maps the pixel value of each pixel point in the adjusted training set HDR picture to the LDR color value range of 0 - 255, and obtains the training set LDR low dynamic range picture sequence F2 with uniform exposure;
[0021] S31: Obtain the training set HDR picture sequence F1 and normalize the n training set HDR pictures included therein, use the sigmoid function to map the pixel value of each pixel point in the training set HDR picture to between 0 and 1, and obtain the normalized training set HDR picture, as shown in the following formula:
[0022] ;
[0023] wherein is the pixel value of the pixel point in the training set HDR picture;
[0024] S32: Adjust the brightness and contrast of the normalized training set HDR images to obtain the adjusted training set HDR images;
[0025] Use the tone mapping module adjustment layer with brightness and contrast controlled by hyperparameters to adjust the brightness of the training set HDR images normalized in S31, as shown in the following formula:
[0026] ;
[0027] where is the brightness hyperparameter, is the i th training set HDR image normalized before brightness adjustment, is the i th training set HDR image after brightness adjustment;
[0028] Adjust the contrast of the training set HDR images after brightness adjustment, as shown in the following formula:
[0029] ;
[0030] where is the contrast hyperparameter, is the i th training set HDR image after contrast adjustment, is the function to enhance contrast, as shown in the following formula:
[0031] ;
[0032] ;
[0033] ;
[0034] where is the picture brightness function defined based on the human eye's sensitivity to the three primary colors, is the function to enhance brightness, is the training set HDR image after brightness adjustment, and is the pixel value of the pixel point on the red channel of the training set HDR image after brightness adjustment, is the pixel value of the pixel point on the green channel of the training set HDR image after brightness adjustment,
[0035] S33: Map the pixel values of each pixel in the adjusted training set HDR images to the LDR color value range of 0 - 255 to obtain a sequence of evenly-exposed training set LDR low-dynamic-range images F2;
[0036] S4: Use a two-dimensional vision segmentation auxiliary model to segment the sequence of evenly-exposed training set LDR low-dynamic-range images F2 to obtain a two-dimensional segmentation mask of the target object in the training set HDR dynamic scene;
[0037] S41: Select a two-dimensional vision large model as the two-dimensional vision segmentation auxiliary model, and input the sequence of training set LDR low-dynamic-range images F2 into the two-dimensional vision large model;
[0038] S42: Set a number of segmentation hint points and background points on the target object in the training set HDR dynamic scene, and use the two-dimensional vision segmentation auxiliary model to obtain a two-dimensional segmentation mask of the target object in the training set HDR dynamic scene;
[0039] For the RGB three-channel image with dimensions H×W×3 in the sequence of training set LDR low-dynamic-range images F2, where H is the height of the training set LDR low-dynamic-range image and W is the width of the training set LDR low-dynamic-range image, obtain its two-dimensional segmentation mask which is H×W×1, and the two-dimensional segmentation mask stores a single-channel boolean value for each pixel point in H×W×1. The area with the mask is True, and the boolean value of the pixel points in this area is 1. The area without the mask is False, and the boolean value of the pixel points in this area is 0;
[0040] S43: If the two-dimensional vision segmentation auxiliary model fails to achieve the desired segmentation effect, then adjust the brightness hyperparameter and the contrast hyperparameter in S32 to optimize the segmentation effect, and repeat S41 to S42 to obtain the sequence of two-dimensional segmentation masks with the best segmentation effect for the target object in the HDR dynamic scene;
[0041] S5: Construct a mask deformation field network, input the sequence of two-dimensional segmentation masks with the best segmentation effect for the target object in the HDR dynamic scene obtained by the two-dimensional segmentation model auxiliary model into the mask deformation field network for mask inverse rendering, project the two-dimensional segmentation mask into the four-dimensional space, and obtain a four-dimensional segmentation mask of the target object in the training set HDR dynamic scene to achieve the segmentation of the target object in the training set HDR dynamic scene;
[0042] The mask deformation field network includes a spatio-temporal interpolation encoder and a mask deformation decoder. The spatio-temporal interpolation encoder network includes a HexPlane sub-module and an MLP sub-module. The mask deformation decoder includes a lightweight multi-layer perceptron network. The input of the mask deformation field network is the position coordinates, ray direction, and timestamp information of the upsampling points on the rays in the training set HDR dynamic scene. The output is the mask confidence score of the upsampling points on the rays in the training set HDR dynamic scene, representing the confidence score that the upsampling points on the rays belong to the target object. Upsampling points of the position coordinates and the ray direction as well as the timestamp information, and the output is the mask confidence score of the upsampling points on the rays in the training set HDR dynamic scene Upsampling points indicating the confidence score that the upsampling points on the ray belong to the target object; Upsampling points ;
[0043] S51: Input the position coordinates, ray direction, and timestamp information of each sampling point on the ray in the training set HDR dynamic scene into the HexPlane sub-module. Based on the HexPlane plane interpolation method, obtain the features of each sampling point on the ray in the training set HDR dynamic scene. Through the MLP sub-module, fuse the features of all sampling points on the ray in the training set HDR dynamic scene to obtain the fused features ; Features of each sampling point , and through the MLP sub-module, fuse the features of all sampling points on the ray in the training set HDR dynamic scene to obtain the fused features ; ;
[0044] Based on the HexPlane plane interpolation method, decompose the 4D voxel space into six learnable parameter planes, as shown in the following formula:
[0045] ;
[0046] Among them, , , , represent the plane bilinear interpolation sequence numbers of the sampling point position coordinates and timestamp, , , , , , are all learnable parameter planes, , , are respectively the vectors orthogonal to the two planes of the bilinear interpolation, , , are respectively the dimensions of the planes for bilinear interpolation, represents element-wise multiplication;
[0047] Six parameter planes are used to represent the four-dimensional scene to reduce memory and computational overhead. Upsampling point Location coordinates , ray direction And the timestamp information is input into the HexPlane submodule to obtain the features output by the HexPlane submodule , as shown in the following formula:
[0048] ;
[0049] in, For sampling point exist Coordinate-based Bilinear interpolation of For sampling point In plane Based on coordinates Bilinear interpolation of For sampling point exist Coordinate-based Bilinear interpolation of For sampling point In plane Based on coordinates Bilinear interpolation of For sampling point In plane Based on coordinates Bilinear interpolation of For sampling point exist Coordinate-based Bilinear interpolation of , ";" is the concatenation operation of the tensor;
[0050] The MLP submodule uses a multi-layer perceptron network Rays in the training set HDR dynamic scene The characteristics of each sampling point Fusion is performed to obtain the fused features , as shown in the following formula:
[0051] ;
[0052] S52: Mask Deformable Decoder uses a lightweight multi-layer perceptron network to predict rays in the training set HDR dynamic scene The confidence score of each sampling point is obtained by mask inverse rendering, and the two-dimensional prediction mask of the target object in the HDR dynamic scene of the training set is obtained;
[0053] Mapping function in the HDR dynamic neural radiance field model used based on S2 Fuse the features obtained based on the spatio-temporal interpolation encoder and the ray direction and map them to a periodic function as shown in the following formula:
[0054] ;
[0055] where is a hyperparameter that controls the highest frequency of the sampling points ;
[0056] Use a lightweight multi-layer perceptron network to predict the confidence score of the upsampling points on the rays in the training set HDR dynamic scene as shown in the following formula:
[0057] ;
[0058] where is the set of confidence scores of all sampling points on the rays in the training set HDR dynamic scene is the number of sampling points;
[0059] S53: Use the pre-trained HDR dynamic neural radiance field to query the volume density of each sampling point on the rays in the training set HDR dynamic scene to obtain the opacity of the target object in the training set HDR dynamic scene, and get the weight values of the sampling points on the rays in the pre-trained HDR dynamic neural radiance field. Combine the confidence scores of the upsampling points on the rays in the training set HDR dynamic scene predicted by the mask deformation field network to obtain the two-dimensional predicted mask of the target object in the training set HDR dynamic scene through mask inverse rendering
[0060] ;
[0061] Use the two-dimensional segmentation mask of the target object in the training set HDR dynamic scene as the ground truth and the two-dimensional predicted mask to calculate the loss, and implement dynamic mask inverse rendering. The loss function is as shown in the following formula:
[0062] ;
[0063] ;
[0064] ;
[0065] wherein, is the mask projection loss, is the total variance loss, is the ray set of picture I, , , are all hyperparameters, is a hyperparameter, is the plane with coordinates in it, is the plane c with coordinates in it, is the plane c with coordinates in it, is the plane c with coordinates in it;
[0066] Update the mask deformation field network parameters by the gradient descent method, where the model parameters of the neural radiance field are frozen during the segmentation process;
[0067] S6: Use the training set to train the mask deformation field network. After reaching the preset number of training rounds, obtain the four-dimensional mask of the segmented training set's HDR dynamic scene. Obtain the segmentation result of the target object in the training set's HDR dynamic scene by querying. Use the test set to evaluate the quality of scene reconstruction and the quality of segmentation. Set the part greater than 0 in the queried two-dimensional mask to True and the part less than 0 to False to obtain the segmentation result of the target object in the test set's HDR dynamic scene. Based on the segmentation result of the test set, obtain the trained mask deformation field network to achieve semantic segmentation of any object in the HDR dynamic neural radiance field.
[0068] The beneficial effects of adopting the above technical solutions are as follows: A semantic segmentation method based on HDR dynamic neural radiance field provided by the present invention makes full use of the advantages of two-dimensional vision large models in segmenting picture sequences, expands the segmentation ability of two-dimensional models to four-dimensional space; proposes a hyperparameter-controlled tone mapping module to control the mapping from HDR views to LDR, greatly improving the segmentation ability in HDR scenes; proposes a mask deformation field network and a dynamic mask inverse rendering method, which do not require expensive annotation of four-dimensional scenes and can perform segmentation of any object in four-dimensional dynamic scenes using two-dimensional picture masks. Description of the Drawings
[0069] Figure 1 It is a flowchart of a semantic segmentation method based on HDR dynamic neural radiance fields provided by an embodiment of the present invention;
[0070] Figure 2 It is a structural diagram of a mask deformation field network provided by an embodiment of the present invention;
[0071] Figure 3 It is a schematic diagram of the mask inverse rendering process provided by an embodiment of the present invention. Detailed Embodiments
[0072] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0073] A semantic segmentation method based on HDR dynamic neural radiance fields in this embodiment aims to solve the problem that the geometric information of objects in HDR dynamic neural radiance fields is implicitly stored in model parameters and it is difficult to achieve semantic segmentation. As Figure 1 shown, taking an HDR dynamic neural radiance field as a prior, using the HDR dynamic neural radiance field, reconstructing an HDR scene according to input images with different exposure rates, constructing a tone mapping module, mapping the multi-view reconstructed pictures of the HDR scene to LDR, using a two-dimensional vision segmentation auxiliary model to segment the target object masks of the LDR picture set, constructing a mask deformation field network, and inversely rendering the masks segmented by the two-dimensional vision model to the four-dimensional dynamic neural radiance field to achieve the segmentation of any object in the four-dimensional dynamic neural radiance field, including the following steps:
[0074] S1: Construct a scene data set, including scene picture files and scene label files, and divide the scene data set into a training set and a test set according to a ratio;
[0075] The scene label file includes the file path, rotation matrix, translation vector, and timestamp information corresponding to each scene picture file. The training set is used for scene reconstruction based on the HDR dynamic neural radiance field, and the test set is used to evaluate the quality of scene reconstruction and the quality of segmentation;
[0076] The methods for obtaining picture files include but are not limited to downloading HDR dynamic neural radiance field data sets from the Internet, constructing virtual scene data sets based on 3D mapping software such as blender, and using pose estimation methods such as COLMAP to estimate the poses of pictures taken in the real world to generate data in the input format of neural radiance fields;
[0077] In this embodiment, the scene data set is divided into a training set and a test set according to a ratio of 8:2;
[0078] S2: Use the HDR dynamic neural radiance field model to reconstruct the images with different exposure rates in the training set into a training set HDR dynamic scene with uniform exposure, and output a training set HDR image sequence F1 based on the training set HDR dynamic scene, including n training set HDR images;
[0079] Input the images with different exposure rates in the training set into the HDR dynamic neural radiance field model to obtain a training set HDR dynamic scene with uniform exposure; in the training set HDR dynamic scene, any ray is:
[0080] ;
[0081] where is the camera origin, is the ray direction, is the distance from a point on the ray to the origin along the ray direction;
[0082] For the sampling points on the ray and the ray direction in the training set HDR dynamic scene, calculate the spatial color value and the volume density of the ray based on the mapping function , and use the volume rendering method to obtain the integral of the ray in the training set HDR dynamic scene on the imaging plane and the integral of the opacity of the sampling points on the ray in the ray direction, which is used to represent the weight of the sampling points on the ray , as shown in the following formula:
[0083] ;
[0084] ;
[0085] where and respectively represent the near point and the far point of the sampling ray ;
[0086] Output a training set HDR image sequence F1 based on the training set HDR dynamic scene, including n training set HDR images;
[0087] The exposure rate of the HDR dynamic neural radiance field in a single learning scenario is learned, and the HDR dynamic scene of the training set with uniform exposure is reconstructed. In this embodiment, the HDR-HexPlane dynamic neural radiance field is used to reconstruct the HDR dynamic scene of the training set with uniform exposure, and based on the reconstructed HDR dynamic scene of the training set, a training set HDR image sequence F1 is output, including n training set HDR images;
[0088] S3: Construct a tone mapping module, input the training set HDR image sequence F1 into the tone mapping module, and obtain a training set LDR low dynamic range image sequence F2 with uniform exposure;
[0089] Since scene capture in the real world usually involves inputs with different exposure rates, the HDR neural radiance field can contain more information about the scene. Since the segmentation ability of two-dimensional vision large models can only accept the input of LDR image sequences, this embodiment proposes a tone mapping module controlled by hyperparameters to map the HDR scene view to LDR, and the scene exposure and contrast mapping are controlled by hyperparameters to solve the problem of local over-dark or over-bright that may be learned in the HDR scene.
[0090] The tone mapping module includes a normalization layer, an adjustment layer, and an output layer; the normalization layer obtains the training set HDR image sequence F1 and normalizes the n training set HDR images included in the training set HDR image sequence F1, and uses the sigmoid function to map the pixel value of each pixel point in the training set HDR image to between 0 and 1 to obtain the normalized training set HDR image; the adjustment layer adjusts the brightness and contrast of the normalized training set HDR image, where the adjustment of brightness and contrast is controlled by hyperparameters to reduce the model training parameters; the output layer maps the pixel value of each pixel point in the adjusted training set HDR image to the LDR color value range of 0-255 to obtain a training set LDR low dynamic range image sequence F2 with uniform exposure;
[0091] S31: Obtain the training set HDR image sequence F1 and normalize the n training set HDR images included therein, and use the sigmoid function to map the pixel value of each pixel point in the training set HDR image to between 0 and 1 to obtain the normalized training set HDR image, as shown in the following formula:
[0092] ;
[0093] where, is the pixel value of the pixel point in the training set HDR image;
[0094] S32: Adjust the brightness and contrast of the normalized training set of HDR images to obtain the adjusted training set of HDR images;
[0095] Since certain regions may be overexposed or underexposed when the reconstructed training set of HDR dynamic scenes is compressed to the LDR low dynamic range, therefore, use the tone mapping module adjustment layer of brightness and contrast controlled by hyperparameters to adjust the brightness of the normalized training set of HDR images in S31, as shown in the following formula:
[0096] ;
[0097] Among them, is the brightness hyperparameter, is the i th normalized training set of HDR images before brightness adjustment, is the i th training set of HDR images after brightness adjustment;
[0098] Adjust the contrast of the training set of HDR images after brightness adjustment, as shown in the following formula:
[0099] ;
[0100] Among them, is the contrast hyperparameter, is the i th training set of HDR images after contrast adjustment, is the function to enhance contrast, as shown in the following formula:
[0101] ;
[0102] ;
[0103] ;
[0104] Among them, is the image brightness function defined based on the sensitivity of the human eye to the three primary colors, is the function to enhance brightness, is the training set of HDR images after brightness adjustment at the pixel value of the pixel point on the red channel, is the training set of HDR images after brightness adjustment at the pixel value of the pixel point on the green channel, is the training set of HDR images after brightness adjustment at the pixel value of the pixel point on the blue channel;
[0105] S33: Map the pixel values of each pixel in the adjusted training set of HDR images to the LDR color value range of 0 - 255 to obtain a sequence of training set LDR low - dynamic - range images F2 with uniform exposure;
[0106] S4: Use a two - dimensional vision segmentation assistance model to segment the sequence of training set LDR low - dynamic - range images F2 with uniform exposure to obtain a two - dimensional segmentation mask of the target object in the training set HDR dynamic scene;
[0107] In dynamic interactive segmentation, the cue points that belong to one object at timestamp t0 may belong to other objects at timestamp t1. In this embodiment, a two - dimensional vision model assistance module and a mask deformation field network are proposed to implement dynamic mask inverse rendering to solve the problem of geometric inconsistency in the three - dimensional segmentation process.
[0108] S41: Select a two - dimensional vision large model as the two - dimensional vision segmentation assistance model and input the sequence of training set LDR low - dynamic - range images F2 into the two - dimensional vision large model;
[0109] Since the two - dimensional vision large model SAM2 can segment any object in the image sequence, in this embodiment, the SAM2 model is used as the two - dimensional vision segmentation assistance model. Based on the segmentation ability of the two - dimensional vision large model SAM2 for video or image sequences, auxiliary segmentation is realized. The sequence of training set LDR low - dynamic - range images F2 in.jpg format is input into the two - dimensional vision large model SAM2, and the model weight is sam2_hiera_large.pt;
[0110] S42: Set a number of segmentation cue points and background points on the target object in the training set HDR dynamic scene, and use the two - dimensional vision segmentation assistance model to obtain a two - dimensional segmentation mask of the target object in the training set HDR dynamic scene;
[0111] In this embodiment, when selecting a number of cue points and background points on the target object in the HDR dynamic scene, based on the two - dimensional vision large model SAM2, the number of segmentation cue points is set to be no less than 4. The more background points, the more accurate the segmentation mask can be obtained;
[0112] For the RGB three - channel image with size H×W×3 in the sequence of training set LDR low - dynamic - range images F2, where H is the height of the training set LDR low - dynamic - range image and W is the width of the training set LDR low - dynamic - range image, obtain its two - dimensional segmentation mask is H×W×1, the two - dimensional segmentation mask Store a single-channel Boolean value for each pixel in H×W×1. The masked area is True, and the Boolean value of the pixel in this area is 1. The unmasked area is False, and the Boolean value of the pixel in this area is 0. The segmentation mask of each image in the training set LDR low dynamic range image sequence F2 is stored in a numpy file in .npy format.
[0113] S43: If the two-dimensional visual segmentation auxiliary model cannot achieve the desired segmentation effect, adjust the brightness hyperparameter in S32 and the contrast hyperparameter To optimize the segmentation effect, S41 to S42 are repeated to obtain a two-dimensional segmentation mask sequence with the best segmentation effect of the target object in the HDR dynamic scene;
[0114] S5: construct a mask deformation field network, input the best 2D segmentation mask sequence obtained by the 2D segmentation model auxiliary model into the mask deformation field network for mask inverse rendering, project the 2D segmentation mask into the four-dimensional space, obtain the four-dimensional segmentation mask of the target object in the HDR dynamic scene of the training set, and achieve the segmentation of the target object in the HDR dynamic scene of the training set;
[0115] The mask deformation field network is as follows Figure 2 As shown in the figure, it includes a spatiotemporal interpolation encoder and a mask deformation decoder. The spatiotemporal interpolation encoder network includes a HexPlane submodule and an MLP submodule. The mask deformation decoder includes a lightweight multi-layer perceptron network. The input of the mask deformation field network is the ray in the HDR dynamic scene of the training set. Upsampling point Location coordinates , ray direction And timestamp information, output as rays in the training set HDR dynamic scene Upsampling point The mask confidence score of , which means ray Upsampling point The confidence score of the target object; the mask inverse rendering process is as follows Figure 3 shown.
[0116] S51: Rays in the training set HDR dynamic scene The position coordinates, ray direction and timestamp information of each sampling point are input into the HexPlane submodule, and the ray in the HDR dynamic scene of the training set is obtained based on the HexPlane plane interpolation method. The characteristics of each sampling point , through the MLP submodule, the rays in the HDR dynamic scene of the training set Fuse the features of all the sampling points above to obtain the fused features ;
[0117] Decompose the 4D voxel space into six learnable parameter planes based on the HexPlane plane interpolation method, as shown in the following formula:
[0118] ;
[0119] where , , , represent the plane bilinear interpolation sequence numbers of the sampling point position coordinates and the timestamp, , , , , , are all learnable parameter planes, , , are respectively the vectors orthogonal to the two planes of the bilinear interpolation, , , are respectively the dimensions of the planes for the bilinear interpolation, represents element-wise multiplication;
[0120] Use the six parameter planes to represent the four-dimensional scene to reduce the memory and computational overhead. Input the position coordinates of the sampling points on the ray , the ray direction and the timestamp information in the training set HDR dynamic scene into the HexPlane sub-module to obtain the feature output by the HexPlane sub-module, as shown in the following formula:
[0121] ;
[0122] where is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point on the plane based on the coordinate , is the bilinear interpolation of the sampling point On the plane bilinear interpolation based on coordinates is performed for the sampling points on the plane bilinear interpolation based on coordinates is performed for the sampling points on the plane bilinear interpolation based on coordinates is performed for the sampling points on the plane based on coordinates ; ";" is for the concatenation operation on the tensor
[0123] The MLP sub-module uses a multi-layer perceptron network to fuse the features of each sampling point on the ray in the training set HDR dynamic scene to obtain the fused features as shown in the following formula
[0124] ;
[0125] S52: The mask deformation decoder uses a lightweight multi-layer perceptron network to predict the confidence score of each sampling point on the ray in the training set HDR dynamic scene, and obtains the two-dimensional predicted mask of the target object in the training set HDR dynamic scene through mask inverse rendering
[0126] Only using the position coordinates of the sampling points on the ray in the training set HDR dynamic scene and the ray direction is not sufficient to fully represent the geometric information in the HDR dynamic scene. Therefore, based on the mapping function in the HDR dynamic neural radiance field model used in S2 the fused features obtained based on the spatio-temporal interpolation encoder and the ray direction are mapped to a periodic function
[0127] ;
[0128] where is the hyperparameter that controls the highest frequency of the sampling point ;
[0129] A lightweight multi-layer perceptron network is used to predict the confidence score of the sampling points on the ray in the training set HDR dynamic scene
[0130] ;
[0131] Among them, is the set of confidence scores of all sampling points on the ray in the training set HDR dynamic scene , and is the number of sampling points;
[0132] S53: Use the pre-trained HDR dynamic neural radiance field to query the volume density of each sampling point on the ray in the training set HDR dynamic scene, so as to obtain the opacity of the target object in the training set HDR dynamic scene, and obtain the weight value of the sampling points on the ray in the pre-trained HDR dynamic neural radiance field. Combine the confidence score of the sampling point on the ray in the training set HDR dynamic scene predicted by the mask deformation field network, and obtain the two-dimensional prediction mask of the target object in the training set HDR dynamic scene through mask inverse rendering, as shown in the following formula:
[0133] ;
[0134] Use the two-dimensional segmentation mask of the target object in the training set HDR dynamic scene as the ground truth and the two-dimensional prediction mask to calculate the loss, and implement dynamic mask inverse rendering. The loss function is as shown in the following formula:
[0135] ;
[0136] ;
[0137] ;
[0138] Among them, is the mask projection loss, is the total variance loss, is the ray set of picture I, , , are all hyperparameters, is a hyperparameter, is the point with coordinates in the plane , is the point with coordinates c in the plane , is the point with coordinates cThe point with coordinates in is the point with coordinates c in the plane ;
[0139] Update the parameters of the mask deformation field network by the gradient descent method, where the model parameters of the neural radiance field are frozen during the segmentation process.
[0140] S6: Use the training set to train the mask deformation field network. After reaching the preset number of training rounds, obtain the four-dimensional mask of the segmented HDR dynamic scene of the training set. Through querying, obtain the segmentation result of the target object in the HDR dynamic scene of the training set. Use the test set to evaluate the quality of scene reconstruction and the quality of segmentation. Set the part greater than 0 in the queried two-dimensional mask to True and the part less than 0 to False to obtain the segmentation result of the target object in the HDR dynamic scene of the test set. Based on the segmentation result of the test set, obtain the trained mask deformation field network to achieve semantic segmentation of any object in the HDR dynamic neural radiance field;
[0141] After 5000 rounds of iterative training, obtain the segmented four-dimensional mask. The projection under each query view is a single-channel confidence score map. Set the part with a score greater than 0 to True and the rest to False. Thus, the mask of the target object can be obtained under any view.
[0142] In this embodiment, the HDR-HexPlane dataset is used to verify the semantic segmentation method based on the HDR dynamic neural radiance field, with the average IoU reaching 86.63% and the average precision reaching 99.49%.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.
Claims
1. A semantic segmentation method based on HDR dynamic neural radiance fields, characterized in that: It includes the following steps: S1: Construct a scene dataset, including scene picture files and scene label files, and divide the scene dataset into a training set and a test set according to a ratio; S2: Use the HDR dynamic neural radiance field model to reconstruct the images with different exposure rates in the training set into a training set HDR dynamic scene with uniform exposure, and output a training set HDR image sequence F1 based on the training set HDR dynamic scene, including n training set HDR images; S3: Construct a tone mapping module, input the training set HDR image sequence F1 into the tone mapping module, and obtain a training set LDR low dynamic range image sequence F2 with uniform exposure; S4: Use a two-dimensional vision segmentation auxiliary model to segment the training set LDR low dynamic range image sequence F2 with uniform exposure, and obtain a two-dimensional segmentation mask of the target object in the training set HDR dynamic scene; S5: Construct a mask deformation field network, input the two-dimensional segmentation mask sequence with the best segmentation effect of the target object in the HDR dynamic scene obtained by the two-dimensional vision segmentation auxiliary model into the mask deformation field network for mask inverse rendering, project the two-dimensional segmentation mask into the four-dimensional space, and obtain a four-dimensional segmentation mask of the target object in the training set HDR dynamic scene, realizing the segmentation of the target object in the training set HDR dynamic scene; S6: Use the training set to train the mask deformation field network. After reaching the preset number of training rounds, obtain the four-dimensional mask of the HDR dynamic scene of the segmented training set. Through querying, obtain the segmentation result of the target object in the HDR dynamic scene of the training set. Use the test set to evaluate the quality of scene reconstruction and the quality of segmentation. Set the part greater than 0 in the queried two-dimensional mask to True, and the part less than 0 to False, obtain the segmentation result of the target object in the HDR dynamic scene of the test set, and based on the segmentation result of the test set, obtain the trained mask deformation field network to realize the semantic segmentation of any object in the HDR neural radiance field.
2. The semantic segmentation method based on HDR dynamic neural radiance fields according to claim 1, wherein: The scene label file described in S1 includes the file path, rotation matrix, translation vector, and timestamp information corresponding to each scene picture file. The training set is used for scene reconstruction based on the HDR neural radiance field, and the test set is used to evaluate the quality of scene reconstruction and the quality of segmentation.
3. The semantic segmentation method based on HDR dynamic neural radiance fields according to claim 2, characterized in that: The specific method of S2 is as follows: Input images with different exposure rates in the training set into the HDR dynamic neural radiance field model to obtain a training set HDR dynamic scene with uniform exposure; in the training set HDR dynamic scene, any ray is as follows: ; Among them, is the origin of the camera, is the ray direction, is the distance from a point on the ray along the ray direction to the origin; For the sampling points on the rays in the training set HDR dynamic scene and the ray directions , based on the mapping function , calculate the spatial color value of the ray and the volume density . Using the volume rendering method, obtain the integral of the ray in the training set HDR dynamic scene on the imaging plane and the integral of the opacity of the sampling points on the ray in the ray direction , which is used to represent the weight of the sampling points on the ray , as shown in the following formula: ; ; Among them, and respectively represent the near point and the far point of the sampling ray ; Based on the training set HDR dynamic scene, output the training set HDR image sequence F1, including n training set HDR images.
4. A semantic segmentation method based on HDR dynamic neural radiance fields according to claim 3, characterized in that: The tone mapping module described in S3 includes a normalization layer, an adjustment layer, and an output layer; the normalization layer obtains the training set HDR image sequence F1 and normalizes the n training set HDR images included in the training set HDR image sequence F1, uses the sigmoid function to map the pixel values of each pixel point in the training set HDR images to between 0 and 1, and obtains the normalized training set HDR images; the adjustment layer adjusts the brightness and contrast of the normalized training set HDR images, where the adjustment of the brightness and contrast is controlled by hyperparameters to reduce the model training parameters; the output layer maps the pixel values of each pixel point in the adjusted training set HDR images to the LDR color value range of 0-255, and obtains the training set LDR low dynamic range image sequence F2 with uniform exposure.
5. A semantic segmentation method based on HDR dynamic neural radiance fields according to claim 4, characterized in that: S3 includes: S31: Obtain the training set HDR image sequence F1, and normalize the n training set HDR images contained therein. Use the sigmoid function to map the pixel value of each pixel point in the training set HDR images to between 0 and 1, and obtain the normalized training set HDR images, as shown in the following formula: ; Among them, is the pixel value of the pixel point in the training set HDR image; S32: Adjust the brightness and contrast of the normalized training set HDR images to obtain the adjusted training set HDR images; Use the tone mapping module adjustment layer of brightness and contrast controlled by hyperparameters to adjust the brightness of the normalized training set HDR images in S31, as shown in the following formula: ; Among them, is the brightness hyperparameter, is the i th training set HDR image before brightness normalization, is the i th training set HDR image after brightness adjustment; The training set HDR images after brightness adjustment Perform contrast adjustment as shown in the following formula: ; Among them, is the contrast hyperparameter, is the i th training set HDR image after contrast adjustment, is the contrast enhancement function, as shown in the following formula: ; ; ; Among them, is the picture brightness function defined based on the human eye's sensitivity to the three primary colors, is the enhanced brightness function, is the training set HDR picture after brightness adjustment is the pixel value of the pixel point on the red channel, is the training set HDR picture after brightness adjustment is the pixel value of the pixel point on the green channel, is the training set HDR picture after brightness adjustment is the pixel value of the pixel point on the blue channel; S33: Map the pixel values of each pixel point in the adjusted training set HDR images to the LDR color value range of 0 - 255 to obtain a training set LDR low dynamic range image sequence F2 with uniform exposure.
6. The semantic segmentation method based on HDR dynamic neural radiance fields according to claim 5, characterized in that: S4 includes: S41: Select a two-dimensional vision large model as the two-dimensional vision segmentation auxiliary model, and input the training set LDR low dynamic range image sequence F2 into the two-dimensional vision large model; S42: Set a number of segmentation hint points and background points on the target object in the training set HDR dynamic scene, and use the two-dimensional vision segmentation auxiliary model to obtain a two-dimensional segmentation mask of the target object in the training set HDR dynamic scene; S43: If the two-dimensional visual segmentation assistance model fails to achieve the desired segmentation effect, adjust the brightness hyperparameter in S32 and the contrast hyperparameter to optimize the segmentation effect, and repeat S41 to S42 to obtain the two-dimensional segmentation mask sequence with the best segmentation effect of the target object in the HDR dynamic scene.
7. A semantic segmentation method based on HDR dynamic neural radiance fields according to claim 6, characterized in that: The specific method of S42 is as follows: For the RGB three-channel image with size H×W×3 in the LDR (Low Dynamic Range) image sequence F2 of the training set, where H is the height of the LDR images in the training set and W is the width of the LDR images in the training set, obtain its two-dimensional segmentation mask which is H×W×1, and the two-dimensional segmentation mask stores a single-channel boolean value for each pixel in H×W×1. The masked area is True and the boolean value of the pixel in this area is 1, while the unmasked area is False and the boolean value of the pixel in this area is 0.
8. A semantic segmentation method based on HDR dynamic neural radiance fields according to claim 7, characterized in that: The mask deformation field network described in S5 includes a spatio-temporal interpolation encoder and a mask deformation decoder. The spatio-temporal interpolation encoder network includes a HexPlane sub-module and an MLP sub-module. The mask deformation decoder includes a lightweight multi-layer perceptron network. The input of the mask deformation field network is the position coordinates, ray direction, and timestamp information of the upsampling points on the rays in the training set HDR dynamic scene, and the output is the mask confidence score of the upsampling points on the rays in the training set HDR dynamic scene, representing the confidence score that the upsampling points on the rays belong to the target object. upsampling points of the ray direction and timestamp information, and the output is the mask confidence score of the upsampling points on the rays in the training set HDR dynamic scene, upsampling points representing the confidence score that the upsampling points on the rays belong to the target object. upsampling points 9. A semantic segmentation method based on HDR dynamic neural radiance fields according to claim 8, characterized in that: S5 includes: S51: Input the position coordinates, ray directions, and timestamp information of each sampling point on the ray in the training set HDR dynamic scene into the HexPlane sub-module. Based on the HexPlane plane interpolation method, obtain the features of each sampling point on the ray in the training set HDR dynamic scene. Fuse the features of all sampling points on the ray in the training set HDR dynamic scene through the MLP sub-module to obtain the fused features ; S52: The mask deformation decoder uses a lightweight multi-layer perceptron network to predict the confidence scores of each sampling point on the rays in the training set HDR dynamic scene, and obtains the two-dimensional predicted mask of the target object in the training set HDR dynamic scene through mask inverse rendering; For each sampling point on the ray, the mask deformation decoder uses a lightweight multi-layer perceptron network to predict the confidence score, and obtains the two-dimensional predicted mask of the target object in the training set HDR dynamic scene through mask inverse rendering; Mapping function in the HDR dynamic neural radiance field model based on S2 usage , the fused features obtained based on the spatio-temporal interpolation encoder and the ray direction are mapped to a periodic function , as shown in the following formula: ; Among them, is the hyperparameter for controlling the sampling point with the highest frequency; Use a lightweight multi-layer perceptron network to predict the confidence scores of rays at upsampling points in the HDR dynamic scene of the training set , as shown in the following formula: ; Among them, is the set of confidence scores of all sampling points on the ray in the training set HDR dynamic scene , and is the number of sampling points; S53: Query the volume density of each sampling point on the ray in the training set HDR dynamic scene using the pre-trained HDR dynamic neural radiance field to obtain the opacity of the target object in the training set HDR dynamic scene, and get the weight values of the sampling points on the ray in the pre-trained HDR dynamic neural radiance field , combine with the confidence scores of the sampling points on the ray in the training set HDR dynamic scene predicted by the mask deformation field network to obtain the two-dimensional predicted mask of the target object in the training set HDR dynamic scene through mask inverse rendering , as shown in the following formula: ; The two-dimensional segmentation mask of the target object in the training set HDR dynamic scene is used as the ground truth and compared with the two-dimensional prediction mask to calculate the loss and achieve dynamic mask inverse rendering. The loss function is as shown in the following formula: ; ; ; Among them, is the mask projection loss, is the total variance loss, is the ray set of image I, , , are all hyperparameters, is a hyperparameter, is the plane with coordinates in it, is the plane c with coordinates in it, is the plane c with coordinates in it, is the plane c with coordinates in it; Update the mask deformation field network parameters by the gradient descent method, where the model parameters of the neural radiance field are frozen during the segmentation process.
10. A semantic segmentation method based on HDR dynamic neural radiance fields according to claim 9, characterized in that: The specific method of S51 is as follows: Based on the HexPlane plane interpolation method, the 4D voxel space is decomposed into six learnable parameter planes as shown in the following formula: ; Among them, , , , represent the plane bilinear interpolation sequence numbers of the sampling point position coordinates and time stamps, , , , , , are all learnable parameter planes, , , are respectively the vectors orthogonal to the two planes of the bilinear interpolation, , , are respectively the dimensions of the plane for performing bilinear interpolation, represents element-wise multiplication; Using six parameter planes to represent a four-dimensional scene to reduce memory and computational overhead, the position coordinates of the upsampling points on the ray in the training set HDR dynamic scene, the ray direction and the timestamp information are input into the HexPlane sub-module to obtain the features output by the HexPlane sub-module, as shown in the following formula: as follows: ; Among them, is a sampling point In bilinear interpolation based on coordinates on the plane ; is a sampling point In the plane bilinear interpolation based on coordinates ; is a sampling point In bilinear interpolation based on coordinates on the plane ; is a sampling point In the plane bilinear interpolation based on coordinates ; is a sampling point In the plane bilinear interpolation based on coordinates ; is a sampling point In bilinear interpolation based on coordinates on the plane ; ";" is an operation for splicing tensors; The MLP sub-module uses a multi-layer perceptron network to fuse the features of each sampling point on the rays in the training set HDR dynamic scene to obtain the fused features as shown in the following formula: 。
Citation Information
Patent Citations
Image high dynamic range reconstruction method based on deep learning
CN111292264A
High-dynamic three-dimensional measurement method based on overexposure connected domain intensity adaptive distribution
CN116310101A