Neural radiance field based underwater scene representation method
By designing a neural radiation field model based on physical laws and adopting hybrid progressive sampling and hybrid volume rendering methods, the problem of dynamic changes in underwater scenes is solved, and high-quality underwater scene reconstruction is achieved, especially on underwater data sets with and without temporal information.
Patent Information
- Application Number
- CN202410812616.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-22
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-06-22
AI Technical Summary
Existing neural radiance field methods cannot effectively handle dynamic changes caused by water absorption, scattering, illumination changes, and object motion in underwater scene reconstruction, resulting in poor reconstruction results.
A neural radiation field model based on physical laws was designed. By simulating underwater lighting changes and object motion, hybrid progressive sampling and hybrid volume rendering methods were used to learn the parameters of static objects and dynamic objects respectively, and a tone mapper was used to process lighting conditions to achieve accurate characterization of underwater scenes.
It achieves high-quality reconstruction of underwater scenes, can handle lighting changes and object motion, improves reconstruction accuracy and speed, and expands the application range of the algorithm, especially performing well on underwater datasets with and without temporal information.
Smart Images

Figure CN118710807B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network structure and neural radiation field design, and in particular to an underwater scene characterization method based on neural radiation field. Background Art
[0002] Neural Radiance Field (NeRF) is a method for 3D reconstruction and new perspective generation proposed in recent years. For a scene, given some 2D images taken from different perspectives, NeRF can learn a 3D representation of the entire scene and generate high-quality 2D images from a new perspective. Unlike the various 3D reconstruction methods that display representations proposed previously, NeRF implicitly stores the learned 3D information in an MLP network. This MLP network takes the 3D coordinates of a point in space and the 2D camera ray direction as input information, and outputs the voxel density and color of the point. When generating a new perspective, NeRF first calculates the camera ray corresponding to each pixel, samples some 3D points on each ray, inputs the information of these points into the MLP to obtain the voxel density and color of each point, and then calculates the color of the corresponding pixel through volume rendering.
[0003] NeRF performs well in most 3D reconstruction tasks for above-water scenes, but its performance for underwater scenes is less than satisfactory. This is mainly because NeRF is a 3D reconstruction method for static scenes, while underwater scenes often contain many dynamic components that NeRF cannot handle, such as:
[0004] First, because water absorbs light, the visibility and color of objects in the water change with the distance between them and the camera, which leads to the first dynamic factor of underwater scenes.
[0005] Secondly, the temporal changes in lighting conditions and the light scattering properties of water make the lighting conditions of the observed scene vary significantly over time, thus generating a second dynamic factor;
[0006] Finally, the movement of the organisms that are prevalent in the water creates a third dynamic element in the scene.
[0007] The above three dynamic factors combine to form a complex dynamic underwater scene, which makes it impossible for existing technologies to accurately represent or reconstruct underwater scenes; therefore, it is crucial to find an accurate representation method for underwater scenes based on NeRF.
[0008] In order to overcome the above-mentioned shortcomings, two solutions are designed in the prior art to optimize the underwater solution, including:
[0009] Prior Art 1: SeaThru-NeRF: Neural Radiance Fields in Scattering Media; see Figure 1 As shown;
[0010] Main method: Underwater images are represented as a synthesis of a clear image absorbed by the water body and backscattered images. Independent parameter representations are designed for the water body and backscattered images. Medium MLP is added to learn and estimate these parameters, realizing 3D reconstruction of static underwater scenes.
[0011] The specific implementation process includes:
[0012] (1) Calculate the camera ray and sample a series of points with weights on each ray.
[0013] (2) Using two MLPs to predict the voxel density σ of objects in the underwater scene obj With color
[0014] c obj , using Medium MLP to predict the voxel density σ of water bodies med With color
[0015] c med , and the voxel density σ of the backscattered part bs .
[0016] (3) Using the volume rendering method, the voxel density and color parameters of each point predicted by the network are used to calculate the pixel color corresponding to the ray. The rendering formula is:
[0017]
[0018] However, existing technology 1 still has technical deficiencies. It can only solve the dynamic changes in the surface color of objects caused by the absorption of light by water in underwater scenes, and cannot handle the dynamic changes caused by changes in illumination and moving objects. Since the color of the water body is related to the viewing angle in the model of this method, the model will mistakenly interpret the changes in illumination under different viewing angles as color changes under different viewing angles, resulting in a decrease in the quality of the generated new perspective images.
[0019] Prior Art 2: Dynamic View Synthesis from Dynamic Monocular Video; see Figure 2 As shown;
[0020] The main method uses a pre-trained model to distinguish between static and dynamic objects in the input data. Two NeRF networks with similar structures are used to learn 3D representations of the static and dynamic parts of the scene respectively. The learned coefficients are combined in the volume rendering process to obtain a 3D representation of the entire scene.
[0021] Implementation process:
[0022] (1) Calculate the camera ray and sample a series of points with weights on each ray.
[0023] (2) For the pixels in the static part, input them into a time-independent static NeRF network to learn the voxel density σ of the static part in the scene s With color c s .
[0024] (3) For the pixels in the dynamic part, input them into a dynamic
[0025] NeRF network learns the voxel density of the dynamic part of the scene at the corresponding time t
[0026] Degree σ d With color c d .
[0027] (4) Using the volume rendering method, the voxel density and color parameters of each point predicted by the network are used to calculate the pixel color corresponding to the ray.
[0028] Existing technology 2 also has technical defects. It can only solve the dynamic changes caused by moving objects in underwater scenes, and cannot handle the dynamic changes caused by the absorption and scattering properties of water to light and changes in lighting conditions. Moreover, since this method requires the use of a pre-trained model to distinguish between the dynamic object part and the static object part in the input image, the dynamic changes caused by other factors in the underwater scene will affect the classification effect of the pre-trained model, and thus affect the reconstruction effect of this method. Summary of the Invention
[0029] In order to solve the above technical problems, the present invention provides an underwater scene characterization method based on neural radiation fields. The invention is based on an underwater imaging model adapted to the physical laws and NeRF scene representation method, which effectively simulates the light absorption properties of water in underwater scenes. By designing an MLP network that predicts scene illumination information and a tone mapper, it simulates the changes in illumination in underwater scenes caused by the scattering properties of water and unstable lighting conditions, avoiding the model from incorrectly modeling such illumination changes as changes in the surface color of objects with changing viewing angles. The present invention also applies two different MLPs to model static objects and dynamic objects in the scene respectively, and adopts a step-by-step training method of first learning the information of static objects in the scene and then learning the information of dynamic objects in the scene, thereby achieving more accurate characterization and reconstruction of underwater scenes.
[0030] The underwater scene representation method based on neural radiation field is realized by designing corresponding modules in the neural network, including:
[0031] Step 1: Hybrid progressive sampling module:
[0032] By sampling on camera rays, both the object and water parts of the scene can be optimized simultaneously.
[0033] For the camera ray r, first select N equidistant m points to predict the parameters of the water body, and then use the multi-scale hash grid sampling N s points to achieve sufficient sampling near the surface of the object;
[0034] This N s The points need to satisfy: σ sta (o(r)+t i ·δ k ·d(r))>τ.
[0035] That is to say, on the camera ray with the starting point o(r) and the direction d(r), k To sample some points at intervals, the voxel density σ in the static scene parameters at these sampling points is satisfied sta Greater than a preset limit τ, t i is a series of positive integer values, a total of N s And satisfy the increasing relationship;
[0036] Dynamic objects contribute to the new perspective image only when they appear between the camera and static objects. Therefore, when modeling dynamic objects, an additional equidistant sampling is performed, starting from the camera's proximity to the depth where static objects appear in the scene. This provides higher accuracy for modeling dynamic objects.
[0037] Step 2: Scene parameter estimation module:
[0038] The sampling points obtained by the hybrid progressive sampling are input into each MLP after bicubic interpolation to predict parameter information of the scene at the sampling point;
[0039] The entire scene is divided into three parts: dynamic objects, static objects and water bodies. This results in the neural network outputting three sets of scene parameter prediction results for each sampling point: static scene parameters (σ sta,i , c sta,i ), dynamic scene parameters (σ dyn,i , c dyn,i ) and water parameters (σ w , c w ), where σ w It is a three-channel variable used to represent the different absorption coefficients of water to the three different colors of RGB light;
[0040] The neural network will simultaneously predict a local illumination parameter λ for each sampling point, which represents the exposure time of the sampling point under the input time and viewing direction;
[0041] Multiply the local illumination parameter λ by the three sets of color results {c sta,i , c dyn,i , c w}, we get the color c′ of the sampling point after considering the lighting conditions {sta,dyn,w},i ,Right now:
[0042] c′ {sta,dyn,w},i =λ·{c sta,i , c dyn,i , c w}.
[0043] As an example, using a tone mapper The colors of the sampling points in the linear space learned by scene parameter estimation are mapped into colors in the nonlinear space, simulating the processing process when the input image is taken, achieving a more realistic reconstruction effect.
[0044] Step 3: Hybrid Volume Rendering Module
[0045] The neural network predicts the three sets of scene parameters. In a real underwater scene, at each sampling point, only one of the three types of objects exists: dynamic objects, static objects, or water. It is difficult to completely separate the three in space, so a hybrid volume rendering model is used to mix the three parameters predicted by the network for the same point.
[0046] Most objects in underwater scenes are opaque, so when considering the blending between objects and water:
[0047] In the part where the object exists: the voxel density of the object is much greater than the voxel density of water;
[0048] Part in the water: The predicted voxel density of the object is close to 0, which is much smaller than the voxel density of water;
[0049] Therefore, there will not be much interference between the object and the water;
[0050] According to the previous assumption, dynamic objects must appear before static objects, so there will be no interference between objects, ensuring the rationality of the hybrid volume rendering model;
[0051] As an example, in order to make the hybrid volume rendering module converge faster, this patent designs a sine blending function to replace the normal linear blending method; for the three sets of voxel density values {σ sta,i ,σ dyn,i ,σ w}, mapped by the sine mixing function β:
[0052]
[0053] Mixed to get the voxel density σ at the sampling point i With color c i :
[0054] σ i =σ sta,i +σ dyn,i +σ w ,
[0055] c i =β sta,i c′ sta,i +β dyn,i c′ dyn,i +β w,i c′ w,i ·
[0056] Then, the pixel color is obtained according to the classic discrete volume rendering equation
[0057]
[0058] Among them: ⊙ represents channel multiplication, T i Indicates the cumulative penetration rate of light at the sampling point, δ i , δ j Indicates the distance between adjacent sampling points.
[0059] As an example, when the underwater scene representation method based on neural radiation field is applied to underwater data without time series information, the neural network is trained using the data without time series information. The training process is as follows:
[0060] ① Training data;
[0061] a) We used the dataset proposed in the existing work Seathru NeRF, which is divided into four scenes. We used the colmap method to perform sparse reconstruction on the images of each scene and obtain the corresponding camera pose for each image. We also separated the training set into the test set in a 4:1 ratio.
[0062] b) Before training, take out all camera rays corresponding to the input and mark the areas that no camera rays pass through in the multi-scale hash grid to accelerate the progressive sampling process.
[0063] Procedure;
[0064] ② Training of neural networks;
[0065] a) Each time a random input image is taken out, some pixels on the input image are randomly selected to obtain a batch;
[0066] b) Map these pixels into corresponding camera rays, place them into a multi-scale hash grid, and use hybrid progressive sampling to obtain a series of sampling points;
[0067] c) The neural network part includes shallow fully connected neural networks MLP-S and MLP-I; when processing input without temporal information, all time-related modules are removed from the complete neural network, including the time-related input of the entire MLP-D and MLP-I; the modified method will output a scene consisting only of static objects and media, and
[0068] The lighting conditions in the same direction at each sampling point are constant;
[0069] d) The color parameters output by the network are passed through the tone mapper, the voxel density parameters output by the network are passed through the sine blending function, the mapped parameters are passed through the hybrid volume rendering, and the corresponding pixel color is obtained through the discrete volume rendering formula; the loss function includes a pixel-level L2 loss ,in
[0070] C(r) represents the true color of the pixel, and a weighted cross entropy function
[0071] This function can be used to constrain the separation of static scenes and water bodies, where w i Represents the weight of each sampling point
[0072] Weight, the calculation formula is:
[0073] w i =T i ·(1-exp(-δ i ·σ i )),
[0074] r v Represents the inverse of visibility, k0 is a preset threshold, n0, n1 constrain the upper and lower bounds, and ε is a very small positive number to ensure that w i = 0, the weighted cross entropy function is still well defined.
[0075] As an example, when the underwater scene representation method based on neural radiation field is applied to underwater data with time series information, the neural network is trained using the time series data. The training process is as follows:
[0076] ① Training data;
[0077] 1. A new underwater scene reconstruction dataset with time-series information is proposed. Clips are captured from underwater videos on the internet, and frames are captured at regular time intervals to obtain the dataset input images. The dataset proposed in this patent is divided into four scenes. The colmap method is used to perform sparse reconstruction on the images of each scene to obtain the camera pose corresponding to each image. A pre-trained model based on optical flow is used to calculate the motion mask for each input image to distinguish between moving and static objects in the image. The training set and test set are separated in a 4:1 ratio.
[0078] 2. Before training, extract all camera rays corresponding to the input and mark the areas where no camera rays pass through them in the multi-scale hash grid to accelerate the progressive sampling process;
[0079] ② Training of neural networks;
[0080] 1. Each time, a random input image is taken out, and some static pixels on the input image are randomly selected according to the motion mask to obtain a batch;
[0081] 2. Map these pixels into corresponding camera rays, place them into a multi-scale hash grid, and use hybrid progressive sampling to obtain a series of sampling points;
[0082] 3. The neural network consists of shallow, fully connected neural networks (MLP-S, MLP-I, and MLP-D). For scenes with temporal information, the MLP-S and water parameters are trained using a method similar to that used for scenes without temporal information to accurately represent static objects in the scene. Furthermore, based on the fixed static object and water parameters, dynamic objects and time-dependent lighting information are further trained.
[0083] 4. The color parameters output by the network are passed through the tone mapper, the voxel density parameters output by the network are passed through the sine blending function, the mapped parameters are passed through the hybrid volume rendering, and the corresponding pixel color is obtained using the discrete volume rendering formula;
[0084] The loss function is a comprehensive L2 loss that mixes static pixels and dynamic pixels: And a weighted cross entropy function where λ sta The coefficient representing the static pixel part, λ dyn Represents the coefficient of the dynamic pixel part, and Represent the calculated static pixel and dynamic pixel values respectively.
[0085] Beneficial effects of the present invention:
[0086] (a) For underwater datasets lacking temporal information, this patent enables high-quality static reconstruction of underwater scenes. This method provides a 3D reconstruction of the scene under average lighting conditions, while removing dynamic objects. Compared to existing methods, this patent improves 3D reconstruction performance on such datasets.
[0087] (b) For underwater datasets with temporal information, this patent can achieve high-quality dynamic reconstruction of underwater scenes, restore the movement trajectory of underwater organisms and real-time changes in illumination, and expand the application scope of the algorithm.
[0088] (c) This patent significantly surpasses other existing methods for 3D reconstruction of underwater scenes in terms of neural network training speed and rendering speed of new perspective images.
[0089] (d) Compared with other existing methods, this patent improves the performance in water body elimination tasks and achieves more realistic results in water body migration tasks.
[0090] (e) A three-dimensional representation and reconstruction method for underwater scenes based on neural radiation fields, including the design of the network structure, the proposed reconstruction process steps, and the network modules specially designed according to the steps.
[0091] (f) Method for training neural networks and testing them by rendering images.
[0092] (g) All images and original video data in the underwater scene 3D reconstruction dataset with temporal information established by this patent, as well as the method of generating a dataset from the original video. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1This is the network structure of the prior art 1 in the background technology of the underwater scene characterization method based on neural radiation field of the present invention.
[0094] Figure 2 This is the network structure of the prior art 2 in the background technology of the underwater scene characterization method based on neural radiation field of the present invention.
[0095] Figure 3 This is the overall architectural design structure of the underwater scene characterization method based on neural radiation field of the present invention.
[0096] Figure 4 This is a system diagram of Example 1 of the underwater scene representation method based on neural radiation field of the present invention. (In the figure: Igt is a color underwater image, is the color result output by the neural network, I mask For I gt Predicted motion mask, black parts represent static pixels, white parts represent dynamic pixels.) DETAILED DESCRIPTION
[0097] Below, reference Figure 3 As shown in FIG, the underwater scene representation method based on neural radiation field is implemented through correspondingly designed modules in the neural network, including:
[0098] Step 1: Hybrid progressive sampling module:
[0099] By sampling on camera rays, both the object and water parts of the scene can be optimized simultaneously.
[0100] For the camera ray r, first select N equidistant m points to predict the parameters of the water body, and then use the multi-scale hash grid sampling N s points to achieve sufficient sampling near the surface of the object;
[0101] This N s The points need to satisfy: σ sta (o(r)+ti·δ k ·d(r))>τ.
[0102] That is to say, on the camera ray with the starting point o(r) and the direction d(r), k To sample some points at intervals, the voxel density σ in the static scene parameters at these sampling points is satisfied sta Greater than a preset limit τ, t i is a series of positive integer values, a total of N s And satisfy the increasing relationship;
[0103] Dynamic objects contribute to the new perspective image only when they appear between the camera and static objects. Therefore, when modeling dynamic objects, an additional equidistant sampling is performed, starting from the camera's proximity to the depth where static objects appear in the scene. This provides higher accuracy for modeling dynamic objects.
[0104] Step 2: Scene parameter estimation module:
[0105] The sampling points obtained by the hybrid progressive sampling are input into each MLP after bicubic interpolation to predict parameter information of the scene at the sampling point;
[0106] The entire scene is divided into three parts: dynamic objects, static objects and water bodies. This results in the neural network outputting three sets of scene parameter prediction results for each sampling point: static scene parameters (σ sta,i , c sta,i ), dynamic scene parameters (σ dyn,i , c dyn,i ) and water parameters (σ w , c w ), where σ w It is a three-channel variable used to represent the different absorption coefficients of water to the three different colors of RGB light;
[0107] The neural network will simultaneously predict a local illumination parameter λ for each sampling point, which represents the exposure time of the sampling point under the input time and viewing direction;
[0108] Multiply the local illumination parameter λ by the three sets of color results {c sta,i , c dyn,i , c w}, we get the color c′ of the sampling point after considering the lighting conditions {sta,dyn,w},i ,Right now:
[0109] c′ {sta,dyn,w},i =λ·{c sta,i , c dyn,i , c w}.
[0110] As an example, using a tone mapper The colors of the sampling points in the linear space learned by scene parameter estimation are mapped into colors in the nonlinear space, simulating the processing process when the input image is taken, achieving a more realistic reconstruction effect.
[0111] Step 3: Hybrid Volume Rendering Module
[0112] The neural network predicts the three sets of scene parameters. In a real underwater scene, at each sampling point, only one of the three types of objects exists: dynamic objects, static objects, or water. It is difficult to completely separate the three in space, so a hybrid volume rendering model is used to mix the three parameters predicted by the network for the same point.
[0113] Most objects in underwater scenes are opaque, so when considering the blending between objects and water:
[0114] In the part where the object exists: the voxel density of the object is much greater than the voxel density of water;
[0115] Part in the water: The predicted voxel density of the object is close to 0, which is much smaller than the voxel density of water;
[0116] Therefore, there will not be much interference between the object and the water;
[0117] According to the previous assumption, dynamic objects must appear before static objects, so there will be no interference between objects, ensuring the rationality of the hybrid volume rendering model;
[0118] As an example, in order to make the hybrid volume rendering module converge faster, this patent designs a sine blending function to replace the normal linear blending method; for the three sets of voxel density values {σ sta,i , σ dyn,i , σ w}, mapped by the sine mixing function β:
[0119]
[0120] Mixed to get the voxel density σ at the sampling point i With color c i :
[0121] σ i =σ sta,i +σ dyn,i +σ w ,
[0122] c i =β sta,i c′ sta,i +β dyn,i c′ dyn,i +β w,i c′ w,i ·
[0123] Then, the pixel color is obtained according to the classic discrete volume rendering equation
[0124]
[0125] Among them: ⊙ represents channel multiplication, T i Indicates the cumulative penetration rate of light at the sampling point, δ i , δ j Indicates the distance between adjacent sampling points.
[0126] As an example, when the underwater scene representation method based on neural radiation field is applied to underwater data without time series information, the neural network is trained using the data without time series information. The training process is as follows:
[0127] ① Training data;
[0128] c) We used the dataset proposed in the existing work Seathru NeRF, which is divided into four scenes. We used the colmap method to perform sparse reconstruction on the images of each scene and obtain the corresponding camera pose for each image. We also separated the training set into the test set in a 4:1 ratio.
[0129] d) Before training, extract all camera rays corresponding to the input and mark the areas that no camera rays pass through in the multi-scale hash grid to accelerate the progressive sampling process;
[0130] ② Training of neural networks;
[0131] e) Each time a random input image is taken out, some pixels on the input image are randomly selected to obtain a batch;
[0132] f) Map these pixels into corresponding camera rays, place them into a multi-scale hash grid, and use hybrid progressive sampling to obtain a series of sampling points;
[0133] g) The neural network consists of shallow fully connected neural networks (MLP-S and MLP-I). When processing inputs without temporal information, all time-dependent modules are removed from the complete neural network, including the entire time-dependent inputs of MLP-D and MLP-I. The modified method outputs a scene consisting only of static objects and media, with constant lighting conditions in the same direction at each sampling point.
[0134] h) The color parameters output by the network are passed through the tone mapper, the voxel density parameters output by the network are passed through the sine blending function, the mapped parameters are passed through the hybrid volume rendering, and the corresponding pixel color is obtained through the discrete volume rendering formula; the loss function includes a pixel-level L2 loss in
[0135] C(r) represents the true color of the pixel, and a weighted cross entropy function
[0136] This function can be used to constrain the separation of static scenes and water bodies, where w i represents the weight of each sampling point, and the calculation formula is:
[0137] w i = T i ·(1-exp(-δ i σ i )),
[0138] r v represents the inverse of the visibility, k0 is a pre-set threshold, n0 and n1 constrain the upper and lower bounds, and ε is a very small positive number to ensure that the weight cross-entropy function is well defined when w i = 0.
[0139] i)
[0140] As an example, when the underwater scene representation method based on neural radiance field is applied to underwater data with time sequence information, the data with time sequence information is used to train the neural network, and the training process is as follows:
[0141] ①Training data;
[0142] 1. A new underwater scene reconstruction data set with time sequence information is proposed, which is obtained by taking a segment from an underwater video on the Internet and taking frames at a certain time interval to obtain input images of the data set. The data set proposed in the patent is divided into four scenes, and the colmap method is used to perform sparse reconstruction on the images of each scene to obtain the camera pose corresponding to each image. A pre-trained model based on optical flow is used to calculate the motion mask for each input image to distinguish moving objects and static objects in the image. According to the 4:1 ratio, the training set and the test set are divided;
[0143] 2. Before training, all the corresponding camera rays are taken out and marked in the multi-scale hash grid without any camera rays passing through the area to speed up the progressive sampling process;
[0144] ②Training of neural network;
[0145] 1. Randomly take out an input image each time, and randomly select some static pixels on the input image according to the motion mask to obtain a batch;
[0146] 2. Map these pixels to the corresponding camera rays and put them into the multi-scale hash grid to obtain a series of sampling points using hybrid progressive sampling;
[0147] 3. The neural network consists of shallow, fully connected neural networks (MLP-S, MLP-I, and MLP-D). For scenes with temporal information, the MLP-S and water parameters are trained using a method similar to that used for scenes without temporal information to accurately represent static objects in the scene. Furthermore, based on the fixed static object and water parameters, dynamic objects and time-dependent lighting information are further trained.
[0148] 4. The color parameters output by the network are passed through the tone mapper, the voxel density parameters output by the network are passed through the sine blending function, the mapped parameters are passed through the hybrid volume rendering, and the corresponding pixel color is obtained using the discrete volume rendering formula;
[0149] The loss function is a comprehensive L2 loss that mixes static pixels and dynamic pixels: And a weighted cross entropy function where λ sta The coefficient representing the static pixel part, λ dyn Represents the coefficient of the dynamic pixel part, and Represent the calculated static pixel and dynamic pixel values respectively.
[0150] In order to better illustrate the design principle of the present invention, the specific process of training a neural network using time series data is briefly introduced through specific embodiment 1:
[0151] Example 1: See Figure 4 As shown;
[0152] a) Dataset construction and data preprocessing: We selected underwater videos from the Internet, which were about 10 to 20 seconds long, and captured 100 frames of underwater images at equal time intervals in the video. We used the colmap method to calculate the camera pose of each input image. We used optical flow to map each image. gt Estimate a motion mask I mask , used to calculate the L2 loss; the input image is divided into a training set and a test set in a 4:1 ratio. The camera rays corresponding to all pixels in the training set are calculated and placed in a multi-level hash grid to mark the space contained in the scene;
[0153] b) Training the neural network: Each time, a random image is selected from the training set. Based on the motion mask of the image, a group of pixels (usually 2048 or 4096) representing static objects are selected. The camera ray equation corresponding to these pixels is calculated. A series of points are mixed and progressively sampled on each ray. After bicubic interpolation, the result is input into the neural network. The neural network outputs the voxel density and color of each point, and then the network output image I and a series of weights w are obtained through hybrid volume rendering. i, and calculate the loss function for back propagation.
[0154] As an example, in the hybrid progressive sampling module, the multi-scale hash grid can be replaced by weighted sampling or layered sampling.
[0155] Alternatively, bicubic interpolation may be replaced with a positional encoding method. As an example, the number of layers or the number of neurons in each layer of the neural network may be changed, or the sinusoidal mixture function may be replaced with a linear function.
[0156] As an example, the abbreviations used in the present invention are summarized as follows:
[0157] NeRF:Neural Radiance Fields;
[0158] MLP:Multi Layer Perceptron;
[0159] Medium MLP: Medium Multi Layer Perceptron predicts medium parameters;
[0160] Colmap: A method for sparse reconstruction of 3D scenes.
[0161] Motion mask: Motion segmentation mask used to distinguish static objects from dynamic objects in an image;
[0162] L2 loss: MSE loss, mean square error;
[0163] The above are only preferred embodiments of the present invention. It should be understood that the description of the above embodiments is only used to help understand the method and core ideas of the present invention, and is not used to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, etc. made within the ideas and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. The underwater scene representation method based on neural radiation field is characterized by: This is achieved by designing corresponding modules in the neural network, including: Step 1: Hybrid progressive sampling module: By sampling on camera rays, both the object and water parts of the scene can be optimized simultaneously. For the camera ray r, first select N equidistant m points to predict the parameters of the water body, and then use the multi-scale hash grid sampling N s points to achieve sufficient sampling near the surface of the object; This N s The points need to satisfy: σ sta (o(r)+t i ·δ k ·d(r))>τ That is to say, on the camera ray with the starting point o(r) and the direction d(r), k Sampling some points at intervals, so that the voxel density of static objects at these sampling points is greater than a preset limit; Dynamic objects contribute to the new perspective image only when they appear between the camera and static objects. Therefore, when modeling dynamic objects, an additional equidistant sampling is performed, starting from the camera's proximity to the depth where static objects appear in the scene. This provides higher accuracy for modeling dynamic objects. Step 2: Scenario parameter estimation: The sampling points obtained by the hybrid progressive sampling are input into each MLP after bicubic interpolation to predict parameter information of the scene at the sampling point; The entire scene is divided into three parts: dynamic objects, static objects and water bodies. This results in the neural network outputting three sets of scene parameter prediction results for each sampling point: static scene parameters (σ sta , c sta ), dynamic scene parameters (σ dyn , c dyn ) and water parameters (σ w , c w ); The neural network will simultaneously predict a local illumination parameter λ for each sampling point, which represents the exposure time of the sampling point under the input time and viewing direction; Multiplying the local illumination parameter λ by the three sets of color results output by the neural network gives the color of the sampling point after considering the illumination conditions, namely: c′ {sta,dyn,w} =λ·c {sta,dyn,w} Step 3: Hybrid Volume Rendering The neural network predicts the three sets of scene parameters. In a real underwater scene, at each sampling point, only one of the three types of objects exists: dynamic objects, static objects, or water. It is difficult to completely separate the three in space, so a hybrid volume rendering model is used to mix the three parameters predicted by the network for the same point. Most objects in underwater scenes are opaque, so when considering the blending between objects and water: In the part where the object exists: the voxel density of the object is much greater than the voxel density of water; Part in the water: The voxel density of the predicted object is close to 0, which is much smaller than the voxel density of water.
2. The underwater scene characterization method based on neural radiation field according to claim 1 is characterized in that: Using a tone mapper The colors of the sampling points in the linear space learned by scene parameter estimation are mapped into colors in the nonlinear space, simulating the processing process when the input image is taken, achieving a more realistic reconstruction effect.
3. The underwater scene characterization method based on neural radiation field according to claim 1 is characterized in that: In order to make the hybrid volume rendering converge faster, a sinusoidal mixing function is designed to replace the normal linear mixing method; the three sets of voxel density values output by the model are mapped by the sinusoidal mixing function: Blend to get the voxel density and color at the sampling point: s i =s sta,i +s dyn,i +s w c i =b sta,i c′ sta,i +b dyn,i c′ dyn,i +b w,i c′ w,i Then, the pixel color is obtained according to the classic discrete volume rendering equation: Among them: ⊙ represents channel-wise multiplication.
4. The underwater scene characterization method based on neural radiation field according to claim 3 is characterized in that: When the underwater scene representation method based on neural radiation field is applied to underwater data without time series information, the neural network is trained using data without time series information. The training process is as follows: ① Training data; a) We used the dataset proposed in the existing work Seathru NeRF, which is divided into four scenes. We used the colmap method to perform sparse reconstruction on the images of each scene and obtain the corresponding camera pose for each image. We also separated the training set into the test set in a 4:1 ratio. b) Before training, all camera rays corresponding to the input are taken out and the areas that are not traversed by any camera rays are marked in the multi-scale hash grid to accelerate the progressive sampling process; ② Training of neural networks; a) Each time a random input image is taken out, some pixels on the input image are randomly selected to obtain a batch; b) Map these pixels into corresponding camera rays and put them into a multi-scale hash grid, A series of sampling points are obtained using mixed progressive sampling; c) The neural network part includes shallow fully connected neural networks MLP-S and MLP-I; when processing input without temporal information, all time-related modules are removed from the complete neural network, including the time-related input of the entire MLP-D and MLP-I; the modified method will output a scene consisting only of static objects and media, and The lighting conditions in the same direction at each sampling point are constant; d) The color parameters output by the network are passed through the tone mapper, the voxel density parameters output by the network are passed through the sine blending function, the mapped parameters are passed through the hybrid volume rendering, and the corresponding pixel color is obtained through the discrete volume rendering formula; the loss function includes a pixel-level L2 loss And a weighted cross entropy function This function can be used to constrain the separation of static scenes and water bodies.
5. The underwater scene characterization method based on neural radiation field according to claim 1 is characterized in that: When the underwater scene representation method based on neural radiation field is applied to underwater data with time series information, the neural network is trained using the time series information data. The training process is as follows: ① Training data; A new underwater scene reconstruction dataset with time-series information is proposed. Clips are taken from underwater videos on the internet, and frames are captured at regular intervals to obtain the dataset input images. The dataset is divided into four scenes. The colmap method is used to perform sparse reconstruction on the images of each scene, obtaining the corresponding camera pose for each image. A pre-trained model based on optical flow is used to calculate a motion mask for each input image to distinguish between moving and static objects in the image. The training and test sets are separated in a 4:1 ratio. Before training, all camera rays corresponding to the input are taken out and the areas that are not traversed by any camera rays are marked in the multi-scale hash grid to speed up the progressive sampling process. ② Training of neural networks; Each time, a random input image is taken out, and some static pixels on the input image are randomly selected according to the motion mask to obtain a batch; Map these pixels into corresponding camera rays, put them into a multi-scale hash grid, and use hybrid progressive sampling to obtain a series of sampling points; The neural network consists of shallow fully connected neural networks (MLP-S, MLP-I, and MLP-D). For scenes with temporal information, MLP-S and water parameters are trained to accurately represent static objects in the scene. Furthermore, based on the fixed static objects and water parameters, dynamic objects and time-related lighting information are further trained. The color parameters output by the network are passed through the tone mapper, the voxel density parameters output by the network are passed through the sine blending function, the mapped parameters are passed through the hybrid volume rendering, and the corresponding pixel color is obtained through the discrete volume rendering formula; The loss function is a comprehensive L2 loss that mixes static pixels and dynamic pixels: And a weighted cross entropy function 6. The underwater scene characterization method based on neural radiation field according to claim 1 is characterized in that: In the hybrid progressive sampling module, the multi-scale hash grid can be replaced by weighted sampling or layered sampling.
7. The underwater scene characterization method based on neural radiation field according to claim 1 is characterized in that: In the hybrid progressive sampling module, bicubic interpolation can be replaced by a position encoding method.
8. The underwater scene characterization method based on neural radiation field according to claim 1 is characterized in that: The neural network may adopt the MLP with the number of layers or the number of neurons in each layer changed, or the sinusoidal mixture function may be replaced with a linear function.
Citation Information
Patent Citations
Sampling image processing method for magnetic anomaly position by underwater magnetic survey robot
CN116137053A
Scene new view generation method based on local space aggregation neural radiation field
CN116993826A