4D scene representation method based on joint pose and radiation field optimization in complex mine environment
By adopting a 4D scene characterization method with joint position and radiation field optimization in complex mine environments, the shortcomings of traditional methods in dynamic and static partial decoupling are solved, and high-accuracy scene reconstruction is achieved, providing reliable information support for mine robot perception and navigation.
Patent Information
- Application Number
- CN202411022864.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-07-29
AI Technical Summary
In the complex environment of the mine, large-scale, non-level road surfaces and dynamic interference lead to the failure of traditional pose estimation methods, and the inability to effectively decouple dynamic and static parts, resulting in inaccurate pose estimation and scene characterization results.
The 4D scene characterization method with joint pose and radiation field optimization is adopted to reconstruct dynamic and static parts, and dynamic separation is achieved, and the pose is jointly optimized to ensure the accuracy of scene reconstruction. Specific steps include collecting video data and lidar scanning data, preprocessing data, building dynamic static sampling point selection fields, spatial static radiation field and spatiotemporal dynamic radiation field models, building static and dynamic losses, progressive joint optimization of pose and radiation field, and finally rendering scene representation.
It realizes effective separation and optimization of dynamic and static parts in complex mine environments, improves the accuracy of pose estimation and scene representation, and provides accurate information for robot perception, positioning and navigation.
Smart Images

Figure CN119006687B_ABST
Abstract
Description
Technical Field
[0001] The present invention is a 4D scene characterization method for joint pose and radiation field optimization in a complex mine environment, specifically a 4D scene characterization method for joint estimation of camera pose and dynamic and static radiation field optimization for a complex mine environment, belonging to the technical field of mine intelligent positioning. Background Art
[0002] In recent years, the implementation of my country's "new infrastructure" construction plan has greatly promoted the improvement of the intelligence level of the coal industry and ensured the construction of smart mines. It can be seen that intelligent coal mining has become one of the energy technologies that the country focuses on supporting.
[0003] Among them, as an important part of coal mine intelligent technology, accurately and effectively characterizing complex mine environments is difficult but very valuable. Under normal circumstances, the classical neural radiation field characterization method uses the Structure-from-Motion (SfM) method to preprocess data to obtain the pose estimation result as the pose truth value. SfM extracts the features contained in the image and matches similar features between image frames. According to the matched features, the pose is solved based on the graphics method.
[0004] However, unlike general small-scale object-level scenes, the large amount of dynamic interference on large-scale and uneven roads in the complex environment of mines poses challenges to the pose estimation, scene perception, and environmental modeling of mine autonomous robots, making the general methods that rely on traditional pose estimation methods and single dynamic field modeling invalid. It is impossible to completely decouple the dynamic and static parts, causing the dynamic and static contents to interfere with each other and become noise, which ultimately affects the results of pose estimation and scene representation. Summary of the invention
[0005] The purpose of the present invention is to provide a 4D scene representation method that combines posture and radiation field optimization in a complex mine environment. The method achieves dynamic and static separation by reconstructing the dynamic and static parts in the mine environment. For posture errors, while achieving scene representation, the posture is jointly optimized to ensure accurate and consistent scene reconstruction results, thereby providing accurate information for the robot's perception, positioning and navigation.
[0006] In order to achieve the above object, the present invention provides a 4D scene representation method for combined posture and radiation field optimization in a complex environment of a mine, comprising the following steps:
[0007] S1: Collect video data V and lidar scanning data L in the complex environment of the target mine;
[0008] S2: pre-processing the video data V and the laser radar scanning data L collected in step S1;
[0009] S3: Construct a dynamic and static sampling point selection field model;
[0010] S4: construct the spatiotemporal dynamic radiation field model and the spatial static radiation field model respectively;
[0011] S5: construct static loss and dynamic loss respectively;
[0012] S6: Progressive joint optimization of camera pose, dynamic and static sampling point selection field, spatial static radiation field and spatiotemporal dynamic radiation;
[0013] S7: Scene representation at the desired viewpoint of the rendered modeled scene.
[0014] In step S2 of the present invention, the collected video data V and laser radar scanning data L are preprocessed, and the specific steps are as follows:
[0015] S2.1: Use the Structure from Motion (SfM) method to read the collected video data V and estimate the pose P of each frame in the video, including the camera's three-dimensional spatial position coordinates X and the camera's shooting direction d. Specifically, first detect the stable feature points of each frame, match the feature points between frames to establish inter-frame associations, estimate the relative pose between frames through the matched inter-frame feature points, and further optimize the estimated pose by combining view geometry principles and reprojection optimization methods to obtain more accurate pose results.
[0016] S2.2: Use the large segmentation model to preprocess the video data V, and manually refine and fine-tune the dynamic mask Mask for the segmentation result;
[0017] S2.3: Use a pre-trained optical flow estimation model to estimate the inter-frame optical flow F representing 2D planar motion. Specifically, first use a convolutional network to encode the features of each frame, extract the corresponding position features of adjacent frames to calculate the visual similarity, and calculate the visual similarities of different scales to form a similarity pyramid in order to focus on both the global and local aspects. Use the same encoder to only extract the contextual features of the previous frame between adjacent frames and the features extracted from the similarity pyramid as the input of the decoding CNN to regress the inter-frame optical flow between adjacent frames.
[0018] S2.4: The laser radar sampling points representing the distance in the laser radar scanning data L are first transferred from the laser radar coordinate system to the data collection vehicle body coordinate system through coordinate conversion, and then the 3D point cloud data is projected onto the 2D image plane to obtain a scaled depth map D;
[0019] S2.5: Define a normalized scene local space S i , normalized in the range of [-2, 2]; construct the local space S iThe training sample set of video data V randomly samples m pixel rays for each frame, corresponding to m viewing angles {V i |i=1,2,...,m}, each pixel ray uniformly samples n spatial sampling points to obtain a set of m×n sampling point positions, and records the viewing angle of each sampling point to form a data sample set.
[0020] The dynamic and static sampling point selection field of the present invention is defined as a learnable four-dimensional feature field;
[0021] Decompose the four-dimensional feature field, decompose the four-dimensional feature field into a feature matrix combination on the XYZT basis, wherein each sampling point A=(x, y, z, t) with a timestamp can query the corresponding scene feature in the decomposed feature combination, and for the sampling point that is not in the grid point, bilinear interpolation is used to obtain the corresponding feature by interpolating from the surrounding grid point features. The dynamic and static sampling point selection field modeling formula is:
[0022] The dynamic and static sampling point selection field model constructed in step S3 for:
[0023]
[0024] Where: R m Represents the number of decomposed plane basis combinations, where each set of plane basis contains three components, and each component is two mutually perpendicular characteristic planes;
[0025] M represents the characteristic plane;
[0026] r represents the serial number of the current plane basis combination;
[0027] and They represent the rth plane decomposed on the two vertical axes X and Y and the rth plane decomposed on the two vertical axes Z and T, respectively. The two planes also maintain a vertical relationship in the feature space;
[0028] and They represent the rth plane decomposed on the two vertical axes X and Z and the rth plane decomposed on the two vertical axes Y and T, respectively. The two planes also maintain a vertical relationship in the feature space;
[0029] and They represent the rth plane decomposed on the two vertical axes Y and Z and the rth plane decomposed on the two vertical axes X and T, respectively. The two planes also maintain a vertical relationship in the feature space;
[0030] Each M is a feature plane to be optimized, and each feature plane is orthogonal to each other. This decomposition expression can show the modeling symmetry and strike a balance between representation ability and storage efficiency. Each sampling point A = (x, y, z, t) is mapped to an f-dimensional feature, and the formula is expressed as:
[0031]
[0032] Among them: ⊙ represents the element product;
[0033] and Respectively represent the sampling point A in all R m The combination of all features indexed at coordinates (x, y) on the XY plane and the sampling point A in all R m The combination of all features indexed at coordinate (z, t) on the ZT plane;
[0034] and Respectively represent the sampling point A in all R m The combination of all features indexed at coordinates (x, z) on the XZ plane and the sampling point A in all R m The combination of all features indexed at coordinate (y, t) on the YT plane;
[0035] and Respectively represent the sampling point A in all R m The combination of all features indexed at coordinates (y, z) on the YZ plane and the sampling point A in all R m The combination of all features indexed at coordinate (x, t) on the XT plane;
[0036] The sampled features are then activated by the Sigmoid function to obtain the dynamic and static selection probability θ of the sampling point, which is expressed as:
[0037] θ=Sigmoid(V m (x, y, z, t)) (3)
[0038] When the value of the dynamic or static selection probability θ of a sampling point is greater than 0.5, it is considered to be a dynamic point and the sampling point is input into the spatiotemporal dynamic radiation field. Otherwise, it is input into the spatial static radiation field. The dynamic and static sampling point selection field is supervised and optimized through the dynamic mask Mask of the scene, which can strongly separate the dynamic and static objects in the scene. It does not need to rely on neural scene flow or canonical field to constrain its dynamics, but directly selects the dynamic and static sampling points for dynamic and static selection, which can avoid mutual interference between dynamic and static objects.
[0039] In step S4 of the present invention, a spatial static radiation field model representing only spatial correlation and a spatiotemporal dynamic radiation field model related to both time and space are constructed respectively, specifically:
[0040] Constructing a spatial static radiation field, since the static part of the scene does not change over time, its characteristics are only related to the given position, that is, the spatial static radiation field can be decomposed on the XYZ basis, and its feature volume is decomposed into a combination of feature vectors and feature matrices. The spatial static radiation field is further represented as a static volume density field representing geometric information. and a static appearance field representing appearance information The formula is:
[0041]
[0042]
[0043] Where: R sσ The number of vector-plane bases representing the static volume density field decomposition;
[0044] R sc The number of vector-plane bases representing the static appearance field decomposition;
[0045] sr represents the sequence number of the vector-plane basis of the current decomposition;
[0046] represents the density feature vector on the X-axis in the current vector-plane basis, Represents the density feature plane on the YZ axis in the current vector-plane basis;
[0047] Represents the density feature vector on the Y axis in the current vector-plane basis, Represents the density feature plane on the XZ axis in the current vector-plane basis;
[0048] Represents the density feature vector on the Z axis in the current vector-plane basis, Represents the density feature plane on the XY axis in the current vector-plane basis;
[0049] Represents the color feature vector on the X-axis in the current vector-plane basis, Represents the color feature plane on the YZ axis in the current vector-plane basis;
[0050] Represents the color feature vector on the Y axis in the current vector-plane basis, Represents the color feature plane on the XZ axis in the current vector-plane basis;
[0051] Represents the color feature vector on the Z axis in the current vector-plane basis, Represents the color feature plane on the XY axis in the current vector-plane basis;
[0052] The volume density of each static sampling point can be calculated by using the density feature corresponding to the ReLU activation function inference, and the formula is expressed as:
[0053]
[0054] The color features of static sampling points use a static point color decoder MLP sc Get the color value of the corresponding sampling point. The formula is:
[0055]
[0056] The inferred color value will be used to generate the pixel value;
[0057] Constructing a spatiotemporal dynamic radiation field, which is different from a spatial static radiation field that is only related to the spatial position, the spatiotemporal dynamic radiation field is related to both time and space information. Therefore, its decomposition is consistent with the dynamic and static sampling point selection field. It is also composed of plane bases. Each set of plane bases contains three components, and each component is composed of two mutually perpendicular planes. At the same time, similar to the spatial static field, a spatiotemporal dynamic radiation field is divided into a dynamic density feature field according to density and color attributes. and dynamic appearance feature field The formula is:
[0058]
[0059]
[0060] Where: R dσ The number of plane bases representing the dynamic volume density field decomposition, R dc The number of plane bases representing the decomposition of the dynamic appearance field;
[0061] dr represents the serial number of the plane basis currently decomposed;
[0062] and They represent the density feature planes of the dynamic density field in the XY and ZT dimensions respectively, and the two planes also maintain a vertical relationship in the feature space;
[0063] and They represent the density feature planes of the dynamic density field in the XZ and YT dimensions respectively, and the two planes also maintain a vertical relationship in the feature space;
[0064] and They represent the density feature planes of the dynamic density field in the YZ and XT dimensions respectively, and the two planes also maintain a vertical relationship in the feature space;
[0065] Similar to the volume density decoding process of the sampling point in the spatial static radiation field, the corresponding features are also inferred using the ReLU activation function to obtain the corresponding sampling point density value, which is expressed as:
[0066]
[0067] The color features of the dynamic sampling points use a dynamic point color decoder MLP dc Get the color value of the corresponding sampling point. The formula is:
[0068]
[0069] The static loss and dynamic loss are constructed respectively in step S5 of the present invention as follows:
[0070] S5.1: Static loss construction, minimizing prediction Photometric loss between images captured in the static region and the image captured in the static region:
[0071]
[0072] Among them: Mask is a dynamic mask;
[0073] S5.2: Dynamic loss construction, the photometric training loss of the dynamic part is:
[0074]
[0075] Introduce auxiliary loss in static loss construction to regularize training:
[0076] S5.1.1: Reprojection loss The reprojection operation refers to finding the correspondence between the surface points of the same three-dimensional object in the adjacent frame image level according to the relationship between posture and depth. First, the estimated camera posture and estimated depth projection pixel are used to project the point back to the next frame image plane according to the camera posture of the next frame, and the corresponding pixel is found. The movement of the pixels in the adjacent frames, i.e., the optical flow, is calculated, and the reprojection loss is calculated with the optical flow value directly estimated by the RAFT model.
[0077] S5.1.2: Difference loss Similar to the above reprojection loss, the error is regularized in the Z-axis direction. The specific method is as follows: find the corresponding pixel points according to the 2D optical flow estimated by the RAFT model, use volume rendering to find the positions of their respective corresponding surface points, and calculate the depth difference loss for the Z-axis coordinates of the two;
[0078] S5.1.3: Monocular Depth Loss The above two losses cannot handle pure rotation of the camera, which often leads to incorrect camera pose and geometry. The depth map is rendered by volume rendering, and the depth loss is constructed between the sparse depth map obtained in the above data processing;
[0079] The final loss of the static part is:
[0080]
[0081] Auxiliary losses are introduced in the dynamic loss construction to regulate training. Auxiliary losses include reprojection losses. Difference loss Monocular depth loss
[0082] Three external prior-based losses are introduced to better simulate dynamic motion; these three losses are similar to the static part, but they need to simulate the motion of 3D points, so a scene flow MLP is introduced to compensate for 3D motion;
[0083]
[0084] Among them: SF i→i+1 Indicates time t i 3D scene flow of 3D point (x, y, z) at time instant;
[0085] The 3D motion prediction of the MLP is further regularized by introducing a smoothing and smaller scene flow loss:
[0086]
[0087] Separate the gradient of the dynamic radiation field from the camera pose, and finally use the dynamic mask Mask to supervise the non-rigid mask Mask d :
[0088]
[0089] The overall loss of the dynamic part is:
[0090]
[0091] S5.3: Linearly combine the static and dynamic partial losses into the final result and predict the non-rigid m d :
[0092]
[0093] The total training loss is:
[0094]
[0095] The total loss function is used to adjust the parameters of the neural radiance field network model.
[0096] The progressive joint optimization of camera pose, dynamic and static sampling point selection field, spatial static radiation field and spatiotemporal dynamic radiation proposed in step S6 of the present invention is specifically performed as follows:
[0097] S6.1: Backpropagation loss gradient to optimize local space S i Internal camera pose, dynamic and static sampling point selection field characteristics and corresponding radiation field characteristics;
[0098] S6.2: Determine whether the optimized current frame pose reaches the set local field boundary: if the condition is false, continue to optimize the current field; if the condition is true, allocate a new space S i+1 ;
[0099] S6.3: Local Space S i and new space S i+1 There is an overlapping area between them, and the local space S i The optimized pose of the first few frames before the boundary is used as the new space S i+1 The initial pose value of the previous few frames will be in the new space S i+1 During the space optimization period, the optimization continues. The new space S i+1 Spatial optimization includes dynamic and static sampling point selection fields, radiation field optimization and corresponding camera pose optimization. The radiation field includes spatial static radiation field and spatiotemporal dynamic radiation field.
[0100] The scene representation under the desired viewing angle of the rendered modeling scene in step S7 of the present invention is specifically performed as follows:
[0101] S7.1: Given a desired viewing direction d and a set of camera rays R under the viewing direction, randomly sample spatial points on the rays to obtain a set of sampling point coordinates X;
[0102] According to the sampling and reasoning process of the dynamic and static sampling point selection field introduced in S3, the dynamic possibility of each sampling point is obtained and its dynamic and static affiliation is determined;
[0103] Among them, static points will sample features from the static radiation field in space, and dynamic points will sample features from the dynamic radiation field in space and time;
[0104] Using the sampling characteristics, the density value and color value of the sampling point belonging to the spatial static radiation field are obtained using formula (6) and formula (7), and the density value and color value of the sampling point belonging to the spatiotemporal dynamic radiation field are obtained using formula (10) and formula (11). The above values will be used for scene rendering;
[0105] S7.2: Scene rendering, using the volume rendering formula, accumulates the density and color values of all sampling points on the camera ray to calculate the expected pixel value at the target viewing angle:
[0106]
[0107] in: Represents the expected pixel value obtained by reasoning;
[0108] δ i Represents the distance between two consecutive sampling points;
[0109] N represents the number of sampling points sampled on the entire sampling ray.
[0110] T i Represents the cumulative transmittance along the ray.
[0111] Compared with the prior art, the present invention collects video data and laser radar scanning data in the complex environment of the target mine, pre-processes the collected data to obtain the required information; then constructs dynamic and static sampling point selection fields, spatial static radiation fields, and spatiotemporal dynamic radiation field models respectively; inputs the sampling points in the training samples into the dynamic and static sampling point selection fields, distinguishes the dynamic and static states of the sampling points, and then inputs them into the corresponding network for training; and combines the camera posture and spatial static radiation field optimization. The dynamic and static scenes are rendered and input into the loss function, and the loss function is used to adjust the parameters of the neural radiation field network model. Finally, the scene representation under the desired perspective is obtained through volume rendering. While realizing the scene representation, the posture is jointly optimized to ensure accurate and consistent scene reconstruction results, providing accurate information for the robot's perception, positioning, and navigation. BRIEF DESCRIPTION OF THE DRAWINGS
[0112] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0113] Figure 2 A model diagram of a dynamic and static sampling point selection field in an embodiment of the present invention;
[0114] Figure 3 This is a network model diagram of a space static radiation field in an embodiment of the present invention;
[0115] Figure 4 This is a network model diagram of the spatiotemporal dynamic radiation field in an embodiment of the present invention;
[0116] Figure 5 Schematic diagram of the joint optimization principle of posture and radiation field in an embodiment of the present invention. DETAILED DESCRIPTION
[0117] The present invention will be further described below in conjunction with the accompanying drawings.
[0118] like Figure 1 As shown, a 4D scene representation method for joint pose and radiation field optimization in a complex mine environment includes the following steps:
[0119] S1: Collect video data V and lidar scanning data L in the complex environment of the target mine;
[0120] The present invention collects video data V and laser radar scanning data L of the actual mine environment by running the data acquisition platform. The video data V in this example contains 500 frames of 1080×1920 resolution RGB images, and the timestamp of the video records the time information of each frame t v The laser radar signal scans the scene and generates sparse point cloud data describing geometric information. In addition, in order to ensure the consistency of the data, the image signal and the laser radar signal need to be time synchronized. The specific method of the time synchronization operation of the present invention is as follows: since GPS allows nanosecond-level time synchronization progress, before collecting data, the GPS sends a signal pulse to activate the time signals of the camera and the laser radar to align the GPS time signal.
[0121] S2: pre-processing the video data V and the laser radar scanning data L collected in step S1;
[0122] The purpose of the preprocessing of the present invention is to generate a series of data that can be directly used subsequently from the collected original data.
[0123] This embodiment first uses the Structure from Motion (SfM) method to obtain the rough approximate pose to be optimized corresponding to each frame of the image, including the initial camera 3D coordinates and orientation. Specifically, the stable feature points of each frame of the image are first detected, and the inter-frame association is established by matching the feature points between frames. The relative pose between frames is estimated by matching the inter-frame feature points, and the estimated pose is further optimized by combining the view geometry principle and the reprojection optimization method to obtain a more accurate pose result.
[0124] In order to supervise the dynamic and static separation process, the original segmentation mask M′ of the dynamic object in the image is obtained by segmenting the large model (SAM-Track), and then the dynamic mask Mask is obtained by refining the insufficient segmentation area through annotation software;
[0125] Use the pre-trained optical flow estimation model to obtain the image optical flow estimation OF k→k+1, k∈[1,...,P-1], further constrain the dynamic representation, specifically: first use the convolutional network to encode the features of each frame of the image, and calculate the visual similarity of the feature points at the corresponding positions of the feature maps extracted from adjacent frames (taking the first frame and the second frame as an example). In order to focus on both the global and local aspects, the visual similarities of different scales are calculated to form a similarity pyramid. Use the same encoder to extract only the context features of the first frame and the features extracted from the similarity pyramid as the input of the decoding CNN to regress the optical flow between the first and second frames;
[0126] The laser radar sampling points representing the distance in the laser radar scanning data L are first transferred from the laser radar coordinate system to the data collection vehicle body coordinate system through coordinate transformation, and then the 3D point cloud data is projected onto the 2D image plane to obtain the scaled depth map D.
[0127] Finally, the entire dataset is divided into training set, test set and validation set in a ratio of 8:1:1.
[0128] Define a normalized scene local space S i : Since the complex environment of the mine is a large-scale boundless environment, in order to facilitate sampling and parameterization, the local space is shrunk by the mapping formula, and the shrinkage range is set to [-2, 2]. The corresponding mapping formula is:
[0129]
[0130] Randomly sample m sampling rays, randomly sample n sampling points on each sampling ray, and obtain a set of m×n sampling point positions, which are used as a batch of data for training. The sampling ray is the connecting ray between the camera shooting optical center corresponding to the randomly selected training frame and the randomly sampled pixel, and the points on the sampling ray are sampled using a uniform random sampling method. In this embodiment, 1024 sampling rays are randomly sampled, and 16 to 128 points are randomly selected on each sampling ray for sampling. The number of sampling points for each ray gradually increases with the increase in the number of training iterations. The purpose is to continuously refine the sampling process, improve efficiency, and retain more scene details.
[0131] S3: Construct a dynamic and static sampling point selection field model;
[0132] like Figure 2 As shown, the dynamic and static sampling point selection field is defined as a learnable four-dimensional feature field;
[0133] Decompose the four-dimensional feature field, and decompose the four-dimensional feature field into a feature matrix combination on the XYZT basis. Each sampling point A = (x, y, z, t) with a timestamp can query the corresponding scene feature in the decomposed feature combination. For sampling points that are not in the grid, bilinear interpolation is used to obtain the corresponding features by interpolating from the surrounding grid point features. The dynamic and static sampling point selection field modeling formula is:
[0134]
[0135] Each M is a set of feature planes to be optimized, and each pair of feature planes is orthogonal. This decomposition expression can show the modeling symmetry and strike a balance between representation ability and storage efficiency. Each sampling point A = (x, y, z, t) is mapped to an f-dimensional feature, expressed as:
[0136]
[0137] The sampled features are then activated by the Sigmoid function to obtain the dynamic and static selection probability θ of the sampling point, which is expressed as:
[0138] θ=Sigmoid(V m (x, y, z, t)) (3)
[0139] When the value of the dynamic or static selection probability θ of a sampling point is greater than 0.5, it is considered to be a dynamic point and the sampling point is input into the spatiotemporal dynamic radiation field. Otherwise, it is input into the spatial static radiation field. The dynamic and static sampling point selection field is supervised and optimized through the dynamic mask Mask of the scene, which can strongly separate the dynamic and static objects in the scene. It does not need to rely on neural scene flow or canonical field to constrain its dynamics, but directly selects the dynamic and static sampling points for dynamic and static selection, which can avoid mutual interference between dynamic and static objects.
[0140] S4: Construct spatial static radiation field model and spatiotemporal dynamic radiation field model respectively;
[0141] S4.1: Construct a static radiation field in space, such as Figure 3 As shown in Figure 1, since the static part of the scene does not change with time, its characteristics are only related to the given position, that is, the spatial static radiation field can be decomposed on the XYZ basis, and its feature volume is decomposed into a combination of feature vectors and feature matrices. The spatial static radiation field is further represented as a static volume density field representing geometric information. and a static appearance field representing appearance information Different from the above-mentioned dynamic and static sampling point selection field plane basis, each component of which is represented by two perpendicular characteristic planes, because there is no time dimension, each component of each set of vector-plane basis of the spatial static radiation field is a combination of a vector and a plane, and the formula is expressed as:
[0142]
[0143] The volume density of each static sampling point can be calculated by using the density feature corresponding to the ReLU activation function inference, and the formula is expressed as:
[0144]
[0145] The color features of static sampling points use a static point color decoder MLP sc Get the color value of the corresponding sampling point. The formula is:
[0146]
[0147] The inferred color value will be used to generate the pixel value;
[0148] S4.2: Construct a spatiotemporal dynamic radiation field, such as Figure 4 As shown in the figure, unlike the spatial static radiation field which is only related to the spatial position, the spatiotemporal dynamic radiation field is related to both time and space information. Therefore, its decomposition is consistent with the dynamic and static sampling point selection field. It is also composed of plane bases. Each set of plane bases contains three components, and each component is composed of two mutually perpendicular planes. At the same time, similar to the spatial static field, a spatiotemporal dynamic radiation field is divided into a dynamic density feature field according to the density and color attributes. and dynamic appearance feature field The formula is:
[0149]
[0150] Similar to the volume density decoding process of the sampling point in the spatial static radiation field, the corresponding features are also inferred using the ReLU activation function to obtain the corresponding sampling point density value, which is expressed as:
[0151]
[0152] The color features of the dynamic sampling points use a dynamic point color decoder MLP dc Get the color value of the corresponding sampling point. The formula is:
[0153]
[0154] S5: construct static loss and dynamic loss respectively;
[0155] S5.1: Static loss construction, minimizing prediction Photometric loss between images captured in the static region and the image captured in the static region:
[0156]
[0157] Among them: Mask is a dynamic mask;
[0158] S5.2: Dynamic loss construction, the photometric training loss of the dynamic part is:
[0159]
[0160] Introduce auxiliary loss in static loss construction to regularize training:
[0161] S5.1.1: Reprojection loss The reprojection operation refers to finding the correspondence between the surface points of the same three-dimensional object in the adjacent frame image level according to the relationship between posture and depth. First, the estimated camera posture and estimated depth projection pixel are used to project the point back to the next frame image plane according to the camera posture of the next frame, and the corresponding pixel is found. The movement of the pixels in the adjacent frames, i.e., the optical flow, is calculated, and the reprojection loss is calculated with the optical flow value directly estimated by the RAFT model.
[0162] S5.1.2: Difference loss Similar to the above reprojection loss, the error is regularized in the Z-axis direction. The specific method is as follows: find the corresponding pixel points according to the 2D optical flow estimated by the RAFT model, use volume rendering to find the positions of their respective corresponding surface points, and calculate the depth difference loss for the Z-axis coordinates of the two;
[0163] S5.1.3: Monocular Depth Loss The above two losses cannot handle pure rotation of the camera, which often leads to incorrect camera pose and geometry. The depth map is rendered by volume rendering, and the depth loss is constructed between the sparse depth map obtained in the above data processing;
[0164] The final loss of the static part is:
[0165]
[0166] Auxiliary losses are introduced in the dynamic loss construction to regulate training. Auxiliary losses include reprojection losses. Difference loss Monocular depth loss
[0167] Three external prior-based losses are introduced to better simulate dynamic motion; these three losses are similar to the static part, but they need to simulate the motion of 3D points, so a scene flow MLP is introduced to compensate for 3D motion;
[0168]
[0169] Among them: SF i→i+1 Indicates time t i3D scene flow of 3D point (x, y, z) at time instant;
[0170] The 3D motion prediction of the MLP is further regularized by introducing a smoothing and smaller scene flow loss:
[0171]
[0172] Separate the gradient of the dynamic radiation field from the camera pose, and finally use the dynamic mask Mask to supervise the non-rigid mask Mask d :
[0173]
[0174] The overall loss of the dynamic part is:
[0175]
[0176] S5.3: Linearly combine the static and dynamic partial losses into the final result and predict the non-rigid m d :
[0177]
[0178] The total training loss is:
[0179]
[0180] The total loss function is used to adjust the parameters of the neural radiance field network model.
[0181] S6: Progressive joint optimization of camera pose, dynamic and static sampling point selection field, spatial static radiation field and spatiotemporal dynamic radiation;
[0182] like Figure 5 As shown, the radiation field (a general term for fields, including the above-mentioned dynamic and static sampling point selection fields, spatial static radiation fields, and spatiotemporal dynamic radiation) is modeled within the respective allocated spatial ranges, such as S i Corresponding to F i As the position changes and the space expands, new radiation fields will be assigned to the corresponding space (expressed as an increase in serial number), and there will be partial overlapping areas between the radiation fields to maintain a smooth transition.
[0183] When continuously optimizing the scene, the camera pose and radiation field are progressively jointly optimized. At the same time, the optimized pose of the previous radiation field will be used as the initial pose of the next field in the overlapping area of the field, and the initial pose will be further optimized as the optimization process of the next field progresses. During the pose optimization process, in order to avoid the influence of dynamic interference on the pose, the pose optimization process is only synchronized with the optimization process of the spatial static radiation field to ensure the accuracy of the pose.
[0184] The progressive joint optimization of camera pose, dynamic and static sampling point selection field, spatial static radiation field and spatiotemporal dynamic radiation proposed in step S6 is as follows:
[0185] S6.1: Backpropagation loss gradient to optimize local space S i Internal camera pose, dynamic and static sampling point selection field characteristics and corresponding radiation field characteristics;
[0186] S6.2: Determine whether the optimized current frame pose reaches the set local field boundary: if the condition is false, continue to optimize the current field; if the condition is true, allocate a new space S i+1 ;
[0187] S6.3: Local Space S i and new space S i+1 There is an overlapping area between them, and the local space S i The optimized pose of the first few frames before the boundary is used as the new space S i+1 The initial pose value of the previous few frames will be in the new space S i+1 During the space optimization period, the optimization continues. The new space S i+1 Spatial optimization includes dynamic and static sampling point selection fields, radiation field optimization and corresponding camera pose optimization. The radiation field includes spatial static radiation field and spatiotemporal dynamic radiation field.
[0188] S7: Scene representation at the desired viewpoint of the rendered modeled scene
[0189] The specific steps are as follows:
[0190] S7.1: Given a desired viewing direction d and a set of camera rays R under the viewing direction, randomly sample spatial points on the rays to obtain a set of sampling point coordinates X;
[0191] According to the sampling and reasoning process of the dynamic and static sampling point selection field introduced in S3, the dynamic possibility of each sampling point is obtained and its dynamic and static affiliation is determined;
[0192] Among them, static points will sample features from the static radiation field in space, and dynamic points will sample features from the dynamic radiation field in space and time;
[0193] Using the sampling characteristics, the density value and color value of the sampling point belonging to the spatial static radiation field are obtained using formula (6) and formula (7), and the density value and color value of the sampling point belonging to the spatiotemporal dynamic radiation field are obtained using formula (10) and formula (11). The above values will be used for scene rendering;
[0194] S7.2: Scene rendering, using the volume rendering formula, accumulates the density and color values of all sampling points on the camera ray to calculate the expected pixel value at the target viewing angle:
[0195]
[0196]
[0197] in: Represents the expected pixel value obtained by reasoning;
[0198] δ i Represents the distance between two consecutive sampling points;
[0199] N represents the number of sampling points sampled on the entire sampling ray; T i Represents the cumulative transmittance along the ray.
Claims
1. A 4D scene representation method combining pose and radiation field optimization in a complex mine environment, characterized in that: The following steps are involved: S1: Collecting video data in complex environments of target mines , LiDAR scan data ; S2: Video data collected in step S1 and LiDAR scan data Perform pre-processing; S3: Construct a dynamic and static sampling point selection field model; S4: construct the spatiotemporal dynamic radiation field model and the spatial static radiation field model respectively; S5: construct static loss and dynamic loss respectively; S6: Progressive joint optimization of camera pose, dynamic and static sampling point selection field, spatial static radiation field and spatiotemporal dynamic radiation field; S7: Scene representation at the desired viewpoint of the rendered modeled scene; In step S2, the collected video data and LiDAR scan data Perform preprocessing. The specific steps are as follows: S2.1: Reading acquired video data using structure-from-motion methods , estimate the corresponding pose of each frame in the video , including the camera's three-dimensional space coordinates and camera shooting direction Specifically, firstly, the stable feature points of each frame are detected, and the inter-frame association is established by matching the feature points between frames. The relative pose between frames is estimated by matching the feature points between frames, and the estimated pose is further optimized by combining the view geometry principle and the reprojection optimization method to obtain a more accurate pose result. S2.2: Using a large segmentation model for video data Perform preprocessing, further manually refine the segmentation results and fine-tune the dynamic mask ; S2.3: Estimate inter-frame optical flow representing 2D planar motion using a pre-trained optical flow estimation model Specifically, we first use a convolutional network to encode the features of each frame, extract the corresponding position features from adjacent frames to calculate the visual similarity, and calculate the visual similarities of different scales to form a similarity pyramid in order to focus on both the global and local aspects. We use the same encoder to only extract the contextual features of the previous frame between adjacent frames and the features extracted from the similarity pyramid as the input of the decoding CNN to regress the optical flow between adjacent frames. S2.4: LiDAR scanning data The laser radar sampling points representing the distance in the image are first converted from the laser radar coordinate system to the data collection vehicle body coordinate system through coordinate transformation. Then, the 3D point cloud data is projected onto the 2D image plane to obtain a scaled depth map. ; S2.5: Define a normalized scene local space , normalized in Within the scope; construct local space Training sample set, video data Each frame of data is randomly sampled pixel rays, corresponding to Perspective, each pixel ray is sampled evenly spatial sampling points, we get A set of sampling point locations is created, and the viewing angle of each sampling point is recorded to form a data sample set; The dynamic and static sampling point selection field is defined as a learnable four-dimensional feature field; Decompose the four-dimensional feature field and convert the four-dimensional feature field into The basis is decomposed into a combination of feature matrices, where each sampling point with a timestamp The corresponding scene features can be queried in the decomposed feature combination. For sampling points that are not in the grid, bilinear interpolation is used to obtain the corresponding features by interpolating from the surrounding grid features. The dynamic and static sampling point selection field modeling formula is: The dynamic and static sampling point selection field model constructed in step S3 for: (1) in: Represents the number of decomposed plane basis combinations, where each set of plane basis contains three components, and each component is two mutually perpendicular characteristic planes; represents the characteristic plane; Indicates the sequence number of the current plane basis combination; and Respectively decompose in The two perpendicular axes plane and decomposition in The two perpendicular axes The two planes also maintain a vertical relationship in the feature space; and Respectively decompose in The two perpendicular axes plane and decomposition in The two perpendicular axes The two planes also maintain a vertical relationship in the feature space; and Respectively decompose in The two perpendicular axes plane and decomposition in The two perpendicular axes The two planes also maintain a vertical relationship in the feature space; Each sampling point All are mapped to an f-dimensional feature, and the formula is expressed as: (2) in: represents element-wise product; and Respectively represent sampling points In all indivual On the plane at coordinates The combination of all features and sampling points indexed at In all indivual On the plane at coordinates The combination of all features indexed at ; Respectively represent sampling points In all indivual On the plane at coordinates The combination of all features and sampling points indexed at In all indivual On the plane at coordinates The combination of all features indexed at ; Respectively represent sampling points In all indivual On the plane at coordinates The combination of all features and sampling points indexed at In all indivual On the plane at coordinates The combination of all features indexed at ; The sampled features are then activated by the Sigmoid function to obtain the dynamic and static selection probabilities of the sampling points. , the formula is: (3) When the sampling point dynamic and static selection probability When the value is greater than 0.5, it is considered to be a dynamic point and the sampling point is input into the spatiotemporal dynamic radiation field. Otherwise, it is input into the spatial static radiation field. The dynamic and static sampling point selection field is determined by the dynamic mask of the scene. Perform supervised optimization.
2. According to claim 1, a 4D scene representation method combining pose and radiation field optimization in a complex mine environment is characterized in that: In step S4, a spatial static radiation field model representing only spatial correlation and a spatiotemporal dynamic radiation field model related to both time and space are constructed respectively, specifically: Constructing a spatial static radiation field. Since the static part of the scene does not change over time, its characteristics are only related to the given position, that is, the spatial static radiation field can be The characteristic volume is decomposed into a combination of eigenvectors and characteristic matrices, and the spatial static radiation field is further represented as a static volume density field representing geometric information. and a static appearance field representing appearance information , the formula is: (4) (5) in: The number of vector-plane bases representing the static volume density field decomposition; The number of vector-plane bases representing the static appearance field decomposition; Indicates the sequence number of the vector-plane basis of the current decomposition; Represents the current vector-plane basis The density eigenvectors on the axes, Represents the current vector-plane basis density characteristic plane on the axis; Represents the current vector-plane basis The density eigenvectors on the axes, Represents the current vector-plane basis density characteristic plane on the axis; Represents the current vector-plane basis The density eigenvectors on the axes, Represents the current vector-plane basis density characteristic plane on the axis; Represents the current vector-plane basis The color feature vector on the axis, Represents the current vector-plane basis Color feature plane on axis; Represents the current vector-plane basis The color feature vector on the axis, Represents the current vector-plane basis Color feature plane on axis; Represents the current vector-plane basis The color feature vector on the axis, Represents the current vector-plane basis Color feature plane on axis; The volume density of each static sampling point is calculated using the density feature corresponding to the ReLU activation function inference, and the formula is expressed as: (6) The color features of static sampling points use a static point color decoder Get the color value of the corresponding sampling point. The formula is: (7) The inferred color value will be used to generate the pixel value; Constructing a spatiotemporal dynamic radiation field, which is different from a spatial static radiation field that is only related to the spatial position, the spatiotemporal dynamic radiation field is related to both time and space information. Therefore, its decomposition is consistent with the dynamic and static sampling point selection field. It is also composed of plane bases. Each set of plane bases contains three components, and each component is composed of two mutually perpendicular planes. At the same time, similar to the spatial static field, a spatiotemporal dynamic radiation field is divided into a dynamic density feature field according to density and color attributes. and dynamic appearance feature field , the formula is: (8) (9) in: The number of plane bases representing the dynamic volume density field decomposition, The number of plane bases representing the decomposition of the dynamic appearance field; Indicates the sequence number of the plane basis currently decomposed; and They represent the dynamic density field in and The density feature plane of the dimension, the two planes also maintain a vertical relationship in the feature space; and They represent the dynamic density field in and The density feature plane of the dimension, the two planes also maintain a vertical relationship in the feature space; and They represent the dynamic density field in and The density feature plane of the dimension, the two planes also maintain a vertical relationship in the feature space; Similar to the volume density decoding process of the sampling point in the spatial static radiation field, the corresponding features are also inferred using the ReLU activation function to obtain the corresponding sampling point density value, which is expressed as: (10) The color features of the dynamic sampling points use a dynamic point color decoder Get the color value of the corresponding sampling point. The formula is: (11)。 3. According to claim 2, a 4D scene representation method combining pose and radiation field optimization in a complex mine environment is characterized in that: The static loss and dynamic loss are constructed respectively in step S5 as follows: S5.1: Static loss construction, minimizing prediction Photometric loss between images captured in the static region and the image captured in the static region: (12) in: is a dynamic mask; S5.2: Dynamic loss construction, the photometric training loss of the dynamic part is: (13) Introduce auxiliary loss in static loss construction to regularize training: S5.1.1: Reprojection loss : The reprojection operation refers to finding the correspondence between the surface points of the same three-dimensional object in the adjacent frame image level according to the relationship between posture and depth. First, the pixel is projected back to the next frame image plane according to the camera posture of the next frame through the estimated camera posture and the estimated depth. The corresponding pixel is found and the movement of the pixels in the adjacent frames, that is, the optical flow, is calculated. The reprojection loss is calculated with the optical flow value directly estimated by the RAFT model; S5.1.2: Difference loss : Similar to the above reprojection loss, the error is regularized in the Z-axis direction. The specific method is as follows: find the corresponding pixel points according to the 2D optical flow estimated by the RAFT model, use volume rendering to find the positions of their respective corresponding surface points, and calculate the depth difference loss for the Z-axis coordinates of the two; S5.1.3: Monocular Depth Loss : Depth map rendered by volume rendering and scaled depth map obtained in the above data preprocessing Construct a deep loss between them; The final loss of the static part is: (14) Auxiliary losses are introduced in the dynamic loss construction to regulate training. Auxiliary losses include reprojection losses. , Difference loss , Monocular depth loss ; Three external prior-based losses are introduced to better simulate dynamic motion; these three losses are similar to the static part, but they need to simulate the motion of 3D points, so a scene flow MLP is introduced to compensate for 3D motion; (15) in: Indicates time Three-dimensional point at time 3D scene flow; The 3D motion prediction of MLP is further regularized by introducing smoothness and scene flow losses: (16) Separate the gradient of the dynamic radiation field from the camera pose and finally use a dynamic mask Supervised Non-Rigid Masking : (17) The overall loss of the dynamic part is: (18) S5.3: Linearly combine the static and dynamic part losses into the final result and predict a non-rigid mask : (19) in: is the distance between two consecutive sampling points; represents the cumulative transmittance along the sampling ray; The total training loss is: (20) The total loss function is used to adjust the parameters of the neural radiance field network model.
4. According to claim 2, a 4D scene representation method combining pose and radiation field optimization in a complex mine environment is characterized in that: The progressive joint optimization of camera pose, dynamic and static sampling point selection field, spatial static radiation field and spatiotemporal dynamic radiation proposed in step S6 is as follows: S6.1: Backpropagating loss gradients to optimize local space Internal camera pose, dynamic and static sampling point selection field characteristics and corresponding radiation field characteristics; S6.2: Determine whether the optimized current frame pose reaches the set local field boundary: if the condition is false, continue to optimize the current field; if the condition is true, allocate new space ; S6.3: Local Space and new space There are overlapping areas between them, local spaces The optimized pose of the first few frames before the boundary is used as the new space The initial pose values of the previous frames will be in the new space Continue to optimize during space optimization, new space Spatial optimization includes dynamic and static sampling point selection fields, radiation field optimization and corresponding camera pose optimization. The radiation field includes spatial static radiation field and spatiotemporal dynamic radiation field.
5. The 4D scene representation method of combined pose and radiation field optimization in a complex mine environment according to claim 2 is characterized in that: Step S7 is to render the scene representation under the desired viewing angle of the modeling scene, and the specific steps are as follows: S7.1: Given the desired viewing direction and a set of camera rays in the viewing direction , randomly sample spatial points on the ray to obtain a set of sampling point coordinates X; According to the sampling and reasoning process of the dynamic and static sampling point selection field in step S3, the dynamic possibility of each sampling point is obtained, and its dynamic and static attribution is determined; Among them, static points will sample features from the static radiation field in space, and dynamic points will sample features from the dynamic radiation field in space and time; Using the sampling characteristics, the density value and color value of the sampling point belonging to the spatial static radiation field are obtained using formula (6) and formula (7), and the density value and color value of the sampling point belonging to the spatiotemporal dynamic radiation field are obtained using formula (10) and formula (11). The above values will be used for scene rendering; S7.2: Scene rendering, using the volume rendering formula, accumulates the density and color values of all sampling points on the camera ray to calculate the expected pixel value at the target viewing angle: , (21) in: Represents the expected pixel value obtained by reasoning; Represents the distance between two consecutive sampling points; Indicates the number of sampling points sampled on the entire sampling ray; Represents the cumulative transmittance along the ray.