Surface rupture 3d profile reconstruction method and system based on deep learning
The deep learning-based method for reconstructing 3D topography of surface fractures utilizes the joint training of a neural implicit field network and a detail enhancement subnetwork to generate a high-frequency detail displacement field. This solves the problems of insufficient micro-topographic detail accuracy and poor surface continuity in the reconstruction of 3D topography of surface fractures, achieving high-fidelity reconstruction of 3D topography of surface fractures.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA EARTHQUAKE DISASTER PREVENTION CENT
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies for reconstructing the three-dimensional topography of fractured surfaces suffer from insufficient accuracy in reconstructing micro-topographic details and poor surface continuity, making it particularly difficult to achieve high-fidelity reconstruction in open environments.
A deep learning-based method for reconstructing the three-dimensional topography of surface fractures is adopted. Sparse three-dimensional point clouds and multi-scale depth feature maps are generated from multi-view optical remote sensing images. The neural implicit field network and detail enhancement sub-network are jointly trained to generate a high-frequency detail displacement field. The zero isosurface is extracted using the moving cube algorithm to generate a triangular mesh model.
It achieves rich detail and topologically correct 3D topography reconstruction of surface fractures in an open environment, improves the accuracy of micro-topography detail reconstruction and surface continuity, and generates a rich detail 3D topography mesh model of surface fractures.
Smart Images

Figure CN122435181A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional reconstruction technology, and in particular to a method and system for reconstructing three-dimensional morphology of surface fractures based on deep learning. Background Technology
[0002] In existing technologies, multi-view stereo vision-based methods are widely used. These methods recover camera pose from multi-view optical remote sensing images and generate sparse point clouds using structure-of-motion (SOG) technology. Dense point clouds are then generated using dense matching algorithms, and triangular mesh models are generated using algorithms such as Poisson reconstruction or Marching Cubes. These methods rely on the accuracy of feature point matching and the photometric consistency between images, and can achieve good results in texture-rich areas. However, since surface fractures are often located in complex terrain environments, there are challenges such as vegetation cover, shadows, and missing textures. Traditional MVS methods still have limitations in reconstructing fracture edges and micro-topographic details, especially in terms of insufficient reconstruction accuracy for fracture features at the centimeter to decimeter level.
[0003] Existing methods often separate segmentation, reconstruction, and post-processing steps, leading to the gradual accumulation of errors. Most methods are based on discrete representations, making it difficult to achieve high-fidelity reconstruction of continuous surfaces. The emergence of neural implicit field technology has provided a new approach to continuous surface representation, but when directly applied to open environment surface fracture scenarios, it still faces problems of blurred details and unclear boundaries, especially in the absence of high-precision prior data, where its ability to preserve micro-topographic features is limited. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a deep learning-based method for reconstructing the three-dimensional topography of surface fractures, which solves the problems of insufficient accuracy in reconstructing micro-topographic details of surface fractures and poor surface continuity.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for reconstructing the three-dimensional topography of surface fractures based on deep learning, which includes: acquiring multi-view optical remote sensing images, processing the multi-view optical remote sensing images through motion recovery structure technology, generating camera pose parameters and sparse three-dimensional point clouds, and performing preprocessing and multi-scale feature extraction to obtain a multi-scale depth feature map. Multi-scale depth feature maps and camera pose parameters are input into a neural implicit field network, and a neural implicit field model is generated through forward inference. Intermediate features are extracted from the neural implicit field network and input into the detail enhancement subnetwork to generate a high-frequency detail displacement field. The high-frequency detail displacement field is then used to enhance the geometric details of the neural implicit field model, resulting in the enhanced neural implicit field model. Based on the enhanced neural implicit field model, the neural implicit field network and the detail enhancement sub-network are jointly trained and their parameters are optimized to obtain a well-trained deep learning model. Using a trained deep learning model, the symbolic distance field of the target region is obtained. The moving cube algorithm is applied to the symbolic distance field to extract the zero isosurface, generating a triangular mesh model. The triangular mesh model is then post-processed to output a three-dimensional surface fracture topography mesh model.
[0007] As a preferred embodiment of the deep learning-based three-dimensional topography reconstruction method for surface fractures described in this invention, the method includes: acquiring multi-view optical remote sensing images, processing the multi-view optical remote sensing images using structure-reconstruction-motion (SRM) technology to generate camera pose parameters and sparse three-dimensional point clouds, and performing preprocessing and multi-scale feature extraction to obtain a multi-scale depth feature map, comprising the following steps: Acquire multi-view optical remote sensing images, extract feature points from the multi-view optical remote sensing images, obtain feature points of the multi-view optical remote sensing images, match the feature points of the multi-view optical remote sensing images, and obtain feature point matching pairs of the multi-view optical remote sensing images. By matching feature points of multi-view optical remote sensing images, the camera projection geometry is solved, the position and orientation of each multi-view optical remote sensing image in three-dimensional space are recovered, and the camera pose parameters of the multi-view optical remote sensing images are obtained. Sparse 3D point clouds are generated by using camera pose parameters of multi-view optical remote sensing images and feature point matching pairs of multi-view optical remote sensing images. The multi-view optical remote sensing images are preprocessed using camera pose parameters to obtain preprocessed multi-view optical remote sensing images. Multi-scale feature extraction is then performed on the preprocessed multi-view optical remote sensing images to obtain multi-scale depth feature maps.
[0008] As a preferred embodiment of the deep learning-based three-dimensional topography reconstruction method for surface fractures described in this invention, the method includes the following steps: inputting multi-scale depth feature maps and camera pose parameters into a neural implicit field network, and generating a neural implicit field model through forward inference. Multi-scale depth feature maps and camera pose parameters are input into a neural implicit field network. The neural implicit field network fuses the multi-scale depth feature maps and camera pose parameters through a learnable view-dependent feature modulation module to construct features of three-dimensional spatial points. The neural implicit field network uses a multilayer perceptron to decode the features of three-dimensional spatial points and regress the symbolic distance function value and color value of the three-dimensional spatial points. By using a differentiable volume rendering algorithm, a neural implicit field model is generated by accumulating symbolic distance function values and color values along camera rays.
[0009] As a preferred embodiment of the deep learning-based three-dimensional topography reconstruction method for surface fractures described in this invention, the method includes the following steps: extracting intermediate features from a neural implicit field network, inputting the intermediate features into a detail enhancement subnetwork, and generating a high-frequency detail displacement field: Intermediate features are extracted from the neural implicit field network to obtain the intermediate features of the neural implicit field network. The intermediate features of the neural implicit field network are input into the detail enhancement subnetwork, and the intermediate features of the neural implicit field network are processed by a multi-scale convolutional neural network to perform multi-scale feature fusion, resulting in fused intermediate features. The detail enhancement subnetwork processes the fused intermediate features through a multilayer perceptron to regress the displacement vector of a point in three-dimensional space, thereby obtaining a high-frequency detail displacement field.
[0010] As a preferred embodiment of the deep learning-based three-dimensional topography reconstruction method for surface fractures described in this invention, the method includes the following steps: enhancing the geometric details of the neural implicit field model using a high-frequency detail displacement field to obtain the enhanced neural implicit field model. Based on the neural implicit field model, query the initial geometric representation of a point in three-dimensional space within the neural implicit field model; The initial geometric representation of three-dimensional spatial points in the neural implicit field model is adjusted using a high-frequency detail displacement field to obtain the adjusted geometric representation; The neural implicit field model is updated based on the adjusted geometric representation to generate an enhanced neural implicit field model.
[0011] As a preferred embodiment of the deep learning-based three-dimensional topography reconstruction method for surface fractures described in this invention, the method includes the following steps: Based on the enhanced neural implicit field model, the neural implicit field network and the detail enhancement sub-network are jointly trained and their parameters optimized to obtain a trained deep learning model: Based on the enhanced neural implicit field model, the joint loss function value of the neural implicit field network and the detail enhancement sub-network is calculated; Based on the joint loss function value of the neural implicit field network and the detail enhancement subnetwork, an optimization strategy that integrates adaptive momentum estimation, dynamic balancing of multi-task loss weights, and gradient pruning mechanism is adopted to update the parameters of the neural implicit field network and the detail enhancement subnetwork. Iterative execution of loss calculation and parameter update based on the enhanced neural implicit field model satisfies the convergence condition, resulting in a well-trained deep learning model.
[0012] As a preferred embodiment of the deep learning-based surface fracture three-dimensional topography reconstruction method of the present invention, the method includes the following steps: obtaining the symbolic distance field of the target region using a trained deep learning model. Input the multi-scale depth feature maps and camera pose parameters into the trained deep learning model; Based on the neural implicit field network and detail enhancement subnetwork in the trained deep learning model, multi-scale deep feature maps and camera pose parameters are processed to generate symbolic distance values of 3D spatial points in the target region. Based on the symbolic distance values of the three-dimensional spatial points in the target region, construct the symbolic distance field of the target region.
[0013] As a preferred embodiment of the deep learning-based surface fracture three-dimensional topography reconstruction method of the present invention, the method includes the following steps: applying the moving cube algorithm to extract the zero isosurface of the symbolic distance field and generating a triangular mesh model: Discretize the symbolic distance field into a voxel grid to obtain the vertex coordinates and symbolic distance values of the voxel grid; Based on the vertex coordinates and sign distance of the voxel mesh, the intersection of each voxel cube with the zero isosurface is found using the moving cube algorithm; Based on the intersection of the voxel cube and the zero isosurface, the positions of the isosurface points and the normal vectors are calculated by interpolation on the edges of the voxel cube. Based on the positions of the isopleth points and the normal vectors, triangular facets are generated by connecting them to obtain a triangular mesh model.
[0014] As a preferred embodiment of the deep learning-based three-dimensional surface fracture topography reconstruction method of the present invention, the method includes the following steps: post-processing the triangular mesh model to output a three-dimensional surface fracture topography mesh model: Perform a mesh simplification operation on the triangular mesh model to generate a simplified triangular mesh model, and then perform a Laplacian smoothing operation on the simplified triangular mesh model to generate a smoothed triangular mesh model. The smoothed triangular mesh model is subjected to manifold checks and repair operations to generate a watertight triangular mesh model, which is then used as a three-dimensional morphological mesh model of the surface fracture.
[0015] Secondly, the present invention provides a deep learning-based three-dimensional topography reconstruction system for surface fractures, including a feature extraction module that acquires multi-view optical remote sensing images, processes the multi-view optical remote sensing images using motion recovery structure technology, generates camera pose parameters and sparse three-dimensional point clouds, and performs preprocessing and multi-scale feature extraction to obtain a multi-scale depth feature map. The inference module inputs multi-scale depth feature maps and camera pose parameters into the neural implicit field network, and generates a neural implicit field model through forward inference. The enhancement module extracts intermediate features from the neural implicit field network, inputs the intermediate features into the detail enhancement sub-network, generates a high-frequency detail displacement field, and uses the high-frequency detail displacement field to perform geometric detail enhancement on the neural implicit field model, resulting in the enhanced neural implicit field model. The training module, based on the enhanced neural implicit field model, performs joint training and parameter optimization on the neural implicit field network and the detail enhancement sub-network to obtain a well-trained deep learning model. The processing module uses a trained deep learning model to obtain the symbolic distance field of the target area, applies the moving cube algorithm to extract the zero isosurface from the symbolic distance field, generates a triangular mesh model, performs post-processing on the triangular mesh model, and outputs a three-dimensional morphological mesh model of the surface fracture.
[0016] The beneficial effects of this invention are as follows: By acquiring multi-view optical remote sensing images, the camera pose is restored using structure-of-motion (SOG) technology, and a sparse 3D point cloud is generated. Multi-scale depth feature maps are then extracted, and these maps, along with camera pose parameters, are input into a neural implicit field network. Differentiable volume rendering is then used to generate a neural implicit field model capable of continuously representing a 3D scene. To improve detail reconstruction capabilities, intermediate features are extracted from the neural implicit field network and input into a dedicated detail enhancement subnetwork to regress the high-frequency detail displacement field. This displacement field is then used to geometrically enhance the neural implicit field model. Based on the enhanced model, the model is further refined using photometric and geometrical methods. A joint loss function consisting of adversarial and regularization terms is used to jointly train and optimize the neural implicit field network and the detail enhancement subnetwork end-to-end, resulting in a well-trained deep learning model. The trained model is then used to predict the signed distance field of the target region. The moving cube algorithm is applied to extract the zero isosurface to generate the initial triangular mesh. After mesh simplification, smoothing, and manifold repair, a detailed and topologically correct 3D surface fracture morphology mesh model is output. Through the synergistic effect of the neural implicit field network and the detail enhancement network, the problems of insufficient accuracy and poor surface continuity in the reconstruction of surface fracture micro-topography details in open environments are solved. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Fig. 1 This is a flowchart of a deep learning-based method for reconstructing the three-dimensional topography of surface fractures.
[0019] Fig. 2 This is a schematic diagram of a deep learning-based system for reconstructing the three-dimensional topography of surface fractures. Detailed Implementation
[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0021] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0022] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0023] Reference Figs. 1-2 This is one embodiment of the present invention, which provides a method for reconstructing the three-dimensional topography of surface fractures based on deep learning, including the following steps: S1. Acquire multi-view optical remote sensing images, process the multi-view optical remote sensing images using motion reconstruction technology, generate camera pose parameters and sparse 3D point clouds, and perform preprocessing and multi-scale feature extraction to obtain multi-scale depth feature maps.
[0024] S1.1 Acquire multi-view optical remote sensing images, extract feature points from the multi-view optical remote sensing images, obtain feature points of the multi-view optical remote sensing images, match the feature points of the multi-view optical remote sensing images, and obtain feature point matching pairs of the multi-view optical remote sensing images.
[0025] Furthermore, a strategy of fusing complementary features is adopted, such as simultaneously utilizing feature points based on gradient information and feature points based on affine invariant regions. When extracting feature points from multi-view optical remote sensing images, the algorithm acquires local gradient extrema and scale-space stable regions in parallel, ensuring sufficient feature points are obtained in both texture-scarce fracture regions and texture-rich background regions. When matching feature points from multi-view optical remote sensing images, a cascaded matching strategy based on descriptor similarity and geometric constraints is adopted. Initial matching is performed using the nearest neighbor distance ratio of feature descriptors to obtain candidate feature point matching pairs for multi-view optical remote sensing images. Using the idea of random sampling consensus algorithm, the algorithm iteratively searches for interior point sets that satisfy the geometric constraints of the fundamental matrix among the candidate matching pairs, thereby filtering out mismatches caused by repeated textures or lighting variations, and obtaining accurate and reliable feature point matching pairs for multi-view optical remote sensing images.
[0026] Specifically, the two-stage strategy of feature fusion and geometric verification effectively addresses the challenges of weak texture and repetitive patterns commonly found in surface fracture scenarios, providing high-quality input data for subsequent geometric restoration, improving the recall and precision of feature matching, and obtaining accurate feature point matching pairs from multi-view optical remote sensing images.
[0027] S1.2. By matching feature points of multi-view optical remote sensing images, the camera projection geometry is calculated to restore the position and orientation of each multi-view optical remote sensing image in three-dimensional space, and the camera pose parameters of the multi-view optical remote sensing images are obtained.
[0028] Furthermore, when solving the camera projection geometry using feature point matching pairs from multi-view optical remote sensing images, the process begins with an initial two-view relationship based on a high-confidence matching pair. New multi-view optical remote sensing images are then added gradually, and global binding adjustments are optimized. When restoring the position and orientation of each multi-view optical remote sensing image in 3D space, a weight is assigned to each feature point matching pair. This weight is based on a dynamic evaluation of reprojection error, ensuring that matching pairs contributing significantly to overall geometric consistency receive higher weights in the optimization process, thereby suppressing the impact of anomalous matching.
[0029] Specifically, based on its adaptive robust optimization mechanism, it allows for accurate recovery of camera pose parameters from multi-view optical remote sensing images even with a certain proportion of mismatches, avoiding the model collapse problem caused by mismatches in traditional methods. For example, even if some broken regions have sparse features, this method can still stably solve the problem by matching other reliable regions. This yields accurate camera pose parameters for multi-view optical remote sensing images. S1.3. Utilize the camera pose parameters of multi-view optical remote sensing images and feature point matching pairs of multi-view optical remote sensing images to generate sparse three-dimensional point clouds.
[0030] Furthermore, when generating sparse 3D point clouds using camera pose parameters and feature point matching pairs from multi-view optical remote sensing images, a method combining triangulation and uncertainty joint evaluation is employed. For each successfully matched feature point pair from multi-view optical remote sensing images, its 3D spatial position is calculated using a triangulation method that minimizes reprojection error, based on the corresponding camera pose parameters. When generating each 3D point, the spatial distribution covariance matrix of its triangulation error is simultaneously estimated. Based on this covariance matrix, an uncertainty ellipsoid is modeled for the 3D point position. A filtering strategy based on joint constraints of spatial density and uncertainty is applied to eliminate abnormal 3D points located in areas with excessively large uncertainty ellipsoid volumes or overly isolated spatial distributions.
[0031] Specifically, it no longer generates a simple set of three-dimensional points, but a sparse three-dimensional point cloud with confidence evaluation. This point cloud is more uniform in spatial distribution and more reliable in geometric accuracy, providing a high-quality geometric constraint basis for subsequent deep learning feature extraction and generating a high-quality sparse three-dimensional point cloud.
[0032] S1.4. Preprocess the multi-view optical remote sensing image using the camera pose parameters of the multi-view optical remote sensing image to obtain the preprocessed multi-view optical remote sensing image. Extract multi-scale features from the preprocessed multi-view optical remote sensing image to obtain a multi-scale depth feature map.
[0033] Furthermore, the camera pose parameters of the multi-view optical remote sensing images are used for preprocessing. The purpose is not only to perform radiometric normalization, but more importantly, to perform view correction and region attention focusing based on 3D geometry. For example, based on sparse 3D point clouds and camera pose parameters, the multi-view optical remote sensing images can be reprojected onto a unified reference plane, reducing perspective distortion and concentrating resources on spatial regions where surface fractures may exist. When extracting multi-scale features from the preprocessed multi-view optical remote sensing images, a multi-level encoder network with lateral connections is employed. Shallow network convolutional kernels have small receptive fields, responsible for capturing fine texture features such as fracture edges and microcracks in the preprocessed multi-view optical remote sensing images; deep network convolutional kernels have large receptive fields, responsible for capturing the overall orientation and contextual features of the fracture zones in the preprocessed multi-view optical remote sensing images.
[0034] Specifically, by using skip connections to fuse feature maps from different levels, the resulting multi-scale deep feature map contains both local details and global semantic information. This solves the problem of insufficient representation of single-scale features when facing multi-scale morphology of surface fracturing, and provides subsequent neural implicit field networks with multi-scale deep feature maps that contain both rich texture details and strong semantic perception capabilities, resulting in information-rich multi-scale deep feature maps.
[0035] S2. Input the multi-scale depth feature map and camera pose parameters into the neural implicit field network, and generate the neural implicit field model through forward inference.
[0036] S2.1 Input the multi-scale depth feature map and camera pose parameters into the neural implicit field network. The neural implicit field network fuses the multi-scale depth feature map and camera pose parameters through a learnable view-dependent feature modulation module to construct the features of three-dimensional spatial points.
[0037] Furthermore, for a query point in 3D space, its coordinates are projected onto the multi-scale depth feature map corresponding to each multi-view optical remote sensing image using camera pose parameters, sampling a set of visual feature vectors associated with different viewpoints and scales. The learnable view-dependent feature modulation module does not directly concatenate these features, but instead learns a dynamic weight network that uses the camera's line-of-sight direction as an additional conditional input to assign appropriate importance weights to feature vectors from different viewpoints. The weighted features are then fused with the 3D position encoding of the query point, allowing the neural implicit field network to adaptively select the most relevant visual cues describing the appearance of a point in 3D space from the current viewpoint.
[0038] Specifically, for example, for a broken surface point that may be occluded, the module can reduce the feature weight from the occluded viewpoint and increase the feature contribution from the visible viewpoint. This realizes the conditionalization of two-dimensional image features to three-dimensional spatial features and the transformation of viewpoint perception. The features of the constructed three-dimensional spatial point not only contain multi-scale appearance information, but also implicitly contain its visibility relationship under different views, thus providing a more robust and expressive feature representation for subsequent decoding, and constructing the features of three-dimensional spatial points that integrate multi-view and multi-scale information.
[0039] S2.2 The neural implicit field network uses a multilayer perceptron to decode the features of three-dimensional spatial points and regress the symbolic distance function value and color value of the three-dimensional spatial points.
[0040] Furthermore, the Neural Implicit Field Network uses a multilayer perceptron to decode the features of 3D spatial points. This decoding process branches into two parallel prediction heads. One prediction head focuses on geometry reconstruction, receiving the features of 3D spatial points and outputting a scalar value, namely the signed distance function value of the 3D spatial point, which defines the directed distance from the point to the nearest fracture surface. The other prediction head is responsible for appearance modeling. It receives the features of the same or slightly transformed 3D spatial points and outputs a vector representing color. The role of the multilayer perceptron here is to map the features of high-dimensional, abstract 3D spatial points to low-dimensional, physically meaningful geometric and appearance quantities through a series of nonlinear transformations. In this joint decoding strategy, the geometry and appearance decoding share most of the feature extraction but are separated in the final stage. This allows the prediction of the signed distance function value to be indirectly constrained by color and appearance information, because the same features need to explain both the position and appearance of the point, prompting the learned features to contain richer scene information.
[0041] Specifically, for example, at the edge of a break, abrupt changes in color and rapid changes in geometric distance will be correlated at the feature level, which helps the network to more accurately locate the zero isosurface, efficiently regress the geometric and appearance attributes necessary to describe the 3D scene from the unified features, provide complete per-point attributes for subsequent differentiable volume rendering, and regress the signed distance function value and color value of the 3D spatial point.
[0042] S2.3. By using a differentiable volume rendering algorithm, the symbolic distance function value and color value are accumulated along the camera light rays to generate a neural implicit field model.
[0043] Furthermore, a neural implicit field model is generated through a differentiable volume rendering algorithm. This algorithm accumulates and integrates the signed distance function values and color values of a series of 3D spatial points along a virtual ray emanating from the camera center and passing through each pixel. It strategically samples a series of 3D spatial points on each camera ray, queries the signed distance function values and color values at these points, and uses the signed distance function values to estimate the volume density at each sampling point. Typically, a differentiable transformation function maps the signed distance function values to density, resulting in high density near the zero isosurface and low density far from the surface. The volume rendering algorithm obtains the contribution weight of the color of each sampling point to the final pixel color of the ray in order from near to far. This weight is determined by the density of the point and the cumulative transmittance of all previous points. The colors of all sampling points are weighted and summed according to their weights to obtain the rendered color of the pixel. The entire rendering process is completely differentiable, and the gradient between the pixel rendered color and the signed distance function values and color values of all sampling points can flow.
[0044] Specifically, the internally represented continuous neural implicit field is connected to the externally observable 2D image through a physically inspired rendering equation, allowing the 2D image to be used as a supervisory signal to train the 3D field. The neural implicit field model is essentially this neural implicit field network that has learned how to render an image consistent with the input multi-view optical remote sensing image, containing the continuous geometry and appearance of the scene. End-to-end unsupervised or self-supervised learning from 2D images to a 3D continuous field is achieved, generating the neural implicit field model.
[0045] S3. Extract intermediate features from the neural implicit field network, input the intermediate features into the detail enhancement subnetwork, and generate a high-frequency detail displacement field.
[0046] S3.1 Extract intermediate features from the neural implicit field network to obtain the intermediate features of the neural implicit field network.
[0047] Furthermore, before the multilayer perceptron of the neural implicit field network decodes the symbolic distance function and color values of 3D spatial points, the output of its hidden layers at a specific depth is extracted as intermediate features of the neural implicit field network. This operation does not extract the final output, but rather captures an intermediate, semantically rich abstract representation of the network in its understanding of scene geometry and appearance. It allows for strategic selection of feature extraction locations, such as extraction after location encoding or after several layers of shared feature transformations. The intermediate features of the neural implicit field network not only encode the coordinate information of 3D spatial points but also integrate visual context from multi-view, multi-scale depth feature maps, containing rich information about scene geometry, material properties, and local and global relationships.
[0048] Specifically, it leverages the inherent power of neural implicit field networks as a robust feature extractor. The intermediate features learned by these networks are more geometrically conscious and information-rich than the original image features, and more informative than the final signed distance function output. For example, in a scene of surface fracturing, these intermediate features may implicitly encode information such as the orientation of the fracturing edges and the textural differences between different lithological regions, resulting in intermediate features from a neural implicit field network capable of representing the deep semantic and geometric information of the scene.
[0049] S3.2 Input the intermediate features of the neural implicit field network into the detail enhancement subnetwork, and process the intermediate features of the neural implicit field network through a multi-scale convolutional neural network to perform multi-scale feature fusion and obtain the fused intermediate features.
[0050] Furthermore, after inputting the intermediate features of the neural implicit field network into the detail enhancement subnetwork, they are processed through a multi-scale convolutional neural network to construct a processing flow with a parallel multi-branch structure. Each branch corresponds to a specific receptive field scale, for example, constructed using dilated convolutional layers with different dilation rates, and processes the intermediate features of the neural implicit field network separately. Each branch independently extracts feature patterns related to that scale; for example, one branch focuses on high-frequency detail mutations, while another branch focuses on mesoscale structural contours. These multi-scale feature maps from different branches are fused through channel concatenation or weighted summation to obtain fused intermediate features, capturing and integrating information of different spatial frequencies related to the three-dimensional morphology of surface fracturing.
[0051] Specifically, the details of surface fracturing include microcracks and roughness at the millimeter to centimeter level, as well as steps and slope morphology at the decimeter to meter level, which cannot be fully covered by a single-scale convolutional kernel. By extracting and fusing multi-scale features in parallel, the detail enhancement subnetwork can simultaneously understand microscopic detail changes and macroscopic morphological trends, giving the network powerful multi-resolution representation capabilities. This ensures that the fused intermediate features can fully support the detailed modeling needs from subtle undulations to large-scale deformations, resulting in fused intermediate features that integrate multi-scale information.
[0052] S3.3 The detail enhancement subnetwork processes the fused intermediate features through a multilayer perceptron to regress the displacement vector of the three-dimensional spatial point, thus obtaining the high-frequency detail displacement field.
[0053] Furthermore, the detail enhancement subnetwork processes the fused intermediate features through a multilayer perceptron to regress the displacement vector of a point in 3D space. This multilayer perceptron acts as a refined regressor, taking the high-dimensional, fused intermediate features as input and transforming them through a series of fully connected layers and nonlinear activation functions, ultimately outputting a 3D vector, i.e., the displacement vector of the 3D point. This displacement vector defines the small 3D offset that needs to be applied to the basic geometry to characterize the high-frequency geometric details that were smoothed out or failed to reconstruct in the original neural implicit location. The detail enhancement problem is thus formulated as a residual displacement regression task based on rich features. By learning the complex mapping from fused features to displacement vectors, the multilayer perceptron can predict nonlinear detail deformations that are highly correlated with the local context.
[0054] Specifically, for example, at the edge of a fracture, it may predict a displacement in a normal direction to sharpen the boundary; on a rough fracture surface, it may predict a series of seemingly random but statistically consistent tiny displacements to increase surface roughness. The resulting high-frequency detail displacement field is a vector field defined in three-dimensional space, with a displacement vector at each location. This decouples the globally coherent base geometry from the locally flexible high-frequency details. The neural implicit field network is responsible for reconstructing the smooth and continuous main body shape, while the detail enhancement sub-network is responsible for predicting the detail residuals on it. The two work together to achieve high-fidelity reconstruction of surface fractures from macroscopic morphology to microscopic features, resulting in a high-frequency detail displacement field that can finely characterize the high-frequency geometric changes of the surface.
[0055] S4. Use high-frequency detail displacement field to enhance the geometric details of the neural implicit field model to obtain the enhanced neural implicit field model.
[0056] S4.1. Based on the neural implicit field model, query the initial geometric representation of a point in the three-dimensional space within the neural implicit field model.
[0057] Furthermore, the 3D coordinates of a point in 3D space are used as input and passed to a neural implicit field network for forward inference. The neural implicit field network outputs the signed distance function value of the 3D point. This signed distance function value defines the directed distance from the point to the nearest rupture surface, thus serving as the initial geometric representation. Utilizing the continuous function characteristics of the neural implicit field model, real-time and accurate geometric queries for any 3D point in space are achieved, avoiding the resolution limitations of traditional discrete representation methods such as voxels or point clouds. For example, for a point on a ruptured surface, the query process can directly obtain its precise positional relationship relative to the rupture surface, even if the point is located in a complex curved or branching region.
[0058] It features the ability to encode the geometry of the entire 3D scene into a differentiable neural network, making the geometric representation not only globally continuous but also computationally efficient. By parameterizing geometry through the neural network, it supports queries at infinite resolutions, thereby capturing the complex topological structure of surface fractures and obtaining a continuous and accurate basic geometric representation. This allows for subsequent enhancement using high-frequency detail displacement fields, resulting in the initial geometric representation of 3D spatial points in the neural implicit field model.
[0059] S4.2. Use high-frequency detail displacement fields to adjust the initial geometric representation of three-dimensional spatial points in the neural implicit field model to obtain the adjusted geometric representation.
[0060] Furthermore, by querying the high-frequency detail displacement field, the 3D displacement vector of the corresponding 3D spatial point is obtained. Then, this 3D displacement vector is added to the coordinates of the 3D spatial point to obtain the adjusted 3D coordinates. These new coordinates represent the geometric position after detail enhancement, i.e., the adjusted geometric representation. A residual learning strategy is used to model high-frequency details as small local perturbations to the basic geometry, rather than directly relearning the entire geometry. For example, for microcracks that might be smoothed out in the initial geometric representation, by adding a displacement vector perpendicular to the local surface, the depth and width details of the crack can be recovered in the adjusted geometric representation.
[0061] Specifically, the two sub-tasks of macroscopic geometric reconstruction and microscopic detail restoration are separated, allowing the neural implicit field network and the high-frequency detail displacement field to perform their respective functions and work together. This maintains the continuity of the main structure while introducing rich surface details. The geometric details are flexibly enhanced through the superposition of displacement vectors, allowing end-to-end training. It can add high-frequency morphological features locally and precisely without destroying the basic geometric coherence, resulting in an adjusted geometric expression.
[0062] S4.3. Update the neural implicit field model based on the adjusted geometric representation to generate an enhanced neural implicit field model.
[0063] Furthermore, by utilizing the new 3D coordinate set corresponding to the adjusted geometric representation and its theoretical symbolic distance function value of zero, additional training constraints are constructed. The parameters of the neural implicit field network are optimized using the backpropagation algorithm, making the symbolic distance function value predicted by the neural implicit field model for these 3D spatial points closer to zero, while maintaining geometric consistency in other regions, thus generating an enhanced neural implicit field model. An iterative optimization loop internalizes the external detail enhancement results into the internal representation of the neural implicit field model, achieving the solidification of detail information. For example, after multiple rounds of adjustment and network parameter updates using the high-frequency detail displacement field, the neural implicit field model can directly output the symbolic distance field of the fractured surface with micro-topographic details, no longer requiring explicit dependence on the high-frequency detail displacement field.
[0064] Specifically, a feedback loop from detail prediction to model update was established, enabling the neural implicit field model to actively learn and absorb high-frequency geometric features, continuously improving its own expression accuracy. Through a differentiable optimization process, discrete displacement adjustments were integrated into the continuous implicit field, achieving persistence and generalization of the detail enhancement effect. Ultimately, a powerful, self-contained neural implicit field model was obtained, which can independently complete the continuous representation of high-detail surface fracture 3D morphology and generate the enhanced neural implicit field model.
[0065] S5. Based on the enhanced neural implicit field model, the neural implicit field network and the detail enhancement sub-network are jointly trained and their parameters are optimized to obtain a well-trained deep learning model.
[0066] S5.1 Based on the enhanced neural implicit field model, calculate the joint loss function value of the neural implicit field network and the detail enhancement sub-network; Furthermore, based on the enhanced neural implicit field model, the joint loss function value of the neural implicit field network and the detail enhancement sub-network is calculated. The photometric reconstruction loss re-renders the enhanced neural implicit field model into a two-dimensional image through differentiable volume rendering, and calculates the pixel-level differences between these rendered images and the input multi-view optical remote sensing images. Its role is to force the appearance prediction of the enhanced neural implicit field model to be consistent with the actual observation. The geometric consistency loss is achieved by comparing the distance between the zero isosurface extracted from the enhanced neural implicit field model and the sparse 3D point cloud. Its role is to constrain the learned continuous geometry to discrete but precise observation points. The adversarial loss evaluates the local geometric realism of the surface represented by the enhanced neural implicit field model through a pre-trained 3D local geometry discriminator. The discriminator attempts to distinguish the surface patches sampled from the enhanced neural implicit field model from the surface patches sampled from high-precision real scan data. Its role is to drive the enhanced neural implicit field model to generate high-frequency geometry that is indistinguishable from the real data in terms of detail statistical distribution. Regularization loss typically constrains the gradient norm of the symbolic distance function value or the average curvature of the surface predicted by the enhanced neural implicit field model. Its purpose is to encourage the reconstructed surface to be smoother and more reasonable, avoiding unnatural and violent geometric oscillations.
[0067] Specifically, a multi-task loss function was constructed that integrates four different supervision signals: appearance, sparse geometry, detail statistics, and prior smoothness. Addressing the characteristics of incomplete data and high detail requirements in surface fracture reconstruction, complementary optimization objectives were provided from different dimensions. For example, the introduction of adversarial loss specifically addresses the problem of blurred details in traditional methods when dense ground truth supervision is lacking. Through adversarial learning, detail distribution patterns are extracted from the data itself. The joint loss function unifies multiple constraints—two-dimensional image consistency, three-dimensional geometric consistency, detail realism, and geometric priors—within a differentiable optimization framework. This allows the neural implicit field network and the detail enhancement sub-network to be trained collaboratively, jointly striving to generate a three-dimensional model that conforms to multi-view observation, fits sparse point clouds, and possesses rich and realistic details.
[0068] The expression for the joint loss function value is: ; in, for, The weighting coefficients for photometric reconstruction loss are: For the loss of photometric reconstruction, The weighting coefficients for photometric reconstruction loss are: For geometric consistency loss, The regularization loss is the result of the geometric consistency loss. To combat the losses, These are the weighting coefficients for the regularization loss. This is the loss due to regularization.
[0069] S5.2 Based on the joint loss function value of the neural implicit field network and the detail enhancement subnetwork, an optimization strategy that integrates adaptive momentum estimation, dynamic balancing of multi-task loss weights, and gradient pruning mechanism is adopted to update the parameters of the neural implicit field network and the detail enhancement subnetwork.
[0070] Furthermore, the adaptive momentum estimation algorithm estimates the first and second moments of historical gradients, calculates an adaptive learning rate for each parameter, and applies momentum to accelerate convergence. This makes the parameter update direction more stable and faster when optimizing complex joint loss function surfaces. The multi-task loss weight dynamic balancing mechanism continuously monitors the magnitude or rate of change of photometric reconstruction loss, geometric consistency loss, adversarial loss, and regularization loss during training, and dynamically fine-tunes the weight coefficients. and To prevent a single loss term from dominating the optimization process and causing other tasks to be neglected, this ensures that the four optimization objectives are pursued in a balanced manner. The gradient pruning mechanism checks the gradient norm of all parameters to be updated before each parameter update. If it exceeds a preset threshold, the gradient vector is scaled proportionally. This effectively prevents gradient explosion problems that may occur during joint training due to the complexity of the loss function, ensuring the numerical stability of the training process.
[0071] Specifically, joint training based on the Neural Implicit Field Network (NNFR) and the Details Enhancement Subnetwork is a typical complex multi-task optimization problem. Different loss terms have varying dimensions and convergence speeds, and may even compete with each other. Using a basic optimizer with fixed loss weights is unlikely to achieve optimal results. This paper integrates robust updates of adaptive momentum estimation, dynamic balancing of multi-task loss weights, and stable protection of gradient clipping to form a robust and efficient optimization strategy. This allows the training process to automatically adapt to changes in the loss landscape at different training stages and with different data batches. For example, in the early stages of training, the model may rely more on photometric reconstruction loss and geometric consistency loss to establish basic geometry, and the dynamic balancing mechanism will assign them higher weights. As training progresses, the adversarial loss becomes crucial for enhancing details, and its weight may be increased. This ensures that the entire deep learning model can learn smoothly and fully, completing one update of the parameters of the NNFR and Details Enhancement Subnetwork.
[0072] S5.3 Iteratively execute the loss calculation and parameter update based on the enhanced neural implicit field model to meet the convergence condition and obtain the trained deep learning model.
[0073] Furthermore, starting from the initialized neural implicit field network and detail enhancement subnetwork, a complete forward propagation is performed. That is, starting from the input multi-scale depth feature map and camera pose parameters, neural implicit field inference, detail enhancement, and geometric adjustment are performed sequentially to generate the current enhanced neural implicit field model. Then, based on the current enhanced neural implicit field model, the joint loss function value is calculated. Based on the joint loss function value, the gradient is calculated according to the fusion optimization strategy and the parameters of the neural implicit field network and detail enhancement subnetwork are updated. After completing one iteration, the updated network parameters are used to process the next batch of training data, and a new iteration begins until the preset convergence condition is met, such as the joint loss function value no longer decreasing significantly in consecutive iterations, or the preset maximum number of iterations is reached.
[0074] Specifically, the entire training framework is constructed as an end-to-end, differentiable closed loop that continuously improves itself through iterative optimization. It gradually learns the key to complex 3D reconstruction mapping from a randomly initialized state. Each parameter update brings the enhanced neural implicit field model closer to a better solution under the constraints of various loss metrics. Through iterative optimization, a deep synergy is formed between the neural implicit field network and the detail enhancement subnetwork: the neural implicit field network provides increasingly better basic geometry, giving the detail enhancement subnetwork a more accurate working context; while the high-frequency detail displacement fields predicted by the detail enhancement subnetwork are fed back to the neural implicit field network through loss calculation and parameter updates, enabling its basic geometry to better accept and integrate these details. The resulting well-trained deep learning model is a set of stable network parameters that have been fully optimized through numerous iterations. It encapsulates the complete ability to reconstruct the 3D morphology of surface fractures with high fidelity from multi-view images, thus obtaining a well-trained deep learning model.
[0075] S6. Using the trained deep learning model, obtain the symbolic distance field of the target region.
[0076] S6.1 Input the multi-scale depth feature map and camera pose parameters into the trained deep learning model.
[0077] Furthermore, multi-view optical remote sensing images of the target area are acquired through a preprocessing procedure. Using the same motion reconstruction structure technique and multi-scale feature extraction method as in the training phase, multi-scale depth feature maps and camera pose parameters of the target area are generated, consistent with the training data format. The trained deep learning model encapsulates a fully jointly trained and parameter-optimized neural implicit field network and detail enhancement subnetwork. It directly receives these standardized inputs and performs forward inference, utilizing a pre-trained, parameter-fixed deep learning model to process new target area data, avoiding the computational cost and time overhead of retraining the model for each new scene.
[0078] Specifically, the complex mapping relationships learned from a large amount of data during the training phase—that is, how to infer 3D geometry from multi-view features and camera geometry—are solidified into the model weights, giving the model a powerful generalization ability. For example, when faced with an earthquake rupture zone that has never been seen before, it is only necessary to collect its remote sensing images and extract multi-scale depth feature maps and camera pose parameters, and input them into the trained deep learning model to start reconstruction. This improves the efficiency of emergency response, realizes the convenience of plug-and-play model, and ensures the consistency and repeatability of the reconstruction process. It also ensures that the multi-scale depth feature maps and camera pose parameters of the target area are correctly fed into the trained deep learning model.
[0079] S6.2. Based on the neural implicit field network and detail enhancement subnetwork in the trained deep learning model, process multi-scale deep feature maps and camera pose parameters to generate symbolic distance values of 3D spatial points in the target region.
[0080] Furthermore, the trained deep learning model sequentially calls its internal Neural Implicit Field Network and Detail Enhancement Subnetwork for forward inference: The Neural Implicit Field Network fuses multi-scale depth feature maps and camera pose parameters through a learnable view-dependent feature modulation module to construct features of 3D spatial points. It uses a multilayer perceptron to decode these features and initially regresses the symbolic distance function value and color value of the 3D spatial points. The Detail Enhancement Subnetwork extracts intermediate features from the intermediate layers of the Neural Implicit Field Network, performs multi-scale feature fusion through a multi-scale convolutional neural network, and then regresses the high-frequency detail displacement field through a multilayer perceptron. This displacement field is used to perform geometric detail enhancement on the initial symbolic distance function value output by the Neural Implicit Field Network, and finally outputs the adjusted, detail-rich symbolic distance value of the 3D spatial points of the target region.
[0081] Specifically, the cascading and synergistic effects of the neural implicit field network and the detail enhancement subnetwork in the inference process are achieved through the seamless connection of the two through trained parameters. Together, they complete the geometric prediction from coarse to fine, integrating macroscopic geometric reconstruction and microscopic detail enhancement into a single forward propagation process. This allows the generated symbolic distance value to simultaneously encode the overall morphology of the fracture and local micro-topographic features. For example, for a surface fracture, the symbolic distance value can not only reflect its main direction and dip angle, but also accurately indicate the changes in fracture width and subtle undulations of the steep slope along the line. High-fidelity, multi-scale geometric information can be obtained through a single forward propagation, avoiding the information loss or error propagation that may be caused by multi-stage processing in traditional methods, and generating accurate and detailed symbolic distance values for three-dimensional spatial points in the target area.
[0082] S6.3 Construct the symbolic distance field of the target region based on the symbolic distance values of the three-dimensional spatial points of the target region.
[0083] Furthermore, a regular sampling grid is defined within the three-dimensional space of the target region, with each node of the grid corresponding to a three-dimensional spatial point. A trained deep learning model is used to batch query or calculate symbolic distance values for these three-dimensional spatial points. Then, the symbolic distance values of all three-dimensional spatial points are organized according to their spatial location into a three-dimensional array or a continuous scalar field function, thus forming a complete symbolic distance field for the target region. The symbolic distance field mathematically represents the directed distance from each location in the entire spatial domain to the nearest surface rupture, where the zero isosurface corresponds to the three-dimensional morphology of the rupture. The continuous symbolic distance field is constructed using symbolic distance values predicted by the deep learning model, rather than through interpolation or fitting of discrete point clouds. This is because the neural implicit field network essentially learns a continuous function.
[0084] Specifically, it achieves a direct conversion from deep learning model output to a continuous geometric representation that can be used for surface reconstruction, where the symbolic distance field is a differentiable mathematical entity. For example, by increasing the resolution of the sampling grid, a symbolic distance field of arbitrary precision can be constructed, thereby capturing full-spectrum geometric features from meters to centimeters in surface fractures. The symbolic distance field provides a complete, consistent, and high-resolution geometric description, supporting subsequent surface extraction algorithms to generate watertight and accurate triangular mesh models, while ensuring end-to-end differentiability and geometric consistency of the entire reconstruction process.
[0085] S7. Apply the moving cube algorithm to the symbolic distance field to extract the zero isosurface and generate a triangular mesh model.
[0086] S7.1 Discretize the symbolic distance field into a voxel grid to obtain the vertex coordinates and symbolic distance values of the voxel grid.
[0087] Furthermore, discretizing the symbolic distance field into a voxel mesh involves defining a regular 3D sampling lattice within the 3D space of the target region. Each node in this lattice is called a voxel vertex, and the set of all voxel vertices constitutes the geometric framework of the voxel mesh. Each voxel vertex has definite 3D coordinates, i.e., the vertex coordinates of the voxel mesh. Simultaneously, by querying the symbolic distance field, the symbolic distance value corresponding to the coordinates of each voxel vertex can be obtained, thus obtaining a complete correspondence between the vertex coordinates and the symbolic distance values of the voxel mesh. Using the voxel mesh as a bridge connecting the continuous symbolic distance field and the discrete triangular mesh, the resolution of the voxel mesh determines the upper limit of detail capture of the final reconstructed surface. To meet the needs of surface fracture reconstruction, denser sampling can be used in areas where fracture zones may be distributed, while sparser sampling can be used in background areas. This non-uniform discretization strategy can ensure sufficient sampling of fracture micro-topographic details without increasing computational load.
[0088] Specifically, by discretizing with controllable precision, the continuous geometric functions learned by the neural network are transformed into regular data structures that can be processed by classical computer graphics algorithms, providing standardized input for subsequent moving cube algorithms. For example, for a crack with varying width, a fine voxel mesh ensures that there are still enough voxel units to characterize its boundary at the narrowest point of the crack, avoiding the loss of geometric information. This achieves the transformation from an abstract field representation to a concrete, operable intermediate geometric representation, enabling the subsequent surface extraction process to be carried out efficiently and systematically, obtaining the vertex coordinates and signed distance values of the structured voxel mesh.
[0089] S7.2 Based on the vertex coordinates and symbol distance values of the voxel mesh, the intersection of each voxel cube with the zero isosurface is found using the moving cube algorithm.
[0090] Furthermore, based on the vertex coordinates and signed distance values of the voxel mesh, the moving cube algorithm is used to find the intersection of each voxel cube with the zero isosurface. The moving cube algorithm treats the voxel mesh as composed of many small cubic units, i.e., voxel cubes. The algorithm processes each voxel cube sequentially. Its core operation is to compare the sign of the signed distance values at the eight corner points of the voxel cube. Depending on whether each corner point is inside the zero isosurface and the signed distance value is negative or outside and the signed distance value is positive, 256 possible sign combinations can be generated. Each sign combination corresponds to a predefined topological configuration of the voxel cube intersecting with the zero isosurface. By consulting a pre-built lookup table, it can be immediately determined whether the current voxel cube is crossed by the zero isosurface and which edges it crosses. Using the sign change of the signed distance value as a strong and computationally simple criterion for detecting the existence of the surface, the moving cube algorithm achieves extremely high parallel processing efficiency by decomposing the globally complex isosurface extraction problem into a large number of identical and independent local cube configuration judgment problems.
[0091] Specifically, the algorithm transforms the search for 3D isosurfaces into the detection of sign changes on 1D edges. This allows the algorithm to robustly generate triangulated surfaces even for geometries with complex branches, holes, and open boundaries, such as surface fractures, without requiring any prior knowledge of surface topology. For example, in a complex region containing multiple intersecting fractures, the moving cube algorithm can automatically and correctly handle each intersecting voxel cube, ensuring that all local triangular faces can be pieced together to form a coherent global surface. It accurately reconstructs the spatial orientation of continuous isosurfaces from discretely sampled voxel data and classifies all intersections between voxel cubes and zero isosurfaces.
[0092] S7.3. Based on the intersection of the voxel cube and the zero isosurface, interpolate the positions of the isosurface points and the normal vectors on the edges of the voxel cube.
[0093] Furthermore, the precise geometric location and fine local orientation of isosurfaces are accurately reconstructed from low-resolution voxel sampling data. It is recognized that the signed distance field is usually smooth and monotonically changing near the zero isosurface, so the linear interpolation assumption is reasonable and efficient. The surface normal, as the normalized result of the gradient, directly reflects the local geometric properties of the surface. For example, at the location of a steep slope with surface fracturing, the normal vector calculated by the above method will clearly reflect the slope direction, and the isosurface points obtained by interpolation can accurately fall on the slope toe or crest line of the steep slope. With minimal computational overhead, the vertex positions and vertex normal properties of the surface are obtained, laying the foundation for generating a correctly lit and visually smooth triangular mesh model, and calculating all the necessary isosurface point positions and normal vectors.
[0094] The expression for the location of the isopleth point is: ; in, Let be the coordinates of the intersection point of the zero isosurface and the edge of the voxel cube. Let be the coordinates of one endpoint of the voxel cube's edge. As endpoints The symbolic distance value at that location, The other endpoint of the voxel cube's edge The symbolic distance value at that location, Let be the coordinates of the other endpoint of the voxel cube's edge.
[0095] The expression for the normal vector is: ; in, isovalue points The unit normal vector at that point, The sign distance field at the isopleths The gradient vector at that point.
[0096] S7.4. Based on the position of the isopleths and the normal vector, connect them to generate triangular facets, thus obtaining a triangular mesh model.
[0097] Furthermore, based on the positions of the isopleth points and the normal vectors, triangular facets are generated by connecting them. The moving cube algorithm reads the connection method of the triangular facets corresponding to the specific topological configuration of each voxel cube from a predefined triangulation template. This template indicates how to connect the isopleth points calculated on each edge of the voxel cube into one or more triangular facets in a specific order. Each triangular facet consists of three isopleth points, and each isopleth point carries its calculated normal vector. By collecting all the triangular facets generated after processing all voxel cubes, a triangular mesh model describing the entire zero isopleth surface is obtained.
[0098] Specifically, discrete isopleths distributed across millions of independent voxel cubes are automatically connected into a globally consistent, non-self-intersecting, and typically non-porous watertight triangular mesh surface. Through deterministic triangulation within local cubes, the algorithm ensures perfect matching of triangular facets generated between adjacent cubes on shared edges, implicitly maintaining the topological consistency of the global surface without requiring complex global stitching or deduplication algorithms. For example, for a complex, meandering surface fracture zone extending hundreds of meters, the moving cube algorithm can generate a single, continuous triangular mesh model that accurately represents its main fracture and all secondary cracks without producing breaks or incorrect connections. This completes the conversion from a scalar field to a vector surface model. The generated triangular mesh model is a universal 3D representation format that can be directly used by various 3D software, analysis tools, and visualization platforms, completing the key output from data to model of the 3D morphology of surface fractures—the triangular mesh model.
[0099] S8. Post-process the triangular mesh model to output a three-dimensional surface fracture topography mesh model.
[0100] S8.1 Perform a mesh simplification operation on the triangular mesh model to generate a simplified triangular mesh model. Perform a Laplacian smoothing operation on the simplified triangular mesh model to generate a smoothed triangular mesh model.
[0101] Furthermore, an iterative simplification algorithm based on edge folding is adopted. This algorithm traverses every edge in the triangular mesh model. If the geometric error caused by folding the edge is taken as the folding cost, the algorithm prioritizes iterating over the edge with the smallest folding cost. Each folding operation removes a vertex and its associated edges and faces, and updates the local connectivity relationships. This effectively reduces the number of vertices and faces in the triangular mesh model while preserving its original geometric features and contours to the greatest extent, especially important features such as fracture lines and steep slope edges. The simplified triangular mesh model is then subjected to a Laplacian smoothing operation to generate a smoothed triangular mesh model. The Laplacian smoothing operation is achieved by slightly adjusting the position of each vertex in the triangular mesh model towards the average position of its neighboring vertices. This adjustment is performed iteratively, and each round of smoothing makes the vertex position closer to its neighborhood center. This effectively suppresses high-frequency noise and irregular triangular face fluctuations that may be introduced by the moving cube algorithm extraction process or the mesh simplification operation, resulting in a smoother triangular mesh model with a smoother surface.
[0102] Specifically, a sequential processing strategy combining feature-preserving simplification and constrained smoothing was adopted. Triangular mesh models extracted from the signed distance field often have excessively high and unnecessary geometric complexity. Pure simplification may damage features, while pure smoothing may blur details. Simplification, measured by geometric error, reduces the amount of data while preserving both macroscopic and microscopic features. Laplacian smoothing then smooths the surface without excessively shifting features. The combination of these two approaches achieves the goal of reducing storage and computational overhead, improving mesh quality, and strictly maintaining the integrity of key morphological features of surface ruptures. For example, mesh simplification identifies and preserves the edges corresponding to continuous ridges representing the direction of rupture scarps, while smoothing primarily operates on large flat or gentle slope areas to eliminate step artifacts caused by voxel discretization, producing a lightweight and visually superior smoothed triangular mesh model.
[0103] S8.2 Perform manifold checks and repairs on the smoothed triangular mesh model to generate a watertight triangular mesh model, which is then used as the three-dimensional morphological mesh model of the surface fracture.
[0104] Furthermore, manifold checks and repair operations are performed on the smoothed triangular mesh model to generate a watertight triangular mesh model. The manifold check aims to detect non-manifold geometric elements in the smoothed triangular mesh model, including non-manifold vertices (a single vertex connected to multiple unconnected sectors), non-manifold edges (an edge shared by two or more faces), and isolated faces or overhanging edges. The repair operation applies specific algorithms to each detected non-manifold condition, such as splitting non-manifold vertices into multiple manifold vertices, copying and reconstructing duplicate faces at non-manifold edges, deleting isolated geometric elements, and possibly adding triangular faces to fill in the gaps caused by the repair operation. This ensures that each edge in the smoothed triangular mesh model is strictly shared by two faces, and that each vertex and its neighborhood are topologically homeomorphic to a disk, thus satisfying the manifold condition and generating a watertight triangular mesh model. The watertight triangular mesh model is then output as a three-dimensional surface fracture topology mesh model.
[0105] Specifically, manifold inspection and repair are key mandatory steps in generating the final usable model. Recognizing that the moving cube algorithm may produce non-manifold geometry in complex boundaries or undersampled regions, subsequent simplification and smoothing operations may exacerbate or introduce additional topological errors. A non-manifold triangular mesh model can cause serious problems in many downstream applications such as finite element analysis and 3D printing. Therefore, through systematic inspection and repair, it is ensured that the output model is a mathematically well-defined two-dimensional manifold, i.e., a watertight triangular mesh model. For example, at complex nodes where multiple cracks converge in a surface fracture, the initially extracted mesh may show multiple faces sharing an edge. Manifold repair can re-divide the mesh in this area into multiple topologically clear connected parts, ensuring the mathematical rigor and physical realizability of the surface fracture 3D topological mesh model. This allows it to be used seamlessly in various scientific analyses, engineering simulations, and visualization applications, marking the final transformation from geometric data to deliverable 3D assets, generating and outputting a watertight triangular mesh model as the final surface fracture 3D topological mesh model.
[0106] This embodiment also provides a deep learning-based three-dimensional surface fracture topography reconstruction system, including: a feature extraction module, which acquires multi-view optical remote sensing images, processes the multi-view optical remote sensing images through motion restoration structure technology, generates camera pose parameters and sparse three-dimensional point clouds, and performs preprocessing and multi-scale feature extraction to obtain a multi-scale depth feature map. The inference module inputs multi-scale depth feature maps and camera pose parameters into the neural implicit field network, and generates a neural implicit field model through forward inference. The enhancement module extracts intermediate features from the neural implicit field network, inputs the intermediate features into the detail enhancement sub-network, generates a high-frequency detail displacement field, and uses the high-frequency detail displacement field to perform geometric detail enhancement on the neural implicit field model, resulting in the enhanced neural implicit field model. The training module, based on the enhanced neural implicit field model, performs joint training and parameter optimization on the neural implicit field network and the detail enhancement sub-network to obtain a well-trained deep learning model. The processing module uses a trained deep learning model to obtain the symbolic distance field of the target area, applies the moving cube algorithm to extract the zero isosurface from the symbolic distance field, generates a triangular mesh model, performs post-processing on the triangular mesh model, and outputs a three-dimensional morphological mesh model of the surface fracture.
[0107] This embodiment also provides a computer device applicable to the deep learning-based method for reconstructing the three-dimensional topography of surface fractures, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the deep learning-based method for reconstructing the three-dimensional topography of surface fractures as proposed in the above embodiment.
[0108] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0109] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the deep learning-based method for reconstructing three-dimensional topography of surface fractures as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0110] In summary, this invention acquires multi-view optical remote sensing images, utilizes structure-of-motion (SOP) technology to recover camera pose and generate sparse 3D point clouds, and then extracts multi-scale depth feature maps. These multi-scale depth feature maps, along with camera pose parameters, are input into a neural implicit field network (NLP). Differentiable volume rendering is then used to generate a NLP model capable of continuously representing a 3D scene. To enhance detail reconstruction capabilities, intermediate features are extracted from the NLP and input into a dedicated detail enhancement subnetwork to regress high-frequency detail displacement fields. These displacement fields are then used to geometrically enhance the NLP model. Based on the enhanced model, photometric, geometric, and other parameters are used to further refine the model. A joint loss function with adversarial and regularization terms is used to jointly train and optimize the neural implicit field network and the detail enhancement subnetwork end-to-end, resulting in a trained deep learning model. The trained model is then used to predict the signed distance field of the target region. The moving cube algorithm is applied to extract the zero isosurface to generate the initial triangular mesh. After mesh simplification, smoothing, and manifold repair post-processing, a detailed and topologically correct 3D surface fracture morphology mesh model is output. Through the synergistic effect of the neural implicit field network and the detail enhancement network, the problems of insufficient accuracy and poor surface continuity in the reconstruction of surface fracture micro-topography details in open environments are solved.
[0111] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for reconstructing three-dimensional morphology of surface fractures based on deep learning, characterized in that: This includes acquiring multi-view optical remote sensing images, processing the multi-view optical remote sensing images using structure-of-motion reconstruction technology, generating camera pose parameters and sparse 3D point clouds, and performing preprocessing and multi-scale feature extraction to obtain multi-scale depth feature maps. Multi-scale depth feature maps and camera pose parameters are input into a neural implicit field network, and a neural implicit field model is generated through forward inference. Intermediate features are extracted from the neural implicit field network and input into the detail enhancement subnetwork to generate a high-frequency detail displacement field. The high-frequency detail displacement field is then used to enhance the geometric details of the neural implicit field model, resulting in the enhanced neural implicit field model. Based on the enhanced neural implicit field model, the neural implicit field network and the detail enhancement sub-network are jointly trained and their parameters are optimized to obtain a well-trained deep learning model. Using a trained deep learning model, the symbolic distance field of the target region is obtained. The moving cube algorithm is applied to the symbolic distance field to extract the zero isosurface, generating a triangular mesh model. The triangular mesh model is then post-processed to output a three-dimensional surface fracture topography mesh model.
2. The method for reconstructing three-dimensional surface fracture topography based on deep learning as described in claim 1, characterized in that: Acquire multi-view optical remote sensing images, process them using structure-of-motion (SOR) techniques to generate camera pose parameters and sparse 3D point clouds, and perform preprocessing and multi-scale feature extraction to obtain multi-scale depth feature maps. The process includes the following steps: Acquire multi-view optical remote sensing images, extract feature points from the multi-view optical remote sensing images, obtain feature points of the multi-view optical remote sensing images, match the feature points of the multi-view optical remote sensing images, and obtain feature point matching pairs of the multi-view optical remote sensing images. By matching feature points of multi-view optical remote sensing images, the camera projection geometry is solved, the position and orientation of each multi-view optical remote sensing image in three-dimensional space are recovered, and the camera pose parameters of the multi-view optical remote sensing images are obtained. Sparse 3D point clouds are generated by using camera pose parameters of multi-view optical remote sensing images and feature point matching pairs of multi-view optical remote sensing images. The multi-view optical remote sensing images are preprocessed using camera pose parameters to obtain preprocessed multi-view optical remote sensing images. Multi-scale feature extraction is then performed on the preprocessed multi-view optical remote sensing images to obtain multi-scale depth feature maps.
3. The method for reconstructing three-dimensional surface fracture topography based on deep learning as described in claim 2, characterized in that: The multi-scale depth feature maps and camera pose parameters are input into the neural implicit field network, and the neural implicit field model is generated through forward inference, including the following steps: Multi-scale depth feature maps and camera pose parameters are input into a neural implicit field network. The neural implicit field network fuses the multi-scale depth feature maps and camera pose parameters through a learnable view-dependent feature modulation module to construct features of three-dimensional spatial points. The neural implicit field network uses a multilayer perceptron to decode the features of three-dimensional spatial points and regress the symbolic distance function value and color value of the three-dimensional spatial points. By using a differentiable volume rendering algorithm, a neural implicit field model is generated by accumulating symbolic distance function values and color values along camera rays.
4. The method for reconstructing three-dimensional surface fracture topography based on deep learning as described in claim 3, characterized in that: The process involves extracting intermediate features from a neural implicit field network, inputting these features into a detail enhancement subnetwork, and generating a high-frequency detail displacement field. This includes the following steps: Intermediate features are extracted from the neural implicit field network to obtain the intermediate features of the neural implicit field network. The intermediate features of the neural implicit field network are input into the detail enhancement subnetwork, and the intermediate features of the neural implicit field network are processed by a multi-scale convolutional neural network to perform multi-scale feature fusion, resulting in fused intermediate features. The detail enhancement subnetwork processes the fused intermediate features through a multilayer perceptron to regress the displacement vector of a point in three-dimensional space, thereby obtaining a high-frequency detail displacement field.
5. The method for reconstructing three-dimensional surface fracture topography based on deep learning as described in claim 4, characterized in that: The neural implicit field model is geometrically enhanced using high-frequency detail displacement fields to obtain the enhanced neural implicit field model, including the following steps: Based on the neural implicit field model, query the initial geometric representation of a point in three-dimensional space within the neural implicit field model; The initial geometric representation of three-dimensional spatial points in the neural implicit field model is adjusted using a high-frequency detail displacement field to obtain the adjusted geometric representation; The neural implicit field model is updated based on the adjusted geometric representation to generate an enhanced neural implicit field model.
6. The method for reconstructing three-dimensional surface fracture topography based on deep learning as described in claim 5, characterized in that: Based on the enhanced neural implicit field model, the neural implicit field network and the detail enhancement subnetwork are jointly trained and their parameters optimized to obtain a trained deep learning model, including the following steps: Based on the enhanced neural implicit field model, the joint loss function value of the neural implicit field network and the detail enhancement sub-network is calculated; Based on the joint loss function value of the neural implicit field network and the detail enhancement subnetwork, an optimization strategy that integrates adaptive momentum estimation, dynamic balancing of multi-task loss weights, and gradient pruning mechanism is adopted to update the parameters of the neural implicit field network and the detail enhancement subnetwork. Iterative execution of loss calculation and parameter update based on the enhanced neural implicit field model satisfies the convergence condition, resulting in a well-trained deep learning model.
7. The method for reconstructing three-dimensional surface fracture topography based on deep learning as described in claim 6, characterized in that: Using a trained deep learning model, the symbolic distance field of the target region is obtained, including the following steps: Input the multi-scale depth feature maps and camera pose parameters into the trained deep learning model; Based on the neural implicit field network and detail enhancement subnetwork in the trained deep learning model, multi-scale deep feature maps and camera pose parameters are processed to generate symbolic distance values of 3D spatial points in the target region. Based on the symbolic distance values of the three-dimensional spatial points in the target region, construct the symbolic distance field of the target region.
8. The method for reconstructing three-dimensional surface fracture topography based on deep learning as described in claim 7, characterized in that: The moving cubes algorithm is applied to the symbolic distance field to extract the zero isosurface and generate a triangular mesh model, including the following steps: Discretize the symbolic distance field into a voxel grid to obtain the vertex coordinates and symbolic distance values of the voxel grid; Based on the vertex coordinates and sign distance of the voxel mesh, the intersection of each voxel cube with the zero isosurface is found using the moving cube algorithm; Based on the intersection of the voxel cube and the zero isosurface, the positions of the isosurface points and the normal vectors are calculated by interpolation on the edges of the voxel cube. Based on the positions of the isopleth points and the normal vectors, triangular facets are generated by connecting them to obtain a triangular mesh model.
9. The method for reconstructing three-dimensional surface fracture topography based on deep learning as described in claim 8, characterized in that: Post-processing the triangular mesh model to output a three-dimensional surface fracture topography mesh model includes the following steps: Perform a mesh simplification operation on the triangular mesh model to generate a simplified triangular mesh model, and then perform a Laplacian smoothing operation on the simplified triangular mesh model to generate a smoothed triangular mesh model. The smoothed triangular mesh model is subjected to manifold checks and repair operations to generate a watertight triangular mesh model, which is then used as a three-dimensional morphological mesh model of the surface fracture.
10. A deep learning-based three-dimensional surface fracture topography reconstruction system, based on the deep learning-based three-dimensional surface fracture topography reconstruction method according to any one of claims 1 to 9, characterized in that: This includes a feature extraction module, which acquires multi-view optical remote sensing images, processes the multi-view optical remote sensing images using motion reconstruction technology, generates camera pose parameters and sparse 3D point clouds, and performs preprocessing and multi-scale feature extraction to obtain multi-scale depth feature maps. The inference module inputs multi-scale depth feature maps and camera pose parameters into the neural implicit field network, and generates a neural implicit field model through forward inference. The enhancement module extracts intermediate features from the neural implicit field network, inputs the intermediate features into the detail enhancement sub-network, generates a high-frequency detail displacement field, and uses the high-frequency detail displacement field to perform geometric detail enhancement on the neural implicit field model, resulting in the enhanced neural implicit field model. The training module, based on the enhanced neural implicit field model, performs joint training and parameter optimization on the neural implicit field network and the detail enhancement sub-network to obtain a well-trained deep learning model. The processing module uses a trained deep learning model to obtain the symbolic distance field of the target area, applies the moving cube algorithm to extract the zero isosurface from the symbolic distance field, generates a triangular mesh model, performs post-processing on the triangular mesh model, and outputs a three-dimensional morphological mesh model of the surface fracture.