Structure-appearing damage self-mapping and dynamically updated illumination-embedded temporal neural radiance field method

By using the illumination-embedded temporal neural radiation field method, the problems of manual dependence and illumination influence in the traditional 3D structural model update are solved. This method enables autonomous mapping and dynamic updating of apparent structural damage, and improves the reconstruction accuracy and consistency of the model under complex illumination conditions.

CN120782941BActive Publication Date: 2026-02-10HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510877377.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2026-02-10
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Traditional methods for updating 3D structural models rely heavily on manual intervention, have low levels of automation and intelligence, struggle to cope with complex structural changes and large-scale data processing, have limited accuracy in point cloud registration and texture mapping, and cannot achieve autonomous mapping and dynamic updating of structural surface damage under complex lighting conditions.

Method used

We employ a method of autonomous mapping and dynamic updating of structural apparent damage using illumination embedding temporal neural radiation fields. By integrating implicit reconstruction networks, temporal information, and illumination information modeling modules, we establish a multi-scale adaptive illumination encoder network based on an attention mechanism. We design illumination-constrained loss functions and optimization strategies, and construct an illumination embedding temporal neural radiation field model to achieve autonomous mapping and dynamic updating of structural scenes.

Benefits of technology

This approach aims to improve the stability and accuracy of models under complex lighting conditions, enable autonomous mapping and dynamic updating of structural apparent damage, enhance the temporal consistency and applicability of models, improve rendering quality and efficiency, and reduce the requirements for large datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120782941B_ABST
    Figure CN120782941B_ABST
Patent Text Reader

Abstract

The application proposes a light embedding timing neural radiance field method for structural apparent damage autonomous mapping and dynamic updating. The method can effectively decouple the influence of light, view angle and time change on the appearance of the structure by integrating an implicit reconstruction network, a time information modeling module and a light information modeling module, while maintaining high-quality rendering, improving the stability and accuracy of the model in a variable environment, and providing a new solution for structural scene dynamic updating and apparent damage autonomous mapping.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of structural health monitoring, and particularly relates to a light embedding time sequence neural radiation field method for autonomous mapping and dynamic updating of structural apparent damage. BACKGROUND

[0002] Civil infrastructure structures play an important role in social construction and national economic development in China. With the extension of the service life of structures, they are inevitably affected by multiple factors such as environmental erosion, material aging, cyclic dynamic load and natural disasters, and gradually produce apparent damage such as concrete spalling and cracking. Civil infrastructure structure damage evolves and develops over time, and these damages have typical spatio-temporal evolution characteristics, directly affecting the overall structural safety and long-term stability of the structure, making it difficult for the evaluation and analysis based on the old three-dimensional model to reflect the latest structure state, and thus may lead to safety hazards and economic losses. Therefore, in order to ensure the accuracy and reliability of the three-dimensional model of civil infrastructure structures, it is necessary to regularly update the three-dimensional model of civil infrastructure structures, accurately map the structural damage to the three-dimensional model, and realize dynamic updating of the model and autonomous mapping of the damaged structure.

[0003] Traditional three-dimensional model updating of structures mainly adopts the following steps:

[0004] (1) Use image-based three-dimensional reconstruction tools (such as COLMAP, OpenMVS) to generate an initial three-dimensional model through SfM and MVS technology;

[0005] (2) Collect the latest multi-view images or point cloud data, and use point cloud processing software (such as CloudCompare, MeshLab) to register and analyze the differences between the new and old data;

[0006] (3) Perform local model updating on the identified changed areas in three-dimensional modeling and visualization software (such as Blender), and perform texture mapping and display.

[0007] However, the above traditional methods for mapping damage and dynamically updating the three-dimensional model of structures still have the following shortcomings:

[0008] (1) It relies on a large amount of manual intervention and manual operation, the process is tedious and time-consuming, and the professional skills of the operator are required, which is easily affected by human error.

[0009] (2) It is difficult to deal with complex structural changes and large-scale data processing, and the automation and intelligence level and updating efficiency are low.

[0010] (3) The accuracy of point cloud registration and texture mapping is limited, the model local change recognition ability is weak, and it is easy to lead to incomplete or inaccurate observation of structural changes.

[0011] In summary, the traditional structure three-dimensional model updating and appearance damage mapping method has the problems of high degree of artificial dependence, lack of automation mechanism, limited updating precision and integrity, etc., and it is difficult to meet the actual needs of the structure in the complex lighting conditions and time sequence changes for the autonomous mapping of the appearance damage and the dynamic updating of the three-dimensional model.

[0012] With the rapid development of computer vision and deep learning technology, more and more researches try to apply deep neural networks to three-dimensional reconstruction tasks, giving birth to new methods that integrate deep learning and traditional modeling techniques. In this context, neural radiance fields, as a new emerging implicit scene representation method, take advantage of efficient utilization of multi-view data, optimize the underlying continuous volume scene function, and can achieve high-quality novel view synthesis, showing strong flexibility and real-time rendering capability.

[0013] Although neural radiance fields perform well in detail reconstruction and view independence, their original structure lacks a direct mapping mechanism for the expression of three-dimensional model appearance structure damage, which to some extent limits their application effect in structure damage observation scenarios. To overcome this limitation, some researches try to introduce surface texture mapping, manual editing, etc. to realize the presentation of damaged areas in the appearance of the model. However, such methods still rely on a large amount of manual intervention, requiring manual adjustment of the exported mesh model or voxel model, which not only reduces efficiency but also limits the application in practical engineering.

[0014] In addition, in practical scenarios, the appearance of the structure is not only affected by its own damage, but also significantly disturbed by environmental lighting conditions (such as time changes, weather factors, etc.). Changes in lighting angle, intensity and color temperature can cause significant differences in texture and color of the structure appearance, thereby affecting the accuracy of three-dimensional reconstruction results and damage mapping. Traditional methods often ignore the impact of lighting changes on appearance expression during model updating, resulting in inconsistent appearance effects of the generated model at different times, reducing the reconstruction accuracy and usability of the model.

[0015] To solve the above problems and meet the modeling needs of autonomous mapping and dynamic updating in the context of structure appearance damage changes, the present invention proposes a lighting-embedded time sequence neural radiance field method for autonomous mapping and dynamic updating of structure appearance damage. By integrating implicit reconstruction network, time information and lighting information modeling modules, it can effectively decouple the effects of lighting, view angle and time changes on the appearance of the structure while maintaining high-quality rendering, improving the stability and accuracy of the model in a changing environment, and providing a new solution for dynamic updating of structure scenes and autonomous mapping of appearance damage. SUMMARY

[0016] The application aims to solve the problems in the prior art and proposes a structure appearance damage autonomous mapping and dynamic updating light embedding time sequence neural radiance field method.

[0017] The application is implemented through the following technical solutions, and proposes a structure appearance damage autonomous mapping and dynamic updating light embedding time sequence neural radiance field method, which comprises the following steps:

[0018] Step one: camera pose recovery based on motion recovery structure;

[0019] Input the structure appearance damage image dataset, extract the feature points of each image, find the corresponding relationship of the feature points in different images through feature matching, and calculate the transformation matrix of the pixel coordinates to the world coordinates and the relative pose of the camera;

[0020] Step two: establish a multi-scale adaptive light encoder network architecture based on an attention mechanism;

[0021] Input the structure image with different light angles, intensities and weather conditions, extract the light-related features of each scale in the image through the feature extraction network with ConvNeXt as the backbone, fuse the convolution block attention mechanism, enhance the modeling ability of the network architecture to the key area light information, and output the multi-scale adaptive light embedding representation, thereby providing stable light coding support for subsequent appearance damage recognition and three-dimensional modeling;

[0022] Step three: establish a structure appearance damage autonomous mapping and dynamic updating light embedding time sequence neural radiance field model;

[0023] Input the spatial coordinates of the structure scene sampling points, the light direction vector, the time information and the original image used for extracting the light features, and output the color and body density value of the structure scene sampling points;

[0024] Step four: design a light constraint loss function and optimization strategy based on the volume rendering technology, and train the structure appearance damage autonomous mapping light embedding time sequence neural radiance field model;

[0025] Input the image sequence fused with light coding and time information, based on the volume rendering method, weight and accumulate the sampling point color along the line of sight path, calculate the predicted color value of the view pixel, apply the pixel color loss function based on the light appearance to optimize the model, and explicitly model the influence of light changes on the pixel value; use the time sequence light embedding interval constraint loss function to control the consistency of the image light vector at the same time and the difference between different times, and constrain the time sequence distribution of the light coding;

[0026] Step five: use the trained model to perform new view rendering synthesis on the structure scene.

[0027] After the training is completed, the pixel point color value is predicted by combining the body rendering technology and the color value of each pixel block, and the new view synthesis is completed;

[0028] Step six: structure and appearance damage self-mapping and model updating based on new damage image;

[0029] The damaged structure area or the area where the damage occurs and develops is photographed to obtain a new damage image, and the model for estimating the damage structure scene color and body density is retrained to realize appearance damage self-mapping and model updating.

[0030] Further, the step one specifically comprises:

[0031] Step one: feature extraction of damage structure image by scale invariant feature transform;

[0032] Step two: establish damage structure image matching track, and realize two-by-two matching of damage structure image pairs based on KD tree neighbor search;

[0033] Step three: solve the fundamental matrix by using the homogeneous coordinates of eight pairs of matching feature points;

[0034] Step four: solve the rotation matrix and translation vector to realize coordinate conversion from the image coordinate system to the world coordinate system.

[0035] Further, the step two specifically comprises:

[0036] Step two: establish light feature extraction and normalization module;

[0037] Step two: establish deep light modeling mechanism based on ConvNeXt;

[0038] Step three: design convolution block attention mechanism and feature enhancement module;

[0039] Step four: design light feature integration module of down-sampling and pooling normalization.

[0040] Further, the step three specifically comprises:

[0041] Step three: camera light generation and discrete light layered sampling;

[0042] Step three: establish light embedded structure and appearance damage mapping time sequence neural radiation field network model;

[0043] The network model is composed of three parts: light neural network, time sequence neural network and view neural network, which are respectively used for modeling scene light features, dynamic damage evolution and view-dependent appearance changes; the network model first , time information and light direction is encoded, and the illumination feature obtained after the illumination encoder is encoded is taken as input to model the change pattern of the appearance of the structure under the influence of illumination; wherein the spatial coordinates are used to describe the position of the sampling points in the scene, the time information reflects the state of the structure at a specific time, and the light direction describes the propagation direction of the light in the scene, while the illumination feature is used to represent the ambient lighting conditions of the scene; meanwhile, a volume density and color value prediction decoder embedded with illumination information is designed to improve the reconstruction accuracy and appearance consistency of the neural radiance field under complex lighting conditions.

[0044] Further, the network architecture of the volume density and color value prediction decoder includes:

[0045] (1) a point cloud neural network of the volume density prediction decoder structure;

[0046] (2) a volume density prediction linear layer of the volume density prediction decoder structure;

[0047] (3) a time series neural network of the color value prediction decoder structure;

[0048] (4) a view angle neural network of the color value prediction decoder structure;

[0049] (5) a color prediction linear layer of the color value prediction decoder structure;

[0050] (6) an illumination neural network of the color value prediction decoder structure embedded with illumination;

[0051] (7) a color prediction linear layer of the color value prediction decoder structure embedded with illumination.

[0052] Further, the step four specifically includes:

[0053] Step four one: calculating view pixel values based on volume rendering technology;

[0054] Step four two: optimizing the time series neural radiance field model embedded with illumination based on a pixel color loss function based on illumination appearance;

[0055] Step four three: optimizing the time series neural radiance field model embedded with illumination based on a time series illumination embedding interval constraint loss function.

[0056] Further, the step five specifically includes:

[0057] Step five one: selecting a camera center, setting the number of rays based on the resolution of the image to be rendered, after sampling on the rays, the spatial coordinates of the sampling points , direction vectors , time information As input, input into the structure appearance damage autonomous mapping light embedding time sequence neural radiance field network model, output the color value of each sampling point And the body density value ;

[0058] Step five two: using volume rendering technology to generate structure rendering to synthesize new view pixel color value along the line of sight path, finally complete the color value prediction of the whole new view, realize the synthesis of new view.

[0059] Further, the step six specifically comprises:

[0060] Step six one: photographing the damaged structure area or the area where the damage occurs and develops, obtaining the damage image that does not exist in the existing structure scene data set, and prefixing the damage structure image name with the shooting time information;

[0061] Step six two: restoring the camera pose based on the motion recovery structure and solving the coordinate transformation matrix, and then transforming the camera center coordinates and direction vectors to the real world coordinate system;

[0062] Step six three: adding the new damage structure image file to the structure image data set, thereby incrementally training the structure appearance damage autonomous mapping light embedding time sequence neural radiance field model;

[0063] Step six four: selecting an arbitrary time node, viewpoint and view direction and reference light condition, generating a series of rays, and performing discrete layered sampling on the rays, and using the updated trained network model to predict the color value and body density value of the sampling points;

[0064] Step six five: combining each pixel block with a predicted color value together to generate a new view image using volume rendering technology, thereby realizing autonomous mapping, dynamic updating and new time node structure scene rendering for new damage.

[0065] The application also proposes an electronic device comprising a memory and a processor, the memory stores a computer program, and the processor realizes the steps of the structure appearance damage autonomous mapping and dynamic updating light embedding time sequence neural radiance field method when executing the computer program.

[0066] The application also proposes a computer readable storage medium for storing computer instructions, which are executed by a processor to realize the steps of the structure appearance damage autonomous mapping and dynamic updating light embedding time sequence neural radiance field method.

[0067] The beneficial effects of the present application are as follows:

[0068] 1. The present application realizes image matching and association by adopting damaged structure foreground based on camera pose recovery of motion structure, effectively reducing the requirement for the amount of data set.

[0069] 2. The present application adopts a layered sampling strategy, improves rendering speed, saves computing resources, and improves rendering quality, so that the generated three-dimensional scene is more realistic and high-quality.

[0070] 3. The present application constructs an illumination encoding feature extraction network that fuses a ConvNeXt backbone and a convolution block attention module (CBAM), realizes efficient extraction and expression of illumination features through a multi-level feature modeling mechanism, and designs a body density and color value prediction decoder embedded with illumination information, which is used to improve the reconstruction accuracy and apparent consistency of neural radiance fields under complex illumination conditions.

[0071] 4. The present application proposes a structure and appearance damage self-mapping illumination embedding time sequence neural radiance field model, embeds time variables, models the apparent structure state of the building scene at different time nodes, realizes the self-mapping of structure and appearance damage, and effectively decouples the structure features and appearance damage under complex illumination conditions through explicit modeling of the illumination variable, thereby enhancing the time sequence consistency and applicability of the model to complex illumination scenes.

[0072] 5. The present application proposes a time sequence illumination embedding interval constraint loss function optimization strategy, which narrows the embedding distance of images within the same time sequence and widens the embedding distance between images of different time sequences, constructs a structured illumination embedding space, and improves the expression ability of the model to cross-time illumination changes and multi-view consistency.

[0073] 6. The present application proposes a pixel color loss function optimization strategy based on illumination appearance, which optimizes the color prediction results of the model under different sampling branches through color consistency constraint and illumination condition guided color regression, and improves the rendering realism and robustness in complex illumination scenes.

[0074] 7. The present application realizes structure and appearance damage self-mapping through the design of structure and appearance damage self-mapping and dynamic updating of illumination embedding time sequence neural radiance field, and synchronously completes actual scene rendering and model updating.

[0075] 8. The present application realizes the adaptability and time sequence consistency of the model under complex illumination conditions through the design of structure and appearance damage self-mapping and dynamic updating of illumination embedding time sequence neural radiance field, effectively distinguishes illumination interference and structure damage features, and guarantees the accuracy and modeling stability of damage identification. BRIEF DESCRIPTION OF DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the accompanying drawings in the following description only only the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of the provided drawings.

[0077] Figure 1 is a structural appearance damage autonomous mapping and dynamic updating light embedding time sequence neural radiation field method flow chart.

[0078] Figure 2 is a light embedding module overall architecture schematic diagram based on attention mechanism.

[0079] Figure 3 is a deep light pattern learning mechanism core component network structure schematic diagram.

[0080] Figure 4 is a convolution block attention module network structure schematic diagram.

[0081] Figure 5 is a camera light generation and hierarchical sampling process schematic diagram.

[0082] Figure 6 is a light embedding time sequence neural radiation field network architecture schematic diagram of structural appearance damage autonomous mapping.

[0083] Figure 7 is a point cloud neural network schematic diagram of the volume density prediction decoder structure.

[0084] Figure 8 is a volume density prediction linear layer schematic diagram of the volume density prediction decoder structure.

[0085] Figure 9 is a time sequence neural network schematic diagram of the color value prediction decoder structure.

[0086] Figure 10 is a view neural network schematic diagram of the color value prediction decoder structure.

[0087] Figure 11 is a color prediction linear layer schematic diagram of the color value prediction decoder structure.

[0088] Figure 12 is a light neural network schematic diagram of the color value prediction decoder structure of light embedding.

[0089] Figure 13 is a color prediction linear layer schematic diagram of the color value prediction decoder structure of light embedding. DETAILED DESCRIPTION

[0090] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0091] Specifically, in combination with Figures 1-13 The present application proposes a structure appearance damage autonomous mapping and dynamic updating illumination embedded time sequence neural radiance field method, which comprises the following steps:

[0092] Step one: camera pose recovery based on motion structure recovery;

[0093] Input the structure appearance damage image dataset, extract the feature points of each image, find the corresponding relationship of the feature points in different images through feature matching, and calculate the transformation matrix of pixel coordinates to world coordinates and the relative pose of the camera;

[0094] Step two: establish a multi-scale adaptive illumination encoder network architecture based on attention mechanism;

[0095] Input the structure image with different illumination angles, intensities and weather conditions, extract the illumination-related features of each scale in the image through the feature extraction network with ConvNeXt as the backbone, fuse the convolution block attention mechanism, enhance the modeling ability of the network architecture to the key area illumination information, and output the multi-scale adaptive illumination embedding representation, which provides stable illumination coding support for subsequent appearance damage recognition and three-dimensional modeling;

[0096] Step three: establish a structure appearance damage autonomous mapping and dynamic updating illumination embedded time sequence neural radiance field model;

[0097] Input the spatial coordinates of the structure scene sampling points, the light direction vector, the time information and the original image used to extract the illumination features, and output the color and body density value of the structure scene sampling points;

[0098] Step four: design an illumination constraint loss function and optimization strategy based on volume rendering technology, and train the structure appearance damage autonomous mapping illumination embedded time sequence neural radiance field model;

[0099] Input the image sequence fused with illumination encoding and time information, calculate the predicted color value of the view pixel by weighting and accumulating the color of the sampling point along the line-of-sight path based on the volume rendering method, optimize the model by applying the pixel color loss function based on the appearance of illumination, and explicitly model the influence of illumination change on the pixel value; use the time sequence illumination embedding interval constraint loss function to control the consistency of the image illumination vector at the same time and the difference between different times, and constrain the time sequence distribution of the illumination encoding;

[0100] Step five: using the trained model to render and synthesize new view of the structure scene;

[0101] After training, the pixel color value is predicted by combining the volume rendering technology with the illumination embedding time sequence neural radiance field of the structural appearance damage self-mapping, and the color value of each pixel block is combined to complete the new view synthesis;

[0102] Step six: structural appearance damage self-mapping and model updating based on new damage image;

[0103] The damaged structure area or the area where the damage occurs and develops is photographed to obtain a new damage image, and the model for estimating the color and volume density of the damaged structure scene is retrained to realize the self-mapping of the apparent damage and the model updating.

[0104] The step one specifically comprises:

[0105] Step one: feature extraction of damage structure image through scale invariant feature transform;

[0106] (1) scale space extreme value detection

[0107] For all scale images, a Gaussian difference function is used to calculate and search through the spatial position to identify possible interest points that are insensitive to scale and direction:

[0108] (1)

[0109] wherein, is a Gaussian difference function, is a Gaussian blur function for smoothing the damage structure image to reduce noise and detail information, is a pixel coordinate of the damage structure image, is a standard deviation of the Gaussian blur kernel, indicating the degree of image smoothing, is a scale factor, indicating the proportional relationship between adjacent scales in the scale space.

[0110] (2) key invariant feature point detection

[0111] Taylor expansion is performed on the Gaussian difference function:

[0112] (2)

[0113] Differentiating the above equation and setting the derivative to zero, we obtain the coordinates of the extreme points and the corresponding extreme values:

[0114] (3)

[0115] (3) Key invariant feature point orientation matching

[0116] For each key invariant feature point detected in the scale pyramid, a reference orientation is assigned using local features of the damaged structure image. For each pixel... Considering the image of the Gaussian pyramid in which it is located Calculate the gradient magnitude of pixels within the neighborhood window. and direction angle :

[0117] (4)

[0118]

[0119] In the formula, The scale space representing the key invariant feature points.

[0120] Histograms are used to count the gradient magnitude and direction of all pixels in the neighborhood. The maximum direction angle in the histogram is taken as the principal direction angle of the key invariant feature point, and the maximum direction angle value is assigned to the key point. Key points containing position, scale and direction angle information are taken as invariant feature points of the image.

[0121] Steps 1 and 2: Establish the matching trajectory of the damaged structure image, and perform a nearest neighbor search based on the KD tree to achieve pairwise matching of the damaged structure image pairs;

[0122] Representing an image The set of invariant feature points for image pairs and For any feature point KD-tree is used to calculate the nearest neighbor matching to obtain its nearest neighbor feature point, and its nearest neighbor distance is... Then search for the feature vectors of the second nearest neighbor, whose distance is... If the nearest neighbor distance Distance to next nearest neighbor If the ratio is less than a certain set threshold (such as 0.6), then its nearest neighbor is determined to be an acceptable feature point matching pair.

[0123] Step 13: Solve for the fundamental matrix using the homogeneous coordinates of eight pairs of matching feature points;

[0124] (5)

[0125] wherein, F is the fundamental matrix between two adjacent views and , representing the geometric transformation relationship between the corresponding points of the images, is the pixel coordinate of a certain feature point of the image . is the pixel coordinate of a certain feature point of the image .

[0126] Step 4: Solve the rotation matrix and translation vector to realize the coordinate conversion from the image coordinate system to the world coordinate system.

[0127] (1) Calculate the intrinsic matrix based on the camera intrinsic matrix and the fundamental matrix E :

[0128] (6)

[0129] wherein, K represents the camera intrinsic matrix.

[0130] (2) Perform SVD feature decomposition on the intrinsic matrix to obtain the rotation matrix and the translation vector :

[0131] (7)

[0132] wherein, and correspond to the feature matrix obtained by decomposing the intrinsic matrix, is a diagonal matrix, the elements on the diagonal are eigenvalues, and the translation vector is the third column of U.

[0133] (3) Based on the rotation matrix and the translation vector, realize the pixel-space coordinate conversion from the image coordinate system to the world coordinate system:

[0134] (8)

[0135] wherein, is the coordinate conversion matrix from the camera coordinate system to the world coordinate system.

[0136] The step two constructs a multi-scale adaptive light encoding network based on attention mechanism, aiming to effectively extract the light features in the scene, construct a unified and robust light representation, and thus improve the adaptability and generalization performance of neural radiance field under complex lighting conditions. The network realizes efficient and stable light feature learning by constructing a multi-level feature extraction framework combined with convolution block attention mechanism and global light encoding. The overall structure of the light encoder is as shown in Figure 2 The specific steps are as follows:

[0137] Step two one: establish a light feature extraction and normalization module;

[0138] At the input end, the preprocessed and normalized image is first subjected to convolution calculation to construct the basic light features. The standard two-dimensional convolution calculation method is as follows:

[0139] (9)

[0140] In the formula, represents the value of the output feature at pixel coordinates and channel , is the weight of the corresponding convolution kernel, is the bias term.

[0141] To ensure the stability of the features when calculating in the deep layer of the network and reduce the interference caused by the inconsistent distribution between different channels, the research designs layer normalization (LN) to standardize the output features after convolution operation. The layer normalization calculation formula is as follows:

[0142] (10)

[0143] In the formula, is the feature value of the th pixel after unfolding in all output channels, and are the mean and variance in the channel dimension, and are the learnable scaling and translation parameters, is a numerically stable constant, taking a very small positive number, is the index of the pixel within the batch, ranging from , is the total number of pixels within the batch.

[0144] Step two two: establish a deep light modeling mechanism based on ConvNeXt;

[0145] ConvNeXt Block is the core component of the deep light modeling mechanism, and its structure is as shown in Figure 3As shown, it is mainly composed of depthwise separable convolution, layer normalization, pointwise convolution (Pointwise Conv2D), GELU activation, Layer Scale mechanism, Drop Path mechanism and residual connection, etc. These modules can effectively extract and fuse illumination information of different scales, and enhance the illumination modeling capability of neural radiance field in complex environment.

[0146] Firstly, depthwise separable convolution is used to reduce computational complexity while enhancing local feature modeling capability. This module first performs spatial convolution (Depthwise Convolution) on each channel respectively, and then fuses channel information through 1×1 pointwise convolution (Pointwise Convolution). The calculation method is as follows:

[0147] (11)

[0148] In the formula, represents the output of depthwise convolution, represents the output of pointwise convolution, and are the weights of two-stage convolution kernel.

[0149] Traditional CNN structure usually uses batch normalization (Batch Normalization, BN), but BN may lead to unstable feature distribution when small batch training or illumination condition distribution changes greatly. Therefore, Layer Normalization (LN) is adopted, which directly normalizes the features of each sample, thereby improving the generalization ability of the model under different illumination conditions. The calculation formula of LN is as follows:

[0150] (12)

[0151] In the formula, and are the mean and standard deviation of the input features, and are trainable parameters. LN enables the network to stably handle changes in illumination intensity, different shadow distributions, etc.

[0152] After local feature extraction, it is necessary to further enhance the feature fusion capability between channels. The present application adopts 1×1 pointwise convolution (Pointwise Convolution) to integrate illumination information of different channels, and combines with GELU (Gaussian Error Linear Unit) activation function. GELU replaces the traditional ReLU, and its calculation formula is as follows:

[0153] (13)

[0154] wherein, is the cumulative distribution function of the standard normal distribution. Unlike ReLU, GELU can still provide smooth gradient changes when the input value is small, so it can more accurately fit the light feature in the scene where the light changes smoothly (such as the gradual change of sunlight intensity).

[0155] In addition, ConvNeXt Block adopts Layer Scale mechanism, which can make the features of different layers adaptively adjusted through a learnable scaling factor. The calculation method is as follows:

[0156] (14)

[0157] wherein, is a learnable scaling parameter, which enables the network to flexibly control the intensity of the light feature at different levels.

[0158] ConvNeXt Block adopts residual connection, that is, the input features are directly added to the output features after being processed by multiple convolution layers:

[0159] (15)

[0160] Step two and three: design convolution block attention mechanism and feature enhancement module;

[0161] In the process of light feature extraction, the convolution block attention module (CBAM, Convolutional Block Attention Module) is introduced to respectively enhance the channel dimension and spatial dimension of the feature map, so as to improve the feature selectivity and representation quality. The module is composed of channel attention submodule and spatial attention submodule, and the overall network structure is as shown in Figure 4

[0162] (1) Channel attention module (CAM)

[0163] First, perform (Global Average Pooling, GAP) and global maximum pooling (Global Max Pooling, GMP) on the input feature map to obtain two channel description vectors and , wherein:

[0164] (16) ​

[0165] wherein, denotes the input feature map, with size ; is the number of channels, is the height of the feature map, is the width of the feature map; and are the indices of height and width, respectively, with value range and ; is the feature value at position ; denotes taking the maximum value in the height and width directions.

[0166] Then, and are mapped by a shared-weight multi-layer perceptron (MLP) respectively, and are added together and applied with an activation function to obtain the channel attention map :

[0167] (17)

[0168] wherein denotes an activation function to normalize the output to the interval ; the MLP is usually composed of two fully connected operations to learn the nonlinear mapping of channel weights; is the obtained channel attention map, which serves to weight the input feature map (typically the same as ) in the channel dimension.

[0169] Finally, the channel attention map is multiplied element-wise with the original feature map to obtain the weighted result:

[0170] (18)

[0171] wherein, denotes element-wise multiplication operation; is the weighted feature map.

[0172] (2) Spatial attention module (SAM)

[0173] First, global average pooling and global maximum pooling are performed on the channel dimension to obtain two two-dimensional feature maps and :

[0174] (19)

[0175] wherein​​​ denotes the feature value of the th channel at position ; in which denotes taking the maximum value over all channels.

[0176] Subsequently, the two feature maps are spliced in the channel dimension, and a 7x7 convolution kernel is used to extract the spatial attention weight :

[0177] (20)

[0178] in which, denotes the convolution operation using a 7x7 convolution kernel; denotes the channel dimension splicing operation. Finally, the spatial attention map is multiplied element-wise with the feature map : is the extracted spatial attention map, whose size is the same as the spatial dimension of .

[0179] (21)

[0180] in which, denotes element-wise multiplication; is the feature map after spatial attention weighting.

[0181] Step two four: design the down-sampling and pooling normalization integrated illumination feature integration module.

[0182] After completing the multi-level illumination feature extraction, the network performs two-step operations of down-sampling and global average pooling (GAP) on the feature map, and combines normalization processing to improve the numerical stability and generalization ability of the features. Down-sampling is composed of Layer Norm and Conv2D two parts, which realizes the reduction of the size of the feature map by setting the stride greater than 1 in the convolution operation.

[0183] (22)

[0184] in which, denotes the feature value of the feature map after down-sampling at spatial position channel ; denotes the value of the input feature map at position channel ; denotes the convolution kernel at position from the input channel v to the output channel weight parameters of the convolutional layer; bias term of the output channel bias term of the output channel size of the convolution kernel stride of the convolution operation denotes the layer normalization operation on the input feature, which is used to stabilize the numerical distribution denote the horizontal and vertical coordinate indices of the down-sampled feature map, respectively.

[0185] Finally, the spatial feature is integrated by using the global average pooling (GAP) to form a compact illumination vector The GAP calculation method is as follows:

[0186] (23)

[0187] where denotes the illumination vector component on the channel after the global average pooling; are the height and width of the input feature map, respectively; is the down-sampling stride; therefore, the size of the pooling region is ; denotes the value of the down-sampled feature map on the channel at the position .

[0188] To ensure that the illumination vector remains numerically stable in different scenes or training batches, a layer normalization operation is introduced at the end, and its calculation method can be represented as:

[0189] ;

[0190] (24)

[0191] where is the normalized illumination vector component; is the original illumination vector component; is the original illumination vector component; is the variance of the channel; and are learnable parameters, is a stabilization factor, which is taken as ; is the total number of channels of the illumination vector.

[0192] Step three establishes the structure apparent damage autonomous mapping illumination embedding time series neural radiation field model. Input structure damage scene sampling point space coordinates​​ , light direction vector , time information and images of different illumination features, output the color and volume density values of the sampling points of the damaged structure scene, specifically including:

[0193] Step three: camera light generation and discrete light hierarchical sampling;

[0194] (1) Calculate the normalized direction vector

[0195] For the damaged structure image with time information , subtract the optical center coordinates from the pixel coordinates , take the focal length as the axis coordinate, obtain the absolute direction vector , and perform normalization to obtain the normalized direction vector .

[0196] (2) Camera light generation

[0197] Based on the transformation matrix obtained in step one c2w , transform the camera center coordinates and normalized direction vector to the world coordinate system, extract the time information from the damaged structure image name, and form a light with time information , where o represents the optical center, s represents the distance along the light direction, represents the normalized light direction vector.

[0198] (3) Discrete light hierarchical sampling

[0199] A hierarchical sampling (Hierarchical Sampling) strategy is adopted, which is a two-stage sampling mechanism, namely coarse sampling (Coarse Sampling) and fine sampling, to perform higher density sampling in important areas along the light.

[0200] First, coarse sampling is performed within the range of the light , and the goal of this stage is to obtain a rough distribution of the ray in the entire space. Assuming that the parameter along the light is within the interval , it can be divided into equidistant points, that is:

[0201] (25)

[0202] These sampling points constitute the uniform sampling result in the light range. Subsequently, they are input into a neural radiance field model for structure damage autonomous mapping to predict the volume density of each sampling point . Through coarse sampling, the color distribution of the light in space can be preliminarily estimated, and the contribution weight of each sampling point to the final color is calculated. To solve the above problems, the present application further introduces a fine sampling strategy based on importance weight, so that the sampling density of the light in the key area is higher.

[0203] In the fine sampling stage, it is necessary to first calculate the contribution weight of each sampling point. In the coarse sampling stage, the volume density of each sampling point represents the color contribution degree of the point. Based on the volume rendering theory, the transparency of the point and its contribution to the final color can be calculated, wherein The calculation formula of is as follows:

[0204] (26)

[0205] In the formula, d is the interval length between adjacent sampling points.

[0206] Based on these values, the normalized weight of each sampling point can be calculated:

[0207] (27)

[0208] These weights reflect which area of the sampling point has a greater impact on the final rendering result, and the greater the weight of the area, the more intense the color change at the place, or the point is located near the surface of the object. Based on the calculated weight, the present application adopts an inverse transform sampling (Inverse Transform Sampling) method to perform fine sampling in the high-contribution area. Specifically, first, the cumulative distribution function (CDF) of these weights is calculated:

[0209] (28)

[0210] Then, a group of random numbers is uniformly sampled between 0 and 1, and the inverse transform method is used to obtain new sampling points :

[0211] (29)

[0212] In this way, a group of orthogonal points and the spatial coordinates of the coordinate points are obtained ​​​and normalizing the direction vector .

[0213] The camera ray generation and stratified sampling process is shown in Figure 5 , which demonstrates how to convert pixel coordinates in camera space to rays in world coordinates, and how to obtain more sampling points in high contribution areas through inverse transform sampling method in the coarse sampling and fine sampling stages.

[0214] Figure 5 The pixel coordinates , the optical center coordinates , the focal length parameters , the direction vector and the starting point o and direction vector of the light ray are marked in the figure, and s represents the distance along the direction of the light ray. The distribution of sampling points , , , , at different positions on the light ray is also shown.

[0215] Step three two: establish a time sequence neural radiance field network model for appearance damage mapping of embedded lighting structure;

[0216] The network model consists of three parts: a lighting neural network, a time sequence neural network and a view angle neural network, which are used to model the scene lighting features, dynamic damage evolution and view angle related appearance changes respectively; the lighting neural network extracts environmental lighting features through a lighting encoding module, and combines other encoding information to reduce the interference of lighting changes on appearance prediction, making the rendering result more stable, thereby improving the accuracy and temporal consistency of damage mapping, and the overall network architecture is shown in Figure 6 . The network model first encodes the spatial coordinates , time information and light direction , and takes the lighting features obtained after encoding by the lighting encoder as input to model the change pattern of the structure appearance affected by lighting; among them, the spatial coordinates are used to describe the position of the sampling point in the scene, the time information reflects the state of the structure at a certain time, the light direction describes the propagation direction of the light in the scene, and the lighting features are used to represent the environmental lighting conditions of the scene; meanwhile, the body density and color value prediction decoder embedded with lighting information is designed to improve the reconstruction accuracy and appearance consistency of the neural radiance field under complex lighting conditions.

[0217] The encoder and encoding process used in step 3.2 are as follows:

[0218] (1) Position encoder

[0219] The method described in this invention employs positional coding technology to map the original low-dimensional coordinates to a high-dimensional feature space, thereby enriching the spectral characteristics of the input information and enhancing the network's ability to express high-frequency components. This coding strategy can be viewed as a discretized approximation of the Fourier transform of the input signal. Fourier analysis shows that any function can be decomposed into a linear combination of a series of sine and cosine functions, and positional coding utilizes this property to construct a multi-scale orthogonal basis.

[0220] Specifically, for a single scalar input Position encoding works by constructing a series of sine and cosine functions. The mapping is performed, and the calculation formula is as follows:

[0221] (30)

[0222] In the formula, Can be the coordinates of a point in space or light direction vector ,in The selected frequency order and preset scalar hyperparameters. Position encoding parameters (such as the number of frequency bands). The selection of () has a significant impact on the model's performance. A larger () Smaller values ​​can help capture higher-frequency changes, but may also lead to more parameters, increased training difficulty, and the risk of overfitting; smaller values... This may not be sufficient to express the subtle changes in the scene. We set the value to 10 in order to reduce computational complexity while maintaining expressive power.

[0223] (2) Timing encoder

[0224] In the dataset described in this invention, each image filename contains a date prefix, such as "2025-01-20," to identify the image's capture time. To uniformly encode time across the network, the time information of all images must first be parsed into days or date differences, and then normalized. Specifically, the date with the smallest time span in the dataset is selected as the reference starting point. Simultaneously, record the date with the largest time span in the dataset as... Convert the image's capture date (e.g., 2025-01-20) to the number of days it differs from the reference starting point, and then use... The difference was normalized to Interval. In this way, each image in the dataset corresponds to an interval between... and the normalized time value between , which can achieve consistent input in subsequent time series encoding.

[0225] After obtaining the normalized time value , the method of the present application introduces a time series encoder based on Laplace cumulative distribution function, which performs deeper mapping to enhance the response ability of the model to mutation time events. The encoder first performs standardization processing on by a set of learnable "center position" parameters and "transition steepness" parameters , so that the input automatically adapts to the distribution characteristics of different time periods in the scene during training. Subsequently, according to the size relationship with each , the time axis is divided into two segments, and a double exponential form is used to construct a segmented function to respectively depict the rapid growth and rapid decay phenomena on the left and right of the center point, and the specific formula is as follows:

[0226] (31)

[0227] In the formula, are learnable parameters, whose value range is , the present application randomly initializes to 0.5, and randomly initializes to 0.4. And in training, it is updated together with the network weight.

[0228] The function is continuously derivable at , and by controlling the exponential steepness, it maintains a large gradient near the mutation point, and quickly saturates away from the center, thereby capturing the key mutation moment while suppressing irrelevant disturbances in long periods.

[0229] In actual implementation, first generate and through learnable parameters, and then perform standardization processing on the normalized time of the input, and then according to the size relationship with , it is divided into two cases, and the exponential increase or exponential decay form is used to calculate the output value. When is small (i.e. is near ), the gradient of the encoding function is at a high level, allowing the network to accurately locate and learn the time points in the scene that have undergone mutation; when is large (i.e. far from the center position), the function output quickly tends to saturation, thereby generally depicting long-term or large-scale time differences. Such a design not only ensures high sensitivity of the model to the time of damage, but also suppresses irrelevant information in the stationary period, thereby providing effective support for modeling of multi-period structural damage changes. Through vectorization design, multiple groups Parallelly encode the same so that the model can simultaneously and explicitly model multiple time mutations that can occur in the scene without manually presetting the number of turning points. Such a multi-layer time coding strategy not only unifies the time representation of cross-period images, but also dynamically adjusts the sensitivity to damage occurrence with the aid of learnable parameters, thereby improving the ability to capture structural apparent damage changes and providing support for time consistency modeling and damage evolution analysis.

[0230] The network architecture of the volume density and color value prediction decoder includes:

[0231] (1) a point cloud neural network of the volume density prediction decoder structure;

[0232] The point cloud neural network of the volume density prediction module is composed of multiple linear layers, and the activation function is a rectified linear unit (ReLU). The input dimension of the first linear layer is input_ch_point, and the output dimension is W; after that, the input dimension of the linear layer after the skip connection layer is input_ch+W, and the input dimension and the output dimension of the other linear layers are both W. M and N are the number of cycles, which can be 3 or 4, Figure 7 A schematic diagram of the point cloud neural network of the volume density prediction decoder structure.

[0233] (2) a volume density prediction linear layer of the volume density prediction decoder structure;

[0234] The input dimension of the volume density prediction linear layer of the volume density prediction module is W, the output dimension is 1, and the activation function is ReLU. Figure 8 A schematic diagram of the volume density prediction linear layer of the volume density prediction decoder structure.

[0235] (3) a time neural network of the color value prediction decoder structure;

[0236] The time neural network of the time information embedding module is composed of multiple linear layers, and the activation function is ReLU. The input dimension of the first linear layer is input_ch_time=spatial coordinate The dimension after the frequency encoder, the time information The sum of the dimension after the exponential centering encoder and the dimension of the spatial feature vector h is W / 2, and the output dimension is W / 2; the input dimension and the output dimension of the linear layer after that are both W / 2, and the output dimension of the last linear layer is W. K is the number of cycles, which can be 2, Figure 9A timing neural network diagram of the color value prediction decoder structure.

[0237] (4) A view neural network of the color value prediction decoder structure;

[0238] The view neural network of the color value prediction module is composed of multiple linear layers, and the activation function is ReLU. The input dimension of the first linear layer is input_ch_view = normalized direction vector The sum of the dimension after the frequency encoder and the dimension of the spatio-temporal feature vector is W / 2; the input dimension and the output dimension of the subsequent linear layer are both W / 2. Figure 10 A view neural network architecture diagram of the color value prediction module.

[0239] (5) A color prediction linear layer of the color value prediction decoder structure;

[0240] The input dimension of the color prediction linear layer of the color value prediction module is W / 2, and the output dimension is 3, using the Sigmoid activation function. Figure 11 A color prediction linear layer architecture diagram of the color value prediction module.

[0241] (6) An illumination neural network of the color value prediction decoder structure with illumination embedding;

[0242] The input dimension of the illumination neural network of the color value prediction decoder structure with illumination embedding is li_emb, and the output dimension is W / 2; the input dimension and the output dimension of the subsequent linear layer are both W / 2. Figure 12 An illumination neural network diagram of the color value prediction decoder structure with illumination embedding.

[0243] (7) A color prediction linear layer of the color value prediction decoder structure with illumination embedding.

[0244] The input dimension of the color prediction linear layer is W / 2, and the output dimension is 3, using the Sigmoid activation function. This linear layer maps the illumination feature to the output space of the three channels of RGB and normalizes it in the range of [0, 1]. Figure 13 A color prediction linear layer diagram of the color value prediction decoder structure with illumination embedding.

[0245] The step four specifically includes:

[0246] Step four one: view pixel value calculation based on volume rendering technology;

[0247] First, the color contribution of each sampling point is weighted and summed along the line-of-sight path to generate the predicted damaged structure pixel color value. The volume rendering weighted sum calculation method is:

[0248] (32)

[0249] wherein, is the color prediction value of the image pixel point, represents generating a light ray, is the color value at the point on the light ray, s is the volume density value at the point on the light ray, is the color value at the point on the light ray, s is the volume density value at the point on the light ray, and respectively represent the nearest boundary and the farthest boundary of the generated light ray; is the transparency function from the viewpoint to the space point, the greater the volume density, the smaller the transparency, indicating that more light rays are absorbed, calculated by integrating the volume density of all points on the path from the camera to the current space point in a negative exponential:

[0250] (33)

[0251] After the discretization processing of the volume rendering formula, the is divided into N uniform intervals, and a sample is randomly extracted therefrom. The calculation method of the discretized volume rendering is:

[0252] ;

[0253] ;

[0254] ; (34)

[0255] wherein, represents the s value corresponding to the i th sampling point on the light ray after discretization, indicates uniform distribution, represents the distance between adjacent sampling points, is the transparency function from the viewpoint to the i th sampling point, is the volume density prediction value of the i th sampling point, is the color prediction value of the i th sampling point, represents the color prediction value of the image pixel point after the discretized volume rendering.

[0256] Step four two: optimizing the light embedding time series neural radiance field model through a pixel color loss function based on the appearance of light;

[0257] The loss function adopts L2 norm to measure the error between the pixel color prediction value and the true value, and its mathematical expression is:

[0258] (35)

[0259] wherein, is the spatio-temporal set of all sampled pixels from the injured structure image dataset; is the known pixel color value on the injured structure image; is the predicted pixel color value at the th time step; is the L2-norm, i.e., the Euclidean distance.

[0260] To better utilize the hierarchical sampling strategy, the loss function not only constrains the color value prediction generated by the coarse sampling branch, but also supervises the prediction of the fine sampling branch.

[0261] Specifically, the loss function contains two parts. First, it is the color consistency constraint on the prediction results of the coarse sampling and fine sampling branches, and its expression is:

[0262] (36)

[0263] wherein, is the color prediction value of the coarse sampling point, is the color prediction value of the fine sampling point.

[0264] Second, the second loss function of light is introduced to make up for the deficiency of simply based on color consistency constraint in the scene of light change, and further improve the fitting ability of the model to the light effect in the actual environment. By taking the light condition as an input variable, this loss function guides the model to consider the influence of light on color generation during the sampling process, so as to more realistically reproduce the shadow, highlight and brightness change and other light characteristics in the rendered image. The specific expression is:

[0265] (37)

[0266] wherein, denotes the sampling ray, is the sampling point, denotes the set of all sampling points; is the true color value at the pixel under a certain light condition; and denotes the color value predicted by the model under the sampling point and the light condition .

[0267] Finally, the total loss function is constructed in a weighted combination way as:

[0268] (38)

[0269] wherein, The balance parameter is used to adjust the relative contribution of the color consistency constraint and the color constraint based on the appearance of the light in the overall training process. This design not only ensures the consistency of color prediction at different sampling levels, but also maintains the visual consistency of the generated image when the lighting condition changes, thereby realizing more accurate and robust three-dimensional reconstruction of the structural neural radiance field model.

[0270] Step four three: optimizing the time series neural radiance field model of the light embedding by the time series light embedding interval constraint loss function.

[0271] To improve the modeling accuracy of the appearance change of the structure in the long-term observation image in the three-dimensional reconstruction process, a time series-based light appearance embedding loss function is proposed. The loss function introduces a contrast constraint of time series when optimizing the light appearance coding, which narrows the distance between images in the embedding space within the same time series, enhances the expression consistency of similar light features, and at the same time, widens the embedding distance between images at different time series, improves the discrimination ability of the light difference across time. This method helps to effectively separate the static geometric information and dynamic appearance change in the time series, maintain the appearance consistency within the same time series, and improve the stability and expression ability of the structure appearance damage modeling in the long-term observation scene.

[0272] During the training process, the light coding of an image at a certain time is , the light coding of another image at the same time is the positive sample , and the light coding at other time points is the negative sample . To realize "compact within the time series and separation between the time series", an interval constraint function is defined:

[0273] (39)

[0274] In the formula, is the cosine similarity, and the indicator function is used to filter negative samples at different time periods. That is, when the sample is different from the sample in time, it is 1, otherwise it is 0. This loss function makes the light coding at the same time point closer in the feature space, and the coding at different time points is pulled apart, thereby strengthening the time series discrimination ability of the light feature.

[0275] The step five specifically includes:

[0276] Step five one: selecting the camera center, setting the number of light rays based on the resolution of the image to be rendered, sampling on the light rays, and then obtaining the spatial coordinates , the direction vector , and the time information As input, the light embedding time sequence neural radiance field network model of structural appearance damage autonomous mapping is input, and the color value of each sampling point is output And the body density value ;

[0277] Step five two: using volume rendering technology to generate structural rendering to synthesize new view pixel color value along the line of sight path, finally complete the color value prediction of the whole new view, realize the synthesis of new view.

[0278] The step six specifically comprises:

[0279] Step six one: photograph the damaged structure area or the area where the damage occurs and develops, obtain the damage image that does not exist in the existing structure scene data set, add the shooting time information prefix to the name of the damage structure image;

[0280] Step six two: restore the camera pose based on the motion recovery structure and solve the coordinate transformation matrix, and then transform the camera center coordinates and direction vector to the real world coordinate system;

[0281] Step six three: add the new damage structure image file to the structure image data set, thereby carrying out the incremental training of the light embedding time sequence neural radiance field model of structural appearance damage autonomous mapping;

[0282] Step six four: select any time node, viewpoint and view direction and reference light condition, generate a series of light rays, and carry out discrete layered sampling on the light rays, and use the updated trained network model to predict the color value and the body density value of the sampling points;

[0283] Step six five: combining the volume rendering technology, combining each predicted color value pixel block to generate a new view image, thereby realizing the autonomous mapping, dynamic updating and new time node structure scene rendering for the new damage.

[0284] The application also proposes an electronic device comprising a memory and a processor, the memory stores a computer program, and the processor realizes the steps of the light embedding time sequence neural radiance field method of structural appearance damage autonomous mapping and dynamic updating when executing the computer program.

[0285] The application also proposes a computer readable storage medium for storing computer instructions, which realizes the steps of the light embedding time sequence neural radiance field method of structural appearance damage autonomous mapping and dynamic updating when the processor executes the computer instructions.

[0286] The memory in the embodiments of the application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a read only memory (read only memory, ROM), a programmable read only memory (programmable ROM, PROM), an erasable programmable read only memory (erasable PROM, EPROM), an electrically erasable programmable read only memory (electrically EPROM, EEPROM) or a flash memory. The volatile memory can be a random access memory (random access memory, RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (static RAM, SRAM), dynamic random access memory (dynamic RAM, DRAM), synchronous dynamic random access memory (synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (double data rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (synchlink DRAM, SLDRAM) and direct memory bus random access memory (direct rambus RAM, DR RAM). It should be noted that the memory of the method described in the application is intended to include but not limited to these and any other suitable type of memory.

[0287] In the above embodiments, all or part of the methods can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the methods can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available media can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a high-density digital video disc (DVD)), or a semiconductor medium (such as a solid state disc (SSD)), etc.

[0288] In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor or the instruction in the form of software. The steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware processor execution or executed by a combination of hardware and software modules in the processor. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0289] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with a signal processing capability. In the implementation process, each step of the method embodiments can be completed by the integrated logic circuit of hardware or the instruction in the form of software in the processor. The processor mentioned above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the memory, and the processor reads the information in the memory, and combines the hardware to complete the steps of the above method.

[0290] The structure of the method of the present application is introduced in detail above. The principle and implementation mode of the present application are described by applying specific examples in this paper. The above embodiment is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed, and the above description should not be understood as the limitation of the present application.

Claims

1. A method for autonomous mapping and dynamically updated illumination-embedded temporal neural radiation fields of structural apparent damage, characterized in that, The method includes the following steps: Step 1: Camera pose recovery based on motion reconstruction structures; Input a dataset of images showing structural and apparent damage, extract feature points from each image, find the correspondence between feature points in different images through feature matching, and calculate the transformation matrix from pixel coordinates to world coordinates and the camera relative pose. Step 2: Establish a multi-scale adaptive illumination encoder network architecture based on an attention mechanism; The system takes structural images with different illumination angles, intensities, and weather conditions as input, and extracts illumination-related features at various scales in the images through a feature extraction network with ConvNeXt as the backbone. It also integrates a convolutional block attention mechanism to enhance the network architecture's ability to model illumination information in key areas and outputs a multi-scale adaptive illumination embedding representation, providing stable illumination coding support for subsequent appearance damage identification and 3D modeling. Step 3: Establish an autonomous mapping and dynamically updated illumination-embedded temporal neural radiation field model for structural apparent damage; Input the spatial coordinates of the sampling points in the structural scene, the ray direction vector, the temporal information, and the original image used to extract illumination features; output the color and volume density values ​​of the sampling points in the structural scene. Step 4: Design a lighting constraint loss function and optimization strategy based on volume rendering technology, and train a lighting embedding temporal neural radiation field model for autonomous mapping of structural appearance damage; Input an image sequence that integrates illumination encoding and temporal information. Based on a volume rendering method, the colors of sampling points are weighted and accumulated along the view path to calculate the predicted color value of the view pixel. The model is optimized by applying a pixel color loss function based on illumination appearance to explicitly model the impact of illumination changes on pixel values. The temporal series illumination embedding interval constraint loss function is used to control the consistency of the image illumination vector at the same time and the differences between different times, thus constraining the temporal distribution of illumination encoding. Step 5: Use the trained model to render and synthesize a new view of the structural scene; After training, the illumination embedding temporal neural radiation field of the structure appearance damage autonomous mapping is used, combined with volume rendering technology to predict the color value of the pixel, and the color values ​​of each pixel block are combined to complete the synthesis of the new perspective. Step 6: Perform autonomous mapping of structural apparent damage and model update based on the new damage image; The system captures images of damaged structural regions or regions where damage is developing, acquires new damage images, and retrains the models for estimating the color and volume density of the damaged structural scenes to achieve autopilot mapping and model updating of apparent damage.

2. The method according to claim 1, characterized in that, Step one specifically includes: Step 11: Extract features from damaged structure images using scale-invariant feature transformation; Steps 1 and 2: Establish the matching trajectory of the damaged structure image, and perform a nearest neighbor search based on the KD tree to achieve pairwise matching of the damaged structure image pairs; Step 13: Solve for the fundamental matrix using the homogeneous coordinates of eight pairs of matching feature points; Step 14: Solve for the rotation matrix and translation vector to achieve coordinate transformation from the image coordinate system to the world coordinate system.

3. The method according to claim 1, characterized in that, Step two specifically includes: Step 21: Establish a module for light feature extraction and normalization; Step 22: Establish a deep lighting modeling mechanism based on ConvNeXt; Steps two and three: Design the convolutional block attention mechanism and feature enhancement module; Step 24: Design a module to integrate downsampling and pooling normalization of illumination features.

4. The method according to claim 1, characterized in that, Step three specifically includes: Step 31: Camera ray generation and discrete ray layer sampling; Step 32: Establish a temporal neural radiation field network model for mapping apparent damage in an illuminated embedded structure; The network model consists of three parts: an illumination neural network, a temporal neural network, and a viewpoint neural network, which are used to model scene illumination features, dynamic damage evolution, and viewpoint-related appearance changes, respectively. The network model first processes spatial coordinates... Time information and the direction of light Encode the lighting features obtained after the lighting encoder is used. As input, model the variation pattern of the structural appearance under the influence of lighting; where spatial coordinates Used to describe the location and timing information of sampling points in the scene. Reflecting the state of the structure at a specific moment, the direction of light Describes the direction of light propagation in a scene, while lighting features This is used to represent the ambient lighting conditions of the scene; at the same time, a volume density and color value prediction decoder with embedded lighting information is designed to improve the reconstruction accuracy and appearance consistency of the neural radiation field under complex lighting conditions.

5. The method according to claim 4, characterized in that, The network architecture of the volume density and color value prediction decoder includes: (1) A point cloud neural network with a volume density prediction decoder structure; (2) Volume density prediction linear layer of the volume density prediction decoder structure; (3) Temporal neural network for color value prediction decoder structure; (4) A viewpoint neural network for the color value prediction decoder structure; (5) Color prediction linear layer of the color value prediction decoder structure; (6) Illumination neural network with illumination-embedded color value prediction decoder structure; (7) Color prediction linear layer of the illumination-embedded color value prediction decoder structure.

6. The method according to claim 1, characterized in that, Step four specifically includes: Step 41: Calculate view pixel values ​​based on volume rendering technology; Step 42: Optimize the temporal neural radiation field model of illumination embedding using a pixel color loss function based on illumination appearance; Step 43: Optimize the temporal neural radiation field model of illumination embedding using the time-series illumination embedding interval constraint loss function.

7. The method according to claim 1, characterized in that, Step five specifically includes: Step 51: Select the camera center, set the number of rays based on the resolution of the image to be rendered, sample the rays, and then record the spatial coordinates of the sampling points. Direction vector Time information As input, the light-embedded temporal neural radiation field network model of the self-mapping of structural apparent damage is fed into the model, and the color value of each sampling point is output. and volume density value ; Step 52: Using volume rendering technology, generate pixel color values ​​for the structural rendering composite new view along the view path, and finally complete the color value prediction of the entire new view to achieve the compositing of the new view.

8. The method according to claim 1, characterized in that, Step six specifically includes: Step 61: Take pictures of the damaged structural area or the area where damage is developing, and obtain damage images that do not exist in the existing structural scene dataset. Add the shooting time information as a prefix to the name of the damaged structural image. Step 62: Recover the camera pose based on the motion recovery structure and solve the coordinate transformation matrix. Then transform the camera center coordinates and orientation vector to the real-world coordinate system. Step 63: Add the new damaged structure image files to the structure image dataset to perform incremental training of the illumination embedding temporal neural radiation field model for autonomous mapping of structural apparent damage. Step 64: Select any time point, viewpoint, viewing direction, and reference lighting conditions to generate a series of rays, and perform discrete layered sampling on the rays. Use the updated and trained network model to predict the color value and volume density value of the sampling points. Step 65: Combining volumetric rendering technology, each pixel block with predicted color value is combined to generate a new perspective image, thereby achieving autonomous mapping, dynamic updating, and rendering of new time-node structural scenes for newly emerging damage.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-8.

10. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Intelligent three-dimensional structure reconstruction method based on three-plane feature representation and visual angle condition diffusion model

    CN117893691A

  • Neural radiation field-based structure three-dimensional model updating method, apparatus and device, and medium

    CN118365817A