A pipeline backfill settlement monitoring method and system based on image recognition
By fusing multimodal data from UAV aerial imagery and SAR remote sensing images and extracting features using a cross-attention mechanism, the problem of insufficient accuracy in pipeline backfill settlement monitoring was solved, achieving automated settlement and compaction monitoring and improving the accuracy and convenience of monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies for settlement monitoring in pipeline backfill areas lack the ability to integrate and analyze multi-source heterogeneous data, making it difficult to achieve simultaneous inversion of three-dimensional settlement and compaction. Furthermore, traditional models are difficult to adapt to complex terrain and dynamic environmental changes, resulting in insufficient monitoring accuracy and poor convenience.
An image recognition-based approach was adopted, which fused multimodal data from UAV aerial images and SAR remote sensing images. The cross-attention mechanism was used to extract elevation changes, ground texture, color features and topographic deformation features. Combined with a multi-task output head, the settlement progress and compaction degree were predicted simultaneously. A dual-stream multimodal feature fusion architecture was designed for automated monitoring.
It improves the accuracy and convenience of monitoring pipeline backfill settlement, enables comprehensive assessment of backfill quality, and enhances safety management efficiency through a real-time early warning mechanism, breaking through the limitations of traditional technologies.
Smart Images

Figure CN120747863B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing and machine learning, and particularly relates to a pipeline backfill settlement monitoring method and system based on image recognition. BACKGROUND
[0002] With the expansion of urban underground pipe network construction scale, the settlement problem of pipeline backfill area has become a key hidden danger affecting the safety of infrastructure. The traditional monitoring method mainly relies on manual inspection, GNSS positioning or single InSAR technology, but its limitations are increasingly prominent. For example, the GNSS monitoring point layout is costly and difficult to achieve wide coverage; manual inspection is long in cycle and limited by the environment, and it is difficult to respond to sudden settlement risks in time; and the monitoring means based on InSAR technology has wide area observation capability, but it is easily disturbed by signal in the vegetation coverage area, resulting in a decrease in deformation information extraction accuracy. In addition, the existing technology usually only focuses on surface deformation data, and lacks comprehensive analysis of key parameters such as soil compaction degree, and the compaction degree and settlement risk have strong correlation (such as the backfill soil not fully compacted is easy to cause late settlement due to rainwater penetration). Therefore, there is an urgent need for a monitoring system that can fuse multi-source heterogeneous data, simultaneously invert three-dimensional settlement and compaction degree, and has automatic early warning capability.
[0003] In recent years, the coordinated development of image recognition and multi-source remote sensing technology provides a new idea for this problem. Unmanned aerial vehicle aerial images can capture small changes in surface texture, color and other micro changes, providing direct basis for compaction degree analysis; and synthetic aperture radar (SAR) realizes millimeter-level deformation monitoring through interferometric techniques (such as SBAS-InSAR), making up for the deficiency of optical images in environmental shielding. However, how to efficiently fuse the features of the two types of data and establish the implicit correlation between settlement progress and compaction degree still faces technical challenges. Existing researches are mostly limited to single modal feature extraction, lack of deep learning framework for collaborative modeling of the two, and traditional models are difficult to adapt to complex terrain and dynamic environmental changes. SUMMARY
[0004] In view of the above technical problems, the present application provides a pipeline backfill settlement monitoring method and system based on image recognition, which realizes joint identification based on multi-modal data and improves the accuracy and convenience of pipeline backfill settlement monitoring.
[0005] In a first aspect, an embodiment of the present application provides a pipeline backfill settlement monitoring method based on image recognition, comprising:
[0006] acquiring a plurality of aerial images and a plurality of SAR remote sensing images of each preset time point of a target area;
[0007] image preprocessing is performed on the plurality of SAR remote sensing images based on an interferometric technique to obtain a plurality of differential interferograms;
[0008] The plurality of aerial images and the plurality of differential interferograms are input into a preset settlement monitoring model, so that the settlement monitoring model extracts elevation change features, ground texture features and color features from the plurality of aerial images based on an optical image branch, extracts geodetic deformation features from each of the differential interferograms based on a SAR image branch, and finally performs feature fusion on each feature based on a cross-attention mechanism, and outputs the current settlement progress and compaction degree of the target area according to the fused features;
[0009] According to the settlement progress and compaction degree, it is judged whether the current pipeline backfill settlement is abnormal, and corresponding early warning is performed based on the judgment result;
[0010] The settlement monitoring model is obtained by training an initial settlement monitoring model based on a labeled historical image dataset, and the initial settlement monitoring model is a dual-flow multi-modal feature fusion architecture, including an optical image branch, a SAR image branch, a cross-feature fusion module and a multi-task output head.
[0011] The embodiment of the present application provides a pipeline backfill settlement monitoring method based on image recognition. By obtaining aerial images and SAR remote sensing images of each preset time point of a target area and inputting them into a trained settlement monitoring model, the current settlement progress and compaction degree of the target area are obtained through image recognition technology, and automatic settlement monitoring of the target area is realized. Specifically, the embodiment of the present application fuses multi-modal data of unmanned aerial images and SAR remote sensing images, uses multi-modal feature fusion and cross-attention mechanism to realize deep correlation analysis of elevation change, ground texture, color features and geodetic deformation features, solves the problem of insufficient settlement monitoring accuracy of single data source, and effectively improves the accuracy of the model in monitoring the pipeline backfill settlement. At the same time, the embodiment of the present application designs a dual-flow multi-modal feature fusion architecture including an optical image branch and a SAR image branch, integrates the image processing process of aerial images and remote sensing images in the model, and improves the convenience of pipeline backfill settlement monitoring. Finally, the multi-task output head is used to simultaneously predict the settlement progress and compaction degree, which breaks through the limitation of traditional technology for ground deformation monitoring, realizes comprehensive evaluation of backfill quality, and significantly improves the safety control efficiency of pipeline backfill engineering in combination with real-time early warning mechanism.
[0012] In one possible implementation, the image preprocessing is performed on the plurality of SAR remote sensing images based on an interferometric technique to obtain a plurality of differential interferograms, including:
[0013] Each of the SAR remote sensing images is extracted in time sequence;
[0014] If the extracted current SAR remote sensing image is not the SAR remote sensing image of the last moment, then the current SAR remote sensing image is used as the main image, the SAR remote sensing image of the next moment is used as the auxiliary image, and a corresponding differential interferogram is generated based on the main image and the auxiliary image.
[0015] If the extracted current SAR remote sensing image is the SAR remote sensing image of the last moment, then the image preprocessing ends.
[0016] This application provides a method for generating differential interferograms. It uses a time-series sliding window approach to process each SAR remote sensing image sequentially and generate corresponding differential interferograms. This avoids nonlinear deformation interference between multi-temporal SAR data, effectively suppresses atmospheric noise through a master-slave image registration strategy, and ensures the temporal continuity and deformation information integrity of the differential interferogram sequence, providing a high-quality data foundation for subsequent deformation feature extraction and image recognition.
[0017] Furthermore, the step of generating a corresponding differential interferogram based on the main image and the auxiliary image includes:
[0018] Based on preset orbital parameters and pixel-level registration algorithms, the auxiliary image is registered to the coordinate system of the main image to obtain a registered image;
[0019] Perform a complex conjugate product operation on the registered image and the main image to generate the original interferogram;
[0020] The original interferogram is phase-unwrapped using the minimum cost flow algorithm to obtain the differential interferogram.
[0021] In the process of generating differential interferograms from primary and secondary images, this application embodiment uses a phase unwrapping method based on the minimum cost flow algorithm. While ensuring computational efficiency, it significantly reduces residual phase noise in the traditional unwrapping process. At the same time, it combines complex conjugate product operations to enhance the sensitivity of microwave signals to small surface deformations, thereby improving the ability of SAR data to resolve millimeter-level deformations and improving the accuracy of subsequent image recognition by the model.
[0022] In one possible implementation, the extraction of elevation variation features, ground texture features, and color features from the plurality of aerial images based on optical image branching includes:
[0023] Based on a preset computer vision algorithm, the aerial images are reconstructed in three dimensions using feature point detection and matching methods to obtain corresponding digital elevation models.
[0024] Each of the digital elevation models is compared with a preset digital elevation model before settlement, and the comparison results are arranged in chronological order to obtain the elevation change characteristics.
[0025] The latest UAV aerial image from the plurality of aerial images is input into a preset residual network so that the residual network can extract the ground texture features by edge detection. The residual network has frozen the bottom convolutional layers.
[0026] Image segmentation is performed on the latest drone aerial image to determine the subsidence area in the latest drone aerial image;
[0027] Based on a preset HSV mapping network, the subsidence area in the latest UAV aerial imagery is converted into a corresponding HSV histogram, which serves as the color feature.
[0028] This application provides a method for feature extraction from UAV aerial imagery. In extracting elevation change features, a digital elevation model (DEM) is reconstructed, and sub-centimeter-level elevation change quantification is achieved through temporal comparison. In extracting ground texture features, a residual network with frozen bottom convolutional layers is used to focus on detecting subtle surface texture changes, avoiding shallow noise interference. When extracting color features, image segmentation and HSV mapping techniques are combined to extract corresponding HSV histograms, accurately capturing visual representations strongly correlated with compaction, such as soil moisture content and degree of exposure, overcoming the spatiotemporal limitations of traditional methods relying solely on density meters. This application improves the accuracy and comprehensiveness of feature extraction from UAV aerial imagery through these three methods, thereby enhancing the accuracy of subsequent monitoring of pipeline backfill settlement.
[0029] In one possible implementation, the extraction of topographic features from the various differential interferograms based on SAR image branches includes:
[0030] The original amplitude matrix and the original phase matrix are extracted from each of the differential interferograms, and the original amplitude matrix and the original phase matrix are standardized to obtain amplitude matrix and phase matrix with the same coordinate system and the same size.
[0031] After performing multi-layer downsampling encoding and multi-layer upsampling decoding on each of the amplitude matrices and each of the phase matrices through several preset complex convolutional layers, multi-scale fusion features corresponding to each of the differential interferograms are generated through an attention gating mechanism, wherein skip connections are used between each layer of downsampling encoding and upsampling decoding.
[0032] Each of the multi-scale fusion features is processed by a 1×1 convolution to generate a deformation feature map corresponding to each of the difference interferograms;
[0033] The geological deformation features are obtained by integrating the various deformation feature maps in chronological order.
[0034] This application provides a method for feature extraction from differential interferograms (DIAs). The method extracts the amplitude and phase matrices from the DIA, processes them using complex convolutional layers and an attention gating mechanism to generate corresponding multi-scale fusion features, and then obtains topographic deformation features based on these multi-scale fusion features. The complex convolutional layers can simultaneously process amplitude and phase information, preserving the electromagnetic wave phase characteristics of SAR data. The hop-connected encoder-decoder structure enhances the scale invariance of the multi-scale deformation features. The attention gating mechanism dynamically strengthens the feature representation of key deformation regions, significantly improving the modeling accuracy of deformation distribution under complex terrain conditions, thereby improving the accuracy of subsequent monitoring of pipeline backfill settlement.
[0035] In one possible implementation, the feature fusion based on the cross-attention mechanism, and the output of the current settlement progress and compaction degree of the target area based on the fused features, includes:
[0036] Based on the cross-attention mechanism, the elevation change features and the terrain deformation features are fused to obtain the first fused feature;
[0037] By feature stitching, the ground texture features and color features are fused to obtain a second fused feature;
[0038] After performing global average pooling on the first fused feature, it is input into a preset fully connected layer to generate the current settlement progress of the target area.
[0039] Based on the preset compaction degree classification, the second fused feature is subjected to 1×1 convolution processing to generate a classification feature map;
[0040] After the classification feature map is processed by spatial pyramid pooling, it is input into a preset Softmax classifier to generate the current compaction degree of the target area.
[0041] This application provides a feature fusion method that differs from existing methods that directly fuse various features. Instead, this method selectively fuses some features from two branches, specifically fusing elevation change features with temporal characteristics with topographic deformation features, and stitching together ground texture and color features without temporal characteristics. This achieves accurate feature fusion for cross-modal data, improving the accuracy of monitoring pipeline backfill settlement. In the specific feature fusion process, this application utilizes a cross-attention mechanism through cross-modal feature response mapping to eliminate spatial alignment bias between optical imagery and SAR data. Spatial pyramid pooling combined with classification feature map processing enhances the model's spatial localization ability for local compaction anomalies while maintaining computational efficiency, thus improving the accuracy of feature fusion.
[0042] Furthermore, the process of training the initial settlement monitoring model based on a labeled historical image dataset to obtain the settlement monitoring model includes:
[0043] Acquire multi-temporal UAV aerial images, multi-temporal SAR remote sensing images, and monitoring results at each time point from several backfilling projects to construct a labeled historical image dataset.
[0044] The optical image branch is constructed based on computer vision algorithms, image segmentation algorithms, residual networks, and HSV mapping networks; the SAR image branch is constructed based on complex convolutional networks; the cross-feature fusion module is constructed based on a cross-attention mechanism; the multi-task output head is constructed based on global average pooling layers, fully connected layers, convolutional layers, pyramid pooling layers, and a Softmax classifier; and the initial settlement monitoring model is constructed based on the optical image branch, SAR image branch, cross-feature fusion module, and multi-task output head.
[0045] The initial settlement monitoring model is trained using a preset multi-task joint loss function and the historical image dataset to obtain the settlement monitoring model. The multi-task joint loss function consists of a settlement regression loss function and a compaction degree classification loss function.
[0046] Furthermore, the step of training the initial settlement monitoring model based on a preset multi-task joint loss function and the historical image dataset to obtain the settlement monitoring model includes:
[0047] The historical image dataset is divided into a first training dataset and a second training dataset;
[0048] Based on the first training dataset and the multi-task joint loss function, the initial settlement monitoring model is subjected to several rounds of single-branch alternating training to obtain the first settlement monitoring model. In each single-branch alternating training process, the model parameters other than the optical image branch are frozen and the model parameters of the optical image branch are updated through the first training dataset, or the model parameters other than the SAR image branch are frozen and the model parameters of the SAR image branch are updated through the first training dataset.
[0049] Based on the second training dataset and the multi-task joint loss function, the first settlement monitoring model is jointly trained end-to-end to obtain the settlement monitoring model.
[0050] This application provides a training method for a settlement monitoring model, which combines a multi-task joint loss function with image data from several backfilling projects to perform phased alternating training on the initial settlement monitoring model. The multi-task joint loss function, through a gradient sharing mechanism between settlement regression loss and compaction classification loss, constrains the model to simultaneously optimize spatial deformation perception and material property recognition capabilities, avoiding feature representation shifts caused by single-task training. The phased alternating training strategy effectively alleviates the parameter coupling problem of the dual-flow model: the first stage of single-branch pre-training ensures that the optical / SAR branches each converge to a local optimum; the second stage of end-to-end fine-tuning promotes cross-modal feature collaborative optimization, significantly reducing the risk of gradient oscillations in the early stages of joint training, improving model convergence stability, and thus enhancing the accuracy of pipeline backfill settlement monitoring.
[0051] Secondly, embodiments of this application provide a pipeline backfill settlement monitoring system based on image recognition, including an acquisition module, an image preprocessing module, a settlement monitoring module, and an early warning module;
[0052] The acquisition module is used to acquire several aerial images and several SAR remote sensing images of the target area at various preset time points;
[0053] The image preprocessing module is used to perform image preprocessing on the several SAR remote sensing images based on interferometry technology to obtain several differential interferograms;
[0054] The settlement monitoring module is used to input the aerial images and differential interferograms into a preset settlement monitoring model, so that the settlement monitoring model extracts elevation change features, ground texture features and color features from the aerial images based on the optical image branch, extracts topographic deformation features from each differential interferogram based on the SAR image branch, and finally performs feature fusion on each feature based on the cross attention mechanism, and outputs the current settlement progress and compaction degree of the target area according to the fused features.
[0055] The early warning module is used to determine whether there is any abnormality in the current pipeline backfill settlement based on the settlement progress and compaction degree, and to issue a corresponding early warning based on the judgment result;
[0056] The settlement monitoring model is obtained by training an initial settlement monitoring model based on a labeled historical image dataset. The initial settlement monitoring model is a dual-stream multimodal feature fusion architecture, including an optical image branch, a SAR image branch, a cross-feature fusion module, and a multi-task output head.
[0057] Furthermore, the monitoring system also includes a model training module, which includes a dataset construction unit, a model construction unit, and a training unit.
[0058] The dataset construction unit is used to acquire multi-temporal UAV aerial images, multi-temporal SAR remote sensing images and monitoring results corresponding to each time moment in several backfilling projects, and to construct a labeled historical image dataset.
[0059] The model building unit is used to construct the optical image branch based on computer vision algorithms, image segmentation algorithms, residual networks, and HSV mapping networks; construct the SAR image branch based on complex convolutional networks; construct the cross-feature fusion module based on a cross-attention mechanism; construct the multi-task output head based on global average pooling layers, fully connected layers, convolutional layers, pyramid pooling layers, and a Softmax classifier; and construct the initial settlement monitoring model based on the optical image branch, SAR image branch, cross-feature fusion module, and multi-task output head.
[0060] The training unit is used to train the initial settlement monitoring model according to the preset multi-task joint loss function and the historical image dataset to obtain the settlement monitoring model. The multi-task joint loss function consists of a settlement regression loss function and a compaction classification loss function. Attached Figure Description
[0061] Figure 1 A schematic flowchart of a pipeline backfill settlement monitoring method based on image recognition provided in an embodiment of this application;
[0062] Figure 2 This is a schematic diagram of a pipeline backfill settlement monitoring system based on image recognition, provided as an embodiment of this application. Detailed Implementation
[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0064] It should be noted that the step numbers in this document are only for the convenience of explaining the specific embodiments and are not intended to limit the order in which the steps are performed. In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0065] Example 1:
[0066] like Figure 1As shown, Embodiment 1 provides a method for monitoring pipeline backfill settlement based on image recognition, including steps S1-S4:
[0067] Step S1: Acquire several aerial images and several SAR remote sensing images of the target area at various preset time points;
[0068] Step S2: Based on interferometry, perform image preprocessing on the several SAR remote sensing images to obtain several differential interferograms;
[0069] Step S3: Input the aerial images and differential interferograms into a preset settlement monitoring model, so that the settlement monitoring model extracts elevation change features, ground texture features and color features from the aerial images based on the optical image branch, extracts topographic deformation features from each differential interferogram based on the SAR image branch, and finally performs feature fusion on each feature based on the cross attention mechanism, and outputs the current settlement progress and compaction degree of the target area according to the fused features.
[0070] Step S4: Determine whether there is any abnormality in the current pipeline backfill settlement based on the settlement progress and compaction degree, and issue a corresponding warning based on the judgment result;
[0071] The settlement monitoring model is obtained by training an initial settlement monitoring model based on a labeled historical image dataset. The initial settlement monitoring model is a dual-stream multimodal feature fusion architecture, including an optical image branch, a SAR image branch, a cross-feature fusion module, and a multi-task output head.
[0072] This application provides an image recognition-based method for monitoring pipeline backfill settlement. It acquires aerial images and SAR remote sensing images of the target area at various preset time points and inputs them into a trained settlement monitoring model. Image recognition technology is used to obtain the current settlement progress and compaction degree of the target area, achieving automated settlement monitoring. Specifically, this application integrates multimodal data from UAV aerial images and SAR remote sensing images. Multimodal feature fusion and cross-attention mechanisms are used to achieve deep correlation analysis of elevation changes, ground texture, color features, and topographic deformation features. This solves the problem of insufficient accuracy of settlement monitoring from a single data source and effectively improves the model's accuracy in monitoring pipeline backfill settlement. Furthermore, this application designs a dual-stream multimodal feature fusion architecture that includes optical image branches and SAR image branches, integrating the image processing of aerial images and remote sensing images within the model, improving the convenience of monitoring pipeline backfill settlement. Finally, by synchronously predicting settlement progress and compaction degree through multi-task output heads, the limitations of traditional technology that only monitors surface deformation are overcome, enabling a comprehensive assessment of backfill quality. Combined with a real-time early warning mechanism, this significantly improves the safety management efficiency of pipeline backfilling projects.
[0073] In a preferred embodiment, in step S1, a drone (such as the DJI Phantom series) equipped with a high-precision RGB camera is used to conduct multi-temporal aerial photography of the target area at a flight altitude of 50-100 meters, with a ground resolution better than 5cm / pixel, covering the pipe backfill area and surrounding buffer zone, thereby obtaining several aerial images. Based on Sentinel-1 satellite C-band data (wavelength 5.6cm, incident angle 20°~45°), several SAR remote sensing images of the target area are obtained based on preset temporal and spatial resolutions.
[0074] In one possible implementation, step S2 involves preprocessing the plurality of SAR remote sensing images based on interferometry techniques to obtain a plurality of differential interferograms, including:
[0075] Each of the SAR remote sensing images is extracted sequentially based on time order;
[0076] If the extracted current SAR remote sensing image is not the SAR remote sensing image of the last moment, then the current SAR remote sensing image is used as the main image, the SAR remote sensing image of the next moment is used as the auxiliary image, and a corresponding differential interferogram is generated based on the main image and the auxiliary image.
[0077] If the extracted current SAR remote sensing image is the SAR remote sensing image of the last moment, then the image preprocessing ends.
[0078] This application provides a method for generating differential interferograms. It uses a time-series sliding window approach to process each SAR remote sensing image sequentially and generate corresponding differential interferograms. This avoids nonlinear deformation interference between multi-temporal SAR data, effectively suppresses atmospheric noise through a master-slave image registration strategy, and ensures the temporal continuity and deformation information integrity of the differential interferogram sequence, providing a high-quality data foundation for subsequent deformation feature extraction and image recognition.
[0079] Furthermore, the step of generating a corresponding differential interferogram based on the main image and the auxiliary image includes:
[0080] Based on preset orbital parameters and pixel-level registration algorithms, the auxiliary image is registered to the coordinate system of the main image to obtain a registered image;
[0081] Perform a complex conjugate product operation on the registered image and the main image to generate the original interferogram;
[0082] The original interferogram is phase-unwrapped using the minimum cost flow algorithm to obtain the differential interferogram.
[0083] In the process of generating differential interferograms from primary and secondary images, this application embodiment uses a phase unwrapping method based on the minimum cost flow algorithm. While ensuring computational efficiency, it significantly reduces residual phase noise in the traditional unwrapping process. At the same time, it combines complex conjugate product operations to enhance the sensitivity of microwave signals to small surface deformations, thereby improving the ability of SAR data to resolve millimeter-level deformations and improving the accuracy of subsequent image recognition by the model.
[0084] In one possible implementation, step S3, which involves extracting elevation change features, ground texture features, and color features from the plurality of aerial images based on optical image branching, includes:
[0085] Based on a preset computer vision algorithm, the aerial images are reconstructed in three dimensions using feature point detection and matching methods to obtain corresponding digital elevation models.
[0086] Each of the digital elevation models is compared with a preset digital elevation model before settlement, and the comparison results are arranged in chronological order to obtain the elevation change characteristics.
[0087] The latest UAV aerial image from the plurality of aerial images is input into a preset residual network so that the residual network can extract the ground texture features by edge detection. The residual network has frozen the bottom convolutional layers.
[0088] Image segmentation is performed on the latest drone aerial image to determine the subsidence area in the latest drone aerial image;
[0089] Based on a preset HSV mapping network, the subsidence area in the latest UAV aerial imagery is converted into a corresponding HSV histogram, which serves as the color feature.
[0090] This application provides a method for feature extraction from UAV aerial imagery. In extracting elevation change features, a digital elevation model (DEM) is reconstructed, and sub-centimeter-level elevation change quantification is achieved through temporal comparison. In extracting ground texture features, a residual network with frozen bottom convolutional layers is used to focus on detecting subtle surface texture changes, avoiding shallow noise interference. When extracting color features, image segmentation and HSV mapping techniques are combined to extract corresponding HSV histograms, accurately capturing visual representations strongly correlated with compaction, such as soil moisture content and degree of exposure, overcoming the spatiotemporal limitations of traditional methods relying solely on density meters. This application improves the accuracy and comprehensiveness of feature extraction from UAV aerial imagery through these three methods, thereby enhancing the accuracy of subsequent monitoring of pipeline backfill settlement.
[0091] In a preferred embodiment, the SFM (Structure from Motion) algorithm is used to perform 3D reconstruction on each aerial image, generating corresponding digital elevation models (DEMs). Then, a difference analysis is performed between each DEM and the baseline DEM before backfilling to generate a time-series elevation change matrix, which serves as the elevation change feature. A ResNet-50 network with fixed parameters for the first five layers is used to extract features from the latest aerial image, obtaining an edge texture feature map, which serves as the ground texture feature. The subsidence area of the aerial image is segmented using a U-Net model, where the IoU parameter is set to ≥0.85. Then, the RGB image of the subsidence area is converted to the HSV color space, and the HSV histogram of the subsidence area is statistically analyzed, serving as the color feature.
[0092] In one possible implementation, step S3, whereby the SAR image-based branch extracts topographic features from each of the differential interferograms, includes:
[0093] The original amplitude matrix and the original phase matrix are extracted from each of the differential interferograms, and the original amplitude matrix and the original phase matrix are standardized to obtain amplitude matrix and phase matrix with the same coordinate system and the same size.
[0094] After performing multi-layer downsampling encoding and multi-layer upsampling decoding on each of the amplitude matrices and each of the phase matrices through several preset complex convolutional layers, multi-scale fusion features corresponding to each of the differential interferograms are generated through an attention gating mechanism, wherein skip connections are used between each layer of downsampling encoding and upsampling decoding.
[0095] Each of the multi-scale fusion features is processed by a 1×1 convolution to generate a deformation feature map corresponding to each of the difference interferograms;
[0096] The geological deformation features are obtained by integrating the various deformation feature maps in chronological order.
[0097] This application provides a method for feature extraction from differential interferograms (DIAs). The method extracts the amplitude and phase matrices from the DIA, processes them using complex convolutional layers and an attention gating mechanism to generate corresponding multi-scale fusion features, and then obtains topographic deformation features based on these multi-scale fusion features. The complex convolutional layers can simultaneously process amplitude and phase information, preserving the electromagnetic wave phase characteristics of SAR data. The hop-connected encoder-decoder structure enhances the scale invariance of the multi-scale deformation features. The attention gating mechanism dynamically strengthens the feature representation of key deformation regions, significantly improving the modeling accuracy of deformation distribution under complex terrain conditions, thereby improving the accuracy of subsequent monitoring of pipeline backfill settlement.
[0098] In a preferred embodiment, the original amplitude and phase matrices of the differential interferogram are normalized (mean 0, variance 1) to a uniform size of 256×256. Then, a complex convolutional network (Complex CNN) is used to convolve the amplitude and phase matrices to generate corresponding multi-scale fusion features. The complex convolutional network comprises four downsampling layers (3×3 kernels, stride 2) and four upsampling layers (3×3 transposed kernels), with an attention gate (SEBlock) embedded in between. The downsampling layers output complex feature maps (real and imaginary parts), and the upsampling layers generate multi-scale fusion features, preserving multi-scale edge information through skip connections. Finally, a 1×1 convolution compresses the channels to 8 dimensions, outputting a deformation feature map corresponding to the differential interferogram.
[0099] In one possible implementation, step S3, which involves fusing the features based on a cross-attention mechanism and outputting the current settlement progress and compaction degree of the target area according to the fused features, includes:
[0100] Based on the cross-attention mechanism, the elevation change features and the terrain deformation features are fused to obtain the first fused feature;
[0101] By feature stitching, the ground texture features and color features are fused to obtain a second fused feature;
[0102] After performing global average pooling on the first fused feature, it is input into a preset fully connected layer to generate the current settlement progress of the target area.
[0103] Based on the preset compaction degree classification, the second fused feature is subjected to 1×1 convolution processing to generate a classification feature map;
[0104] After the classification feature map is processed by spatial pyramid pooling, it is input into a preset Softmax classifier to generate the current compaction degree of the target area.
[0105] This application provides a feature fusion method that differs from existing methods that directly fuse various features. Instead, this method selectively fuses some features from two branches, specifically fusing elevation change features with temporal characteristics with topographic deformation features, and stitching together ground texture and color features without temporal characteristics. This achieves accurate feature fusion for cross-modal data, improving the accuracy of monitoring pipeline backfill settlement. In the specific feature fusion process, this application utilizes a cross-attention mechanism through cross-modal feature response mapping to eliminate spatial alignment bias between optical imagery and SAR data. Spatial pyramid pooling combined with classification feature map processing enhances the model's spatial localization ability for local compaction anomalies while maintaining computational efficiency, thus improving the accuracy of feature fusion.
[0106] In a preferred embodiment, the elevation variation features output by the optical branch and the topographic deformation features output by the SAR branch are acquired, both with a feature dimension of 512. Linear interpolation is used to align them to a unified time scale T, ensuring consistency in the time dimension. Query(Q), Key(K), and Value(V) are constructed:
[0107]
[0108] Where Q is the query vector, representing the elevation change feature that needs to be focused on; K is the key vector, representing the index of the topographic deformation feature; and V is the value vector, representing the actual information of the topographic deformation feature. For learnable weight matrix, Characteristics of elevation changes, This represents the characteristics of topographic deformation. Then, by calculating QK... T The dot product of the elevation change features and the terrain deformation features is used to obtain the similarity weights. These similarity weights are then converted into a probability distribution, i.e., attention weights, using the Softmax function. Finally, the attention weights are used to weight and sum the Value (deformation features) to output the first fused feature. It contains information relating elevation and deformation.
[0109] The ground texture features (256 channels, size H×W) and color features (3 channels, size H×W) are obtained and then stitched together along the channel dimension to obtain the second fused feature. .
[0110] For the first fused feature, average pooling is performed on the time dimension T to obtain the feature vector. The data is then input into a preset fully connected network to generate the current settlement progress S of the target area. The fully connected network consists of two fully connected layers (512→128→1) with ReLU activation function.
[0111] For the second fusion feature, a 1×1 convolution is used to compress the number of channels from 259 to 64, outputting a classification feature map. Then, pooling layers with different kernel sizes are applied to the feature map to obtain the corresponding pooling results. These pooling layers with different kernel sizes include 1×1, 2×2, and 4×4 kernels. The pooling results are then flattened and concatenated to output the feature vector. Finally, the feature vectors The data is input into a preset Softmax classifier, which then generates the current compaction degree of the target area.
[0112] Furthermore, the process of training the initial settlement monitoring model based on a labeled historical image dataset to obtain the settlement monitoring model includes:
[0113] Acquire multi-temporal UAV aerial images, multi-temporal SAR remote sensing images, and monitoring results at each time point from several backfilling projects to construct a labeled historical image dataset.
[0114] The optical image branch is constructed based on computer vision algorithms, image segmentation algorithms, residual networks, and HSV mapping networks; the SAR image branch is constructed based on complex convolutional networks; the cross-feature fusion module is constructed based on a cross-attention mechanism; the multi-task output head is constructed based on global average pooling layers, fully connected layers, convolutional layers, pyramid pooling layers, and a Softmax classifier; and the initial settlement monitoring model is constructed based on the optical image branch, SAR image branch, cross-feature fusion module, and multi-task output head.
[0115] The initial settlement monitoring model is trained using a preset multi-task joint loss function and the historical image dataset to obtain the settlement monitoring model. The multi-task joint loss function consists of a settlement regression loss function and a compaction degree classification loss function.
[0116] Furthermore, the step of training the initial settlement monitoring model based on a preset multi-task joint loss function and the historical image dataset to obtain the settlement monitoring model includes:
[0117] The historical image dataset is divided into a first training dataset and a second training dataset;
[0118] Based on the first training dataset and the multi-task joint loss function, the initial settlement monitoring model is subjected to several rounds of single-branch alternating training to obtain the first settlement monitoring model. In each single-branch alternating training process, the model parameters other than the optical image branch are frozen and the model parameters of the optical image branch are updated through the first training dataset, or the model parameters other than the SAR image branch are frozen and the model parameters of the SAR image branch are updated through the first training dataset.
[0119] Based on the second training dataset and the multi-task joint loss function, the first settlement monitoring model is jointly trained end-to-end to obtain the settlement monitoring model.
[0120] This application provides a training method for a settlement monitoring model, which combines a multi-task joint loss function with image data from several backfilling projects to perform phased alternating training on the initial settlement monitoring model. The multi-task joint loss function, through a gradient sharing mechanism between settlement regression loss and compaction classification loss, constrains the model to simultaneously optimize spatial deformation perception and material property recognition capabilities, avoiding feature representation shifts caused by single-task training. The phased alternating training strategy effectively alleviates the parameter coupling problem of the dual-flow model: the first stage of single-branch pre-training ensures that the optical / SAR branches each converge to a local optimum; the second stage of end-to-end fine-tuning promotes cross-modal feature collaborative optimization, significantly reducing the risk of gradient oscillations in the early stages of joint training, improving model convergence stability, and thus enhancing the accuracy of pipeline backfill settlement monitoring.
[0121] In a preferred embodiment, the training process of the settlement monitoring model is as follows:
[0122] 1. Data Acquisition and Labeling. Multi-source data from 10 typical pipeline backfilling projects were collected, and settlement labels were generated by comparing GNSS measured coordinates (accuracy ±1cm) with the baseline DEM. Compaction labels were generated using the ring cutter method, and a labeled historical image dataset was constructed.
[0123] 2. Data preprocessing. Aerial images: cropped into 224×224 image blocks according to time series, and subsidence area masks are marked (IoU≥0.85); SAR images: registered to a unified coordinate system, generating differential interferograms (deformation accuracy ±2mm), and extracting amplitude / phase matrices.
[0124] 3. Dataset Partitioning. The training, validation, and test sets are partitioned in a 7:2:1 ratio to ensure time series continuity. Furthermore, within the training set, 50% of the data is separated into single-modal data containing only aerial imagery or SAR images, serving as the first training dataset, while the remaining data serves as the second training dataset.
[0125] 4. Single-branch alternating training. Based on the first training set and a preset loss function, the initial settlement monitoring model undergoes several rounds of single-branch alternating training to obtain the first settlement monitoring model. Each round of single-branch alternating training includes optical branch pre-training and SAR branch pre-training. In optical branch pre-training, the SAR branch parameters are frozen, only the optical branch is updated, and the Adam optimizer is used for training with a learning rate of 1e-4 and a batch size of 16. In SAR branch pre-training, the optical branch parameters are frozen, only the SAR branch is updated, and the optimizer parameters are set as above.
[0126] 5. End-to-end joint training. Based on the second training dataset and the multi-task joint loss function, the first settlement monitoring model is jointly trained end-to-end to obtain the settlement monitoring model. During the training process, all parameters are unfrozen, and the learning rate is optimized (1e-4→1e-5) using cosine annealing. At the same time, an early stopping mechanism is introduced, and the training is conducted for 30 epochs with an early stopping threshold set to 3.
[0127] Example 2:
[0128] like Figure 2 As shown, Embodiment 2 provides a pipeline backfill settlement monitoring system based on image recognition, including an acquisition module 10, an image preprocessing module 20, a settlement monitoring module 30, and an early warning module 40;
[0129] The acquisition module 10 is used to acquire several aerial images and several SAR remote sensing images of the target area at various preset time points;
[0130] The image preprocessing module 20 is used to perform image preprocessing on the plurality of SAR remote sensing images based on interferometry technology to obtain a plurality of differential interferograms;
[0131] The settlement monitoring module 30 is used to input the plurality of aerial images and the plurality of differential interferograms into a preset settlement monitoring model, so that the settlement monitoring model extracts elevation change features, ground texture features and color features from the plurality of aerial images based on the optical image branch, and extracts geomorphic features from each of the differential interferograms based on the SAR image branch. Finally, it performs feature fusion on each feature based on the cross attention mechanism, and outputs the current settlement progress and compaction degree of the target area according to the fused features.
[0132] The early warning module 40 is used to determine whether there is any abnormality in the current pipeline backfill settlement based on the settlement progress and compaction degree, and to issue a corresponding early warning based on the judgment result;
[0133] The settlement monitoring model is obtained by training an initial settlement monitoring model based on a labeled historical image dataset. The initial settlement monitoring model is a dual-stream multimodal feature fusion architecture, including an optical image branch, a SAR image branch, a cross-feature fusion module, and a multi-task output head.
[0134] Furthermore, the image preprocessing module 20 performs image preprocessing on the plurality of SAR remote sensing images based on interferometry technology to obtain a plurality of differential interferograms, including:
[0135] Each of the SAR remote sensing images is extracted sequentially based on time order;
[0136] If the extracted current SAR remote sensing image is not the SAR remote sensing image of the last moment, then the current SAR remote sensing image is used as the main image, the SAR remote sensing image of the next moment is used as the auxiliary image, and a corresponding differential interferogram is generated based on the main image and the auxiliary image.
[0137] If the extracted current SAR remote sensing image is the SAR remote sensing image of the last moment, then the image preprocessing ends.
[0138] Furthermore, the step of generating a corresponding differential interferogram based on the main image and the auxiliary image includes:
[0139] Based on preset orbital parameters and pixel-level registration algorithms, the auxiliary image is registered to the coordinate system of the main image to obtain a registered image;
[0140] Perform a complex conjugate product operation on the registered image and the main image to generate the original interferogram;
[0141] The original interferogram is phase-unwrapped using the minimum cost flow algorithm to obtain the differential interferogram.
[0142] In one possible implementation, the settlement monitoring module 30 extracts elevation change features, ground texture features, and color features from the plurality of aerial images based on optical image branching, including:
[0143] Based on a preset computer vision algorithm, the aerial images are reconstructed in three dimensions using feature point detection and matching methods to obtain corresponding digital elevation models.
[0144] Each of the digital elevation models is compared with a preset digital elevation model before settlement, and the comparison results are arranged in chronological order to obtain the elevation change characteristics.
[0145] The latest UAV aerial image from the plurality of aerial images is input into a preset residual network so that the residual network can extract the ground texture features by edge detection. The residual network has frozen the bottom convolutional layers.
[0146] Image segmentation is performed on the latest drone aerial image to determine the subsidence area in the latest drone aerial image;
[0147] Based on a preset HSV mapping network, the subsidence area in the latest UAV aerial imagery is converted into a corresponding HSV histogram, which serves as the color feature.
[0148] In one possible implementation, the subsidence monitoring module 30 extracts topographic deformation features from each of the differential interferograms based on SAR image branches, including:
[0149] The original amplitude matrix and the original phase matrix are extracted from each of the differential interferograms, and the original amplitude matrix and the original phase matrix are standardized to obtain amplitude matrix and phase matrix with the same coordinate system and the same size.
[0150] After performing multi-layer downsampling encoding and multi-layer upsampling decoding on each of the amplitude matrices and each of the phase matrices through several preset complex convolutional layers, multi-scale fusion features corresponding to each of the differential interferograms are generated through an attention gating mechanism, wherein skip connections are used between each layer of downsampling encoding and upsampling decoding.
[0151] Each of the multi-scale fusion features is processed by a 1×1 convolution to generate a deformation feature map corresponding to each of the difference interferograms;
[0152] The geological deformation features are obtained by integrating the various deformation feature maps in chronological order.
[0153] Furthermore, the settlement monitoring module 30 performs feature fusion on various features based on a cross-attention mechanism, and outputs the current settlement progress and compaction degree of the target area according to the fused features, including:
[0154] Based on the cross-attention mechanism, the elevation change features and the terrain deformation features are fused to obtain the first fused feature;
[0155] By feature stitching, the ground texture features and color features are fused to obtain a second fused feature;
[0156] After performing global average pooling on the first fused feature, it is input into a preset fully connected layer to generate the current settlement progress of the target area.
[0157] Based on the preset compaction degree classification, the second fused feature is subjected to 1×1 convolution processing to generate a classification feature map;
[0158] After the classification feature map is processed by spatial pyramid pooling, it is input into a preset Softmax classifier to generate the current compaction degree of the target area.
[0159] In one possible implementation, the monitoring system further includes a model training module, which includes a dataset construction unit, a model construction unit, and a training unit.
[0160] The dataset construction unit is used to acquire multi-temporal UAV aerial images, multi-temporal SAR remote sensing images and monitoring results corresponding to each time moment in several backfilling projects, and to construct a labeled historical image dataset.
[0161] The model building unit is used to construct the optical image branch based on computer vision algorithms, image segmentation algorithms, residual networks, and HSV mapping networks; construct the SAR image branch based on complex convolutional networks; construct the cross-feature fusion module based on a cross-attention mechanism; construct the multi-task output head based on global average pooling layers, fully connected layers, convolutional layers, pyramid pooling layers, and a Softmax classifier; and construct the initial settlement monitoring model based on the optical image branch, SAR image branch, cross-feature fusion module, and multi-task output head.
[0162] The training unit is used to train the initial settlement monitoring model according to the preset multi-task joint loss function and the historical image dataset to obtain the settlement monitoring model. The multi-task joint loss function consists of a settlement regression loss function and a compaction classification loss function.
[0163] Furthermore, the step of training the initial settlement monitoring model based on a preset multi-task joint loss function and the historical image dataset to obtain the settlement monitoring model includes:
[0164] The historical image dataset is divided into a first training dataset and a second training dataset;
[0165] Based on the first training dataset and the multi-task joint loss function, the initial settlement monitoring model is subjected to several rounds of single-branch alternating training to obtain the first settlement monitoring model. In each single-branch alternating training process, the model parameters other than the optical image branch are frozen and the model parameters of the optical image branch are updated through the first training dataset, or the model parameters other than the SAR image branch are frozen and the model parameters of the SAR image branch are updated through the first training dataset.
[0166] Based on the second training dataset and the multi-task joint loss function, the first settlement monitoring model is jointly trained end-to-end to obtain the settlement monitoring model.
[0167] For a more detailed explanation of the working principle and procedures of this embodiment, please refer to the relevant description in Embodiment 1.
[0168] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.
Claims
1. A method for monitoring settlement of backfill in a pipeline based on image recognition, characterized in that, The method comprises the following steps: acquiring a plurality of aerial images and a plurality of SAR remote sensing images of a target area at a plurality of preset time points; performing image preprocessing on the plurality of SAR remote sensing images based on an interferometric technique to obtain a plurality of differential interferograms; inputting the plurality of aerial images and the plurality of differential interferograms into a preset settlement monitoring model, so that the settlement monitoring model extracts elevation change features, ground texture features and color features from the plurality of aerial images based on an optical image branch, extracts geodetic deformation features from each of the differential interferograms based on a SAR image branch, and finally performs feature fusion on each feature based on a cross-attention mechanism, and outputs the current settlement progress and compaction degree of the target area according to the fused features; judging whether the current pipeline backfill settlement is abnormal based on the settlement progress and the compaction degree, and performing corresponding early warning based on the judgment result; wherein the settlement monitoring model is obtained by training an initial settlement monitoring model based on a labeled historical image dataset, and the initial settlement monitoring model is a dual-flow multi-modal feature fusion architecture, which includes an optical image branch, a SAR image branch, a cross-feature fusion module and a multi-task output head.
2. The image recognition-based pipe backfill settlement monitoring method of claim 1, wherein, The image preprocessing on the plurality of SAR remote sensing images based on the interferometric technique comprises the following steps: extracting each of the SAR remote sensing images in time sequence; if the current SAR remote sensing image extracted is not the SAR remote sensing image at the last time, taking the current SAR remote sensing image as a main image, taking the SAR remote sensing image at the next time as an auxiliary image, and generating a corresponding differential interferogram according to the main image and the auxiliary image; if the current SAR remote sensing image extracted is the SAR remote sensing image at the last time, ending the image preprocessing.
3. The image recognition-based pipe backfill settlement monitoring method of claim 2, wherein, The generation of the corresponding differential interferogram according to the main image and the auxiliary image comprises the following steps: registering the auxiliary image to the main image coordinate system according to a preset orbit parameter and a pixel-level registration algorithm to obtain a registered image; performing complex conjugate multiplication operation on the registered image and the main image to generate an original interferogram; performing phase unwrapping on the original interferogram based on a minimum cost flow algorithm to obtain the differential interferogram.
4. The image recognition-based pipe backfill settlement monitoring method of claim 1, wherein, The extraction of the elevation change features, the ground texture features and the color features from the plurality of aerial images based on the optical image branch comprises the following steps: reconstructing each of the plurality of aerial images in a three-dimensional manner through feature point detection and matching according to a preset computer vision algorithm to obtain a corresponding digital elevation model; comparing each of the digital elevation models with a preset pre-settlement digital elevation model, and arranging the comparison results in time sequence to obtain the elevation change features; inputting the latest unmanned aerial vehicle aerial image in the plurality of aerial images into a preset residual network to extract the ground texture features through edge detection, wherein the residual network has a frozen bottom convolutional layer; performing image segmentation on the latest unmanned aerial vehicle aerial image to determine the settlement area in the latest unmanned aerial vehicle aerial image; The deposition area in the latest time unmanned aerial vehicle aerial image is converted into a corresponding HSV histogram based on a preset HSV mapping network, as the color feature.
5. The image recognition-based pipe backfill settlement monitoring method of claim 1, wherein, The SAR image branch extracts ground deformation features from each of the differential interferograms, including: An original amplitude matrix and an original phase matrix corresponding to each of the differential interferograms are extracted, and the original amplitude matrix and the original phase matrix are normalized to obtain amplitude matrices and phase matrices of the same coordinate system and the same size; After a plurality of preset complex convolution layers are used for multi-layer down-sampling encoding and multi-layer up-sampling decoding of each of the amplitude matrices and each of the phase matrices, each of the multi-scale fusion features corresponding to each of the differential interferograms is generated through an attention gate mechanism, wherein a skip connection is used between each layer of down-sampling encoding and up-sampling decoding; Each of the multi-scale fusion features is subjected to 1×1 convolution processing to generate a deformation feature map corresponding to each of the differential interferograms; The ground deformation features are obtained by integrating each of the deformation feature maps in time sequence.
6. The image recognition based pipe backfill settlement monitoring method of claim 1, wherein, The cross-attention mechanism is used for feature fusion of each feature, and the current deposition progress and compaction degree of the target area are output based on the fused features, including: The elevation change features and the ground deformation features are fused based on the cross-attention mechanism to obtain first fused features; The ground texture features and the color features are fused through feature splicing to obtain second fused features; After global average pooling processing of the first fused features, the initial deposition monitoring model is trained based on a preset compaction degree classification, and the current deposition progress of the target area is generated by further inputting the first fused features into a preset fully connected layer. The classification feature map is generated by performing 1×1 convolution processing on the second fused features based on a preset compaction degree classification. After spatial pyramid pooling processing of the classification feature map, the current compaction degree of the target area is generated by further inputting the classification feature map into a preset Softmax classifier.
7. The image recognition based pipe backfill settlement monitoring method according to any one of claims 1-6, wherein, The deposition monitoring model is obtained by training an initial deposition monitoring model based on a labeled historical image dataset, including: A labeled historical image dataset is constructed by obtaining multi-temporal unmanned aerial vehicle aerial images, multi-temporal SAR remote sensing images, and corresponding monitoring results of each time in a plurality of backfill projects; The optical image branch is constructed based on a computer vision algorithm, an image segmentation algorithm, a residual network, and an HSV mapping network; the SAR image branch is constructed based on a complex convolution network; the cross-feature fusion module is constructed based on a cross-attention mechanism; the multi-task output head is constructed based on a global average pooling layer, a fully connected layer, a convolution layer, a pyramid pooling layer, and a Softmax classifier; and the initial deposition monitoring model is constructed based on the optical image branch, the SAR image branch, the cross-feature fusion module, and the multi-task output head. The initial deposition monitoring model is trained based on a preset multi-task joint loss function and the historical image dataset to obtain the deposition monitoring model, wherein the multi-task joint loss function is composed of a deposition regression loss function and a compaction degree classification loss function.
8. The image recognition-based pipe backfill settlement monitoring method of claim 7, wherein, The initial settlement monitoring model is trained according to the preset multi-task joint loss function and the historical image data set, and the settlement monitoring model is obtained, comprising: The historical image data set is divided into a first training data set and a second training data set; According to the first training data set and the multi-task joint loss function, the initial settlement monitoring model is subjected to single-branch alternating training for several rounds to obtain a first settlement monitoring model, wherein, in the process of each single-branch alternating training, the model parameters of the optical image branch are updated through the first training data set, or the model parameters of the SAR image branch are updated through the first training data set. According to the second training data set and the multi-task joint loss function, the first settlement monitoring model is subjected to end-to-end joint training to obtain the settlement monitoring model.
9. A pipe backfill settlement monitoring system based on image recognition, characterized by, It comprises an acquisition module, an image preprocessing module, a settlement monitoring module and a warning module. The acquisition module is used to acquire a plurality of aerial images and a plurality of SAR remote sensing images at each preset time point of a target area. The image preprocessing module is used to perform image preprocessing on the plurality of SAR remote sensing images based on interferometric measurement technology to obtain a plurality of differential interferograms. The settlement monitoring module is used to input the plurality of aerial images and the plurality of differential interferograms into a preset settlement monitoring model, so that the settlement monitoring model extracts elevation change features, ground texture features and color features from the plurality of aerial images based on an optical image branch, extracts ground deformation features from each of the plurality of differential interferograms based on a SAR image branch, and finally performs feature fusion on each feature based on a cross-attention mechanism, and outputs the current settlement progress and compaction degree of the target area according to the fused features. The warning module is used to determine whether the current pipeline backfill settlement is abnormal according to the settlement progress and compaction degree, and to perform corresponding warning based on the determination result. The settlement monitoring model is obtained by training an initial settlement monitoring model based on a labeled historical image data set, and the initial settlement monitoring model is a dual-flow multi-modal feature fusion architecture comprising an optical image branch, a SAR image branch, a cross-feature fusion module and a multi-task output head.
10. The image recognition based pipe backfill settlement monitoring system of claim 9, wherein, The monitoring system further comprises a model training module, and the model training module comprises a data set construction unit, a model construction unit and a training unit. The data set construction unit is used to acquire a plurality of multi-temporal unmanned aerial vehicle aerial images, a plurality of multi-temporal SAR remote sensing images and corresponding monitoring results at each time in a plurality of backfill projects, and construct a labeled historical image data set. The model construction unit is configured to construct the optical image branch according to a computer vision algorithm, an image segmentation algorithm, a residual network and an HSV mapping network; construct the SAR image branch according to a complex convolution network; construct the cross-feature fusion module based on a cross-attention mechanism; construct the multi-task output head based on a global average pooling layer, a full connection layer, a convolution layer, a pyramid pooling layer and a Softmax classifier; and construct the initial settlement monitoring model according to the optical image branch, the SAR image branch, the cross-feature fusion module and the multi-task output head. The training unit is configured to train the initial settlement monitoring model according to a preset multi-task joint loss function and the historical image dataset to obtain the settlement monitoring model, wherein the multi-task joint loss function is composed of a settlement regression loss function and a compaction degree classification loss function.
Citation Information
Patent Citations
Road surface subsidence detection method based on unmanned aerial vehicle DOM and satellite-borne SAR images
CN117968631A
Wide-area remote sensing detection method for highway facilities in alpine region
CN119919813A