Inspection unmanned aerial vehicle non-aligned two-time-phase image intelligent change detection method

By using a lightweight intelligent registration model and image alignment processing, combined with cross-attention feature fusion, the problems of high false detection rate and high computational complexity of non-aligned images in low-altitude UAV image change detection are solved, achieving efficient and accurate change detection and improving the level of monitoring automation.

CN120913115AActive Publication Date: 2025-11-07CHINA RAILWAY DESIGN GRP CO LTD

Patent Information

Application Number
CN202511452977.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2025-11-07
Estimated Expiration
2045-10-13

AI Technical Summary

Technical Problem

Existing technologies for detecting changes in low-altitude UAV images suffer from high false detection rates due to non-aligned images, insufficient robustness, and high computational complexity, making it difficult to simultaneously achieve accuracy, robustness, and computational efficiency.

Method used

A lightweight intelligent registration model and image alignment processing are adopted, combined with cross-attention feature fusion, to construct a lightweight convolutional neural network for image registration and change detection. Through two-stage learning, fast and accurate feature point matching and change region extraction are achieved.

Benefits of technology

It effectively reduced the false detection rate, improved detection accuracy and computational efficiency, enhanced the automation level of low-altitude UAV image monitoring, and reduced maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913115A_ABST
    Figure CN120913115A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent change detection method for non-aligned two-time-phase images of an inspection unmanned aerial vehicle, and relates to the technical field of unmanned aerial vehicle image processing and change detection, and the method comprises the steps: obtaining two-phase images collected by a low-altitude unmanned aerial vehicle under a fixed route and same sensor parameters; a lightweight registration model is constructed and trained, feature point matching is utilized to predict matching point pairs, a homography matrix is calculated, and accurate registration of non-aligned images is achieved; and constructing and training a change detection model based on image pair interaction feature fusion, analyzing the aligned image after registration, and outputting a change information binary image. The method can effectively solve the problem of non-alignment caused by position and angle differences during two-time-phase image acquisition of an unmanned aerial vehicle, and the technical problems of low precision, poor robustness and insufficient calculation efficiency of a traditional method in change detection, effectively improves the automation level of low-altitude safety monitoring of ground highways and railways, reduces the maintenance cost, and improves the safety of the unmanned aerial vehicle. The important application value is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of unmanned aerial vehicle image processing and change detection, in particular to a kind of intelligent change detection method of non-aligned two-phase images of inspection unmanned aerial vehicle. BACKGROUND

[0002] With the rapid development of highway and railway transportation and infrastructure in China, the monitoring and maintenance of the line and the surrounding environment have become a key link to ensure transportation safety. Traditional monitoring methods rely on manual inspection or fixed sensors, which have the disadvantages of limited coverage, low efficiency and high cost. In recent years, low-altitude unmanned aerial vehicles have been widely used in highway and railway environment remote sensing monitoring due to their flexibility, high-resolution image acquisition capability and efficiency. By comparing two-phase unmanned aerial vehicle images, changes in line infrastructure, environmental hazards, invasions and geological disasters can be extracted to provide decision-making basis for safety management.

[0003] However, due to factors such as terrain differences, unmanned aerial vehicle flight attitude, environmental lighting and changes in viewing angle, the two-phase images before and after often have geometric distortion, scale difference and viewing angle deviation (hereinafter referred to as "non-aligned images"), which leads to high false detection rate and insufficient robustness of traditional change detection methods based on pixel-level comparison. Existing non-aligned image change detection methods mainly include single-stage detection based on regional feature alignment and two-stage detection based on image registration. The single-stage method extracts local semantic features for detection, which is difficult to quickly capture complex linear structures and edge information; the two-stage method detects changes after high-precision registration, but is limited by the dynamic environment of the scene (such as vegetation changes, weather interference) and the diversity of non-aligned images, and the computational complexity and robustness of traditional registration algorithms are difficult to meet the actual needs. Although deep learning technology has made significant progress in image processing, there is still a lack of research on change detection for low-altitude inspection unmanned aerial vehicle non-aligned images, and existing methods cannot simultaneously consider image misalignment, complex environmental interference and lightweight requirements.

[0004] Therefore, how to solve the above problems and overcome the limitations of traditional methods in precision, robustness and computational efficiency has become a technical problem that needs to be solved by personnel in the field. SUMMARY

[0005] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, one purpose of the present application is to propose an intelligent change detection method for non-aligned two-phase images of inspection unmanned aerial vehicle, which effectively overcomes the limitations of traditional methods in precision, robustness and computational efficiency through an innovative lightweight intelligent registration model, image alignment processing and an efficient intelligent change detection model, aiming to improve the automation level of ground highway and railway low-altitude safety monitoring and reduce maintenance costs, and has important application value.

[0006] In order to solve the above problems, the application provides a kind of inspection unmanned aerial vehicle non-alignment two phase image intelligent change detection method, comprising the following steps: Step 1, obtain two period images collected by low-altitude unmanned aerial vehicle in fixed route and same sensor parameters; Step 2, select non-aligned images for normalization processing and registration point labeling to form a registration sample set; Step 3, select aligned images with changes in the region for normalization processing, label change information, and form a change detection sample set; Step 4, construct a lightweight registration model based on feature point matching, and train based on the registration point sample set of step 2 to obtain the optimal training weight parameters of the lightweight registration model; Step 5, construct a change detection model based on the interaction feature fusion of image pairs, and train based on the change sample set of step 3 to obtain the optimal training weight parameters of the change detection model; Step 6, after normalization processing of a pair of non-aligned images, input into the registration model constructed in step 4, extract features and predict matching point pairs, calculate homography matrix, and generate aligned images; Step 7, based on the aligned images obtained in step 6, input into the change detection model trained in step 5, output change information binary graph, and overlay the binary graph to the second image in the two period images to mark the change region.

[0007] Preferably, in step 1, a plurality of pairs of two period unmanned aerial vehicle low-altitude image sets S are required, which include first period images and second period images collected based on fixed flight points and at intervals.

[0008] Preferably, in step 2, two period corresponding image sets with view angle difference are selected for normalization processing and registration point labeling to form a registration sample set, and the specific steps are as follows: Step 2-1: scale and fill processing is adopted for scale normalization of the collected images, assuming that the size of the collected images is a × b, Wherein a ≥ b, a is the length of the long side ,b is the length of the short side, and the adaptive scaling factor is defined as a the inverse of the ratio of the normalized scale d , the length of the long side after scaling is , the length of the short side is , and the length of the short side is , the length of the short side is c The calculation formula is: ; ​Step 2-2: Manually label the key feature pixel points in the normalized image pair; form a sample set containing at least 1000 pairs of registration points, covering two-period image scenes with illumination changes and view angle shifts, and the sample labeling information contains the coordinate label of each pair of points.

[0009] Preferably, in step 3, two-period data sets without view angle differences but with different scale changes are selected , the image scale is normalized to 1024x1024x3 in the manner of step 2-1, and then the change information in each pair of images is labeled to generate a binary graph as the change detection data set label.

[0010] Preferably, in step 4, a lightweight convolutional neural network model for two-period image registration is constructed , and the data set constructed in step 2 is used for model training to obtain the registration model optimal weight parameter set , the specific steps are as follows: Step 4-1: input the registration sample set front and back phase images A and B into the lightweight feature extraction network of the registration model to obtain fine feature maps 、 , and semantic feature maps 、 ; Step 4-1-1: shallow feature extraction is performed on the input image using depth separable convolution, batch normalization and SiLU function activation processing; Step 4-1-2: then stack the GhostELA-DWC2f module multiple times for multi-scale feature extraction, first stack the GhostELA-DWC2f module to process the shallow features to obtain fine feature maps 、 , and then stack the GhostELA-DWC2f module twice to obtain semantic feature maps 、 ; Step 4-2: calculate the matching confidence matrix using local cross-attention encoding on the semantic feature maps 、 ; Step 4-2-1: divide 、 into 8x8 windows, traverse each window, and when the window is m , convert its features into query vector , key vector and value vector , the formula is as follows: ; wherein, , , are embedding matrices for transforming query, key and value vectors, respectively; Step 4-2-2: Calculate window cross-attention scores from to and from to respectively: ; ; wherein h is the number of attention heads, and h = 8, C is the number of channels, C = 256, then the output feature representation after cross-attention calculation is: ; ; Step 4-2-3: The confidence of matching within the window is represented as: ; Splice all windows, and the global matching confidence matrix is represented as: ; wherein: R represents the real number field; Step 4-3: Based on the mutual proximity calculation matching rule, use lightweight MLP to predict and screen high-quality matching point pairs; Step 4-3-1: Extract mutual proximity matching information P from the global matching confidence matrix , set to obtain the matching set M, wherein is the mutual proximity matching threshold and , the specific calculation method is as follows:

[0011]

[0012]

[0013] ; Step 4-3-2: Arrange each pair of matches to obtain spliced features , use the confidence MLP predicted by , the calculation method is as follows: ; wherein , is the matching pair, the high-quality matching set is obtained , wherein is the confidence threshold of the high-quality matching pair; Step 4-4: fuse the cross-attention feature map with the backbone network extraction feature; the specific steps are as follows: Step 4-4-1: first, the cross-attention feature map , is calculated using the GhostELA-DWC2f-x module, and the channel splicing fusion is adopted to obtain the fusion feature , ; Step 4-4-2: continue to calculate the fusion feature , using the GhostELA-DWC2f-x module, and then perform channel splicing fusion with the fine feature , to obtain the fused fine feature , ; Step 4-5: map the points in the high-quality matching set selected in step 4-3 to the fused fine feature map , , calculate , the pixel correlation of the corresponding window and the correlation probability , and calculate the expected coordinates as the sub-pixel position value of the mapping area of the fused fine feature map corresponding to the matching point pair ; Step 4-6: using the correlation probability as the weight, the more accurate sub-pixel displacement can be further calculated: ; ; wherein is a weighted function generated based on a Gaussian weight, is a control center weighting intensity, and is taken ; Step 4-7: update the offset of the effective matching points in the registration model training process in this way to obtain the optimal weight parameter set of the registration model .

[0014] Preferably, a convolutional neural network model for change area extraction needs to be constructed in step 5 It consists of three parts: a lightweight backbone network, an efficient Transformer neck network, and a boundary-aware decoding network; the location and category labels of all targets in the segmentation map are input to train the neural network to obtain... , For change detection model The optimal set of weight parameters is determined by the following steps: Step 5-1: Combine the before and after time-phase images of the change detection sample set. and The input is fed into a lightweight backbone network, which uses the same backbone network as the image registration model to obtain semantic features. , ; Step 5-2: Perform channel feature exchange and spatial feature exchange on the semantic features sequentially; , The input is fed into the channel switching module to obtain switching features, and then the GhostELA-DWC2f module is used to extract deep switching features. , Furthermore, the half-channel features are input into the spatial exchange module for exchange at corresponding positions in the spatial dimension. Then, GhostELA-DWC2f is used to extract deep exchange features to obtain the interaction information features. , ; Step 5-3: Input the interactive features into the feature fusion network to obtain fused features with deep semantic and spatial information. ; Use convolution and 2x upsampling structure to extract dimensionality-reduced features , The initial fused features are obtained by using concatenation and fusion. This feature will be channel-wise concatenated and fused with the third layer of the feature extraction network, and then the GhostELA-DWC2f-x module will be used to obtain the first-layer enhanced fusion feature. , ; Continue fusing twice in this manner to obtain the final enhanced fusion feature. ; Step 5-4: Input the fused features into the decoding network for binary mask prediction of changed and unchanged features; first, use bilinear interpolation to align the spatial scale of the input image, then use 1×1 Conv to map the fused features to the number of target categories. N ,in N=2 The dataset includes two categories: changed and unchanged. Then, 3×3Conv is used to enhance local features and improve spatial details. Finally, the SiLU activation function is used to output a probability map of changed pixels. Step 5-5: Construct an effective loss function during model training; Step 5-5-1: Binary cross-entropy loss is used for pixel-level classification to measure the difference between the predicted change mask and the ground truth mask. The pixel classification loss is calculated as follows: wherein y represents the ground truth label, 0 means no change, and 1 means change; represents the probability predicted by the model, which is predicted based on the SiLU function in step 5-4; Step 5-5-2: Boundary loss is constructed to enhance edge accuracy. The edges of the changed region are explicitly optimized by the Sobel operator. The boundary loss is defined as: wherein represents the confidence of the real edge generated using the Sobel operator; represents the Sobel output value of the model predicted edge map; Step 5-5-3: Obtain the total loss, which is a weighted combination of the boundary loss and the pixel classification loss, for model training weight parameter tuning, denoted as: is the weight of the pixel classification loss, represents the weight of the boundary loss; the optimal weight parameter set of the change detection model is obtained after training .

[0015] Preferably, when step 6 is performed, for non-aligned pre- and post-phase test images , input them into the registration model , use the weight parameter set , extract the registration point information of the images relative to , and obtain the aligned images after registration; the specific steps are as follows: Step 6-1: Assuming that the image representing the first period image is , and the second period image is , input them into the configuration model , and output the matching point coordinate pair set in the non-aligned images; Step 6-2: Map the matching point coordinate pair set to the grayscale images of the non-aligned pre- and post-phase test images to calculate the homography transformation matrix , select 4 pairs of matching points , and each point pair satisfies the following homogeneous coordinate relationship: ​​​​​​​​

[0016] ; Will Transform into an unknown parameter vector Then solve the matrix using the least squares method; Step 6-3: Based on the homography transformation matrix Perspective transformation is used to transform the image Mapping to Image Coordinate system; Obtaining a registered image with irregular black borders. ; Step 6-4: Extract the irregular black border areas and fill them into the image. To obtain the current period image Corresponding registration image Overlapping registration images The irregular black border areas are extracted from the locations where the pixel value is 0 in the detected image.

[0017] Preferably, step 7 requires aligning the images. , Input to change detection model Using the optimal set of weight parameters for the change detection model Extract the information of the binarized change regions and label them onto the original image. .

[0018] The advantages of this invention compared to the prior art are: The present invention provides an intelligent change detection method for non-aligned two-phase images from an inspection drone, based on two non-aligned images acquired by a low-altitude drone. The specific advantages are as follows: 1. Compared with the direct change detection in the prior art, the present invention introduces image registration and image alignment processing, which effectively addresses the complex geometric distortion and viewpoint deviation of low-altitude UAV image pairs.

[0019] 2. This invention constructs a lightweight registration model that combines cross-attention, effectively reducing computational complexity; it adopts a two-stage matching point filtering mechanism to achieve accurate and fast feature point matching.

[0020] 3. This invention constructs a multi-temporal high-resolution image change detection network, which can effectively capture medium and small-scale change areas in traffic routes and improve the detection capability of minor safety hazards.

[0021] 4. The application provides a new idea for extracting change information of non-aligned unmanned aerial vehicle images based on two-stage learning of two-phase data, that is, by pre-training an image registration and change detection model for a non-aligned image "matching-alignment-change detection" inference process, so that the false detection rate can be effectively reduced and the accuracy can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0023] Figure 1 The overall framework diagram of the non-aligned low-altitude remote sensing image change detection method of the present application; Figure 2 The lightweight image registration network model structure diagram of the present application; Figure 3 The lightweight layer aggregation feature extraction module principle diagram of the present application; Figure 4 The non-aligned image registration process visualization result schematic diagram of the present application; Figure 5 The change detection network model structure diagram of the present application; Figure 6 The change detection result display diagram of the present application. DETAILED DESCRIPTION

[0024] The embodiments of the present application will be described in detail below, and examples of the embodiments are shown in the drawings, wherein the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present application, and cannot be understood as a limitation of the present application.

[0025] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connecting" should be understood in a broad sense, for example, it can be fixed connection, or detachable connection, or integral connection, it can be mechanical connection, or electrical connection, it can be direct connection, or indirect connection through an intermediate medium, it can be internal communication of two elements or interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0026] The application will be further described in detail below with reference to the accompanying drawings.

[0027] In combination Figure 1 , the intelligent change detection method for non-aligned two-phase images of the inspection unmanned aerial vehicle of the application comprises the following steps: Step 1: Obtain a plurality of pairs of two-period unmanned aerial vehicle low-altitude image sets S, which contain two-period images collected based on fixed flight points and with a certain time interval. Due to the influence of the environment, especially wind, rain and other interference, when the unmanned aerial vehicle collects images, some image pairs collected at fixed flight points will be offset and distorted.

[0028] Step 2: Select two-period corresponding image sets with different view angles , and perform normalization processing and registration point labeling to form a registration sample set.

[0029] Step 2-1: Scale and fill the collected images for scale normalization, assuming that the size of the collected image is a × b, , wherein a ≥ b, a is the length of the long side ,b is the length of the short side, and the adaptive scaling factor is defined as the inverse of the ratio of the normalized scale a to d , the length of the long side after scaling is , and the length of the short side is , and the length of the short side is , and the length of the short side is c , and the length of the short side is , and the length of the short side is ; Step 2-2: Manually label the key ground feature pixel points in the normalized image pair; form a sample set containing at least 1000 pairs of registration points, covering two-period image scenes such as illumination change and view angle offset, and the sample labeling information contains the coordinate label of each pair of points.

[0030] Step 3: Select two-period data sets without view angle difference but with different scale changes , normalize the image scale to 1024x1024x3 in the manner of step 2-1, and then label the change information in each image pair to generate a binary image as a change detection data set label.

[0031] Step 4: Construct a lightweight convolutional neural network model for two-period image registration , as shown in Figure 2 . Use the data set constructed in step 2 to train the model and obtain the optimal weight parameter set of the registration model.

[0032] Step 4-1: input the pre-processed image pair (size 1024x1024x3) of the registration point sample set into the lightweight feature extraction network of the registration model to obtain fine feature maps 、 and semantic feature maps 、 ; Step 4-1-1: shallow feature extraction is performed on the input image using depth separable convolution, batch normalization and SiLU function activation processing (DW-CBS module); Step 4-1-2: then multiple GhostELA-DWC2f modules are stacked for multi-scale feature extraction, and GhostELA-DWC2f is an improved deep feature lightweight extraction module, the structure of which is shown in Figure 3 . Specifically, the GhostELA-DWC2f module is stacked to process the shallow features to obtain fine feature maps 、 , and the semantic feature maps 、 are obtained by stacking twice; Step 4-2: calculate the matching confidence matrix using local cross-attention encoding on the semantic feature maps 、 ; Step 4-2-1: divide 、 into 8x8 windows, traverse each window, and when the window is m , convert its features into query vector , key vector and value vector , the formula is as follows: ; wherein 、 、 are embedding matrices for converting query, key and value vectors, respectively.

[0033] Step 4-2-2: calculate the window cross-attention score from to and from to respectively: ; ; wherein h is the number of attention heads, and h =8, C is the number of channels, C =256, then the output feature is represented as: ; ; Step 4-2-3: The confidence score of the match within the window is represented as follows: ; The global matching confidence matrix, formed by concatenating all windows, is represented as follows: ; This method can significantly reduce the computational load and roughly obtain the matching confidence value between semantic feature maps.

[0034] Step 4-3: Based on the mutual proximity calculation matching rule, use a lightweight method. MLP Predict and filter high-quality matching pairs; Step 4-3-1: Match the global confidence matrix P Extracting nearest neighbor matching information ,set up M in The threshold for mutual nearest neighbor matching and The specific calculation method is as follows:

[0035]

[0036]

[0037] ; Step 4-3-2: For each pair of matches Arrange to obtain splicing features ,use MLP Prediction confidence The calculation method is as follows: ; in , For learnable weight values, retain Get high-quality matching sets from the matching pairs. .

[0038] Step 4-4: Fuse the cross-attention feature map with the features extracted by the backbone network; Step 4-4-1: First, generate the cross-attention feature map. , Use the GhostELA-DWC2f-x module (an improvement on the GhostELA-DWC2f module). Figure 3The calculation is performed using channel concatenation fusion to obtain the fused feature 、 ; Step 4-4-2: Continue to calculate the fused feature 、 using the GhostELA-DWC2f-x module, and then perform channel concatenation fusion with the fine feature 、 to obtain the fused fine feature 、 .

[0039] Step 4-5: Map the points in the high-quality matching set selected in step 4-3 to the fused fine feature map 、 . Taking a matching point pair as an example, calculate the pixel correlation 、 of the corresponding window and the correlation probability , and calculate the expected coordinates as the sub-pixel position value of the mapping area of the fused fine feature map corresponding to the matching point pair . The specific calculation method is as follows:

[0040]

[0041] ; Step 4-6: Using the correlation probability as a weight, a more accurate sub-pixel displacement can be further calculated: ; ; where is a weighting function that can be generated based on a Gaussian weight, to control the weighting strength, and is generally taken as .

[0042] Step 4-7: Update the offset of the effective matching points in the registration model training process in this way to obtain the optimal weight parameter set of the registration model .

[0043] Step 5: Construct a convolutional neural network model for change area extraction , such as Figure 4 ​As shown, it includes three parts of lightweight backbone network, efficient Transformer neck network, and boundary-aware decoding network. The positions and class labels of all segmentation target are input for training neural network to obtain weight parameters ; Step 5-1: input the dual-phase images and into the lightweight backbone network, which uses the same backbone network as the image registration model to obtain semantic features 、 ; Step 5-2: channel feature exchange and spatial feature exchange are performed on the semantic features. Specifically, input 、 into the channel exchange module (Channel Exchange, Figure 4 as shown) to obtain exchange features, and then use the GhostELA-DWC2f module to extract deep exchange features to obtain 、 . Further, input half of the channel features into the spatial exchange module (Spatial Exchange, Figure 4 as shown) to exchange at the corresponding position in the spatial dimension. Then continue to use the GhostELA-DWC2f to extract deep exchange features to obtain interaction information features 、 ; Step 5-3: input the interaction features into the feature fusion network to obtain fusion features with deep semantic information and spatial information . Specifically, use convolution and 2 times upsampling structure to extract dimension reduction features 、 , use splicing fusion and convolution to obtain initial fusion features . The features will be channel spliced and fused with the third layer network of the feature extraction network, and then use the GhostELA-DWC2f-x module to obtain the first layer enhanced fusion features 、 . In this way, continue to fuse twice to obtain the final enhanced fusion features .

[0044] Step 5-4: input the fusion features into the decoding network to predict the binary mask of the changed and unchanged. The decoding network is as shown in the CD-Decoder Head structure of Figure 4 . First, use bilinear interpolation to align the input image space scale, and then use 1x1 Conv to map the fusion features to the target class number N ( N=2, changes and non-changes); then use 3x3 Conv to enhance local features and improve spatial details, and finally use the activation function SiLU to output the probability map of changed pixels.

[0045] Step 5-5: During the model training process, an effective loss function is constructed.

[0046] Step 5-5-1: Binary cross-entropy loss is used for pixel-level classification to measure the difference between the predicted change mask and the real mask , which is calculated as follows: ; where represents the real label, 0 represents non-change, and 1 represents change; represents the probability predicted by the model, which is predicted based on the SiLU function in step 5-4.

[0047] Step 5-5-2: A boundary loss is constructed to enhance the edge accuracy. Specifically, the edges of the changed area are explicitly optimized by the Sobel operator. The boundary loss is defined as: ; where represents the real edge label generated using the Sobel operator; represents the Sobel output value of the model predicted edge map.

[0048] Step 5-5-3: The total loss is a weighted combination of the boundary loss and the pixel classification loss, and the weight parameters for model training are optimized, represented as: ; where , represents the loss weight. After training, the optimal weight parameter set of the change detection model is obtained .

[0049] Step 6: Take the original non-aligned image as an example, input it into the registration model , use the weight parameter set , extract the registration point information of the image relative to , and obtain the registered aligned image.

[0050] Step 6-1: As shown in Figure 4 , assuming that the image representing the first period image is , and the corresponding second period image is , input it into the configuration model Outputs a set of matching point coordinate pairs in the unaligned image, and its visualization result is as follows. Figure 4 As shown in (c); Step 6-2: Map the set of matching point coordinate pairs to the unaligned image. In the grayscale image, to calculate the homography transformation matrix. Select 4 pairs of matching points Each pair of points satisfies the following homogeneous coordinate relationship:

[0051] ; Will Transform into an unknown parameter vector The matrix is ​​solved using the least squares method.

[0052] Step 6-3: Based on the homography transformation matrix Perspective transformation is used to transform the image Mapping to Image Coordinate system. Obtain a registered image with irregular black borders. ( Figure 4 (d)); Step 6-4: Extract the irregular black border areas and fill them into the image. That is, to obtain the current period image Corresponding registration image Overlapping registration images ( Figure 4 (e) Irregular black border areas are extracted by detecting the locations of pixels with a value of 0 in the image.

[0053] Step 7: Align the image , Input to change detection model Using a set of weight parameters Extract the information of the binarized change regions and label them onto the original image. ,like Figure 5 As shown.

[0054] To more clearly illustrate the specific embodiments of the present invention, two examples are provided below: Example 1: Highway Inspection Application: Two sets of UAV imagery were collected for a section of a highway, with a time interval of 6 months. The image resolution for both images is 0.1m, and there are significant differences in rotation and scale.

[0055] First, the image is preprocessed by scaling and padding the original 4000×3000 pixel image to 1024×1024 pixels. Then, a pre-trained registration network is used to extract feature points and calculate the homography matrix to achieve image alignment.

[0056] The aligned images are input into a change detection network to identify areas of change such as newly added cracks and road surface damage.

[0057] Example 2: Railway Inspection Application: A routine inspection of a railway section was conducted, and two sets of drone images were collected. Due to high wind speeds during the data collection, there was a significant misalignment issue between the two sets of images.

[0058] The method of this invention successfully identified safety hazards such as track deviation and slope slippage. Compared with traditional methods, the detection accuracy and efficiency are significantly improved.

[0059] Finally, for any aspects of this invention not fully described, mature technical means from the prior art are employed.

[0060] The following is a list of definitions for formula characters in this invention:

[0061]

[0062]

[0063]

[0064] This invention is not only applicable to road and railway inspection, but can also be extended to power line inspection, pipeline monitoring, and other fields. Any equivalent substitutions or modifications made based on the essential spirit of this invention should be covered within the scope of protection of this invention.

[0065] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A non-aligned two-phase image intelligent change detection method for inspection of unmanned aerial vehicles, characterized in that, Comprise the following steps: Step 1, obtain a pair of images collected by a low-altitude unmanned aerial vehicle at fixed routes and with the same sensor parameters; Step 2, select non-aligned images for normalization processing and registration point labeling to form a registration sample set; Step 3, select aligned images with changed regions for normalization processing and change information labeling to form a change detection sample set; Step 4, construct a lightweight registration model based on feature point matching, train the model based on the registration point sample set of step 2, and obtain the optimal training weight parameters of the lightweight registration model; Step 5, construct a change detection model based on the fusion of interactive features of an image pair, train the model based on the change sample set of step 3, and obtain the optimal training weight parameters of the change detection model; Step 6, input the normalized pair of non-aligned images into the registration model constructed in step 4, extract features and predict matching point pairs, calculate the homography matrix, and generate aligned images; Step 7, input the aligned images obtained in step 6 into the change detection model trained in step 5, output a change information binary graph, and overlay the binary graph on the second image in the two images to mark the changed region.

2. The method according to claim 1, wherein the method is characterized in that: In step 1, a plurality of pairs of two-period unmanned aerial vehicle low-altitude image sets S need to be obtained, which include a first-period image and a second-period image collected based on fixed flight points and at a certain time interval.

3. The method according to claim 1, wherein the method further comprises: The two corresponding image sets with visual angle difference in step 2 are selected The normalization processing and the registration point labeling are performed to form a registration sample set, and the specific steps are as follows: Step 2-1: Perform scale normalization on the acquired image using adaptive scaling and padding. Assume the size of the acquired image is... a × b, in a ≥ b, a Length of the longer side ,b For the shorter side length, the adaptive scaling factor is defined as 'a' and the normalized scale. d reciprocal of the ratio The length of the longer side after scaling is The length of the shorter side is To convert to a normalized scale, double-sided short-side zero-filling is used, resulting in single-sided short-side fill pixels. c The calculation formula is: ; Step 2-2: Manually label the key ground feature pixel points in the normalized image pair; form a sample set containing at least 1000 registration points, covering two-period image scenes with light changes and view angle shifts, and the sample labeling information includes the coordinate labels of each pair of points.

4. The method according to claim 3, wherein the method further comprises: The two data sets in step 3 need to be selected without visual angle difference but with different scale changes , the image scale is normalized to 1024*1024*3 in the manner of step 2-1, and then the change information in each pair of images is labeled to generate a binary graph as a change detection data set label.

5. The method according to claim 1, wherein the method further comprises: A lightweight convolutional neural network model for two-phase image registration needs to be constructed in step 4 , and the registration model is obtained by using the data set constructed in step 2 for model training The optimal weight parameter set The specific steps are as follows: Step 4-1: input the pre- and post-phase images A, B of the registration sample set into the lightweight feature extraction network of the registration model to obtain fine feature maps 、 , and semantic feature maps 、 ; Step 4-1-1: Perform shallow feature extraction on the input image using depth separable convolution, batch normalization, and SiLU function activation processing; Step 4-1-2: Then multiple times of superimposing the GhostELA-DWC2f module are performed for multi-scale feature extraction, first superimposing the GhostELA-DWC2f module to process the shallow layer features to obtain fine feature maps 、 , and then superimposing twice the GhostELA-DWC2f module to obtain semantic feature maps 、 ; Step 4-2: Semantic feature map 、 The local cross-attention encoding is used to calculate the matching confidence matrix. Step 4-2-1: Divide the image into 8x8 windows, traverse each window, and convert its features when the window is , m ​​ To query the vector , the key vector and the value vector , the formula is as follows: ; wherein, , , are embedding matrices for transforming the query, key and value vectors, respectively; Step 4-2-2: Calculate the window cross-attention scores from to and from to respectively: ; ; In the formula h is the number of attention heads, and h = 8, C is the number of channels, C = 256, then the output feature after cross-attention calculation is represented as: ; ; Step 4-2-3: The confidence of the matching in the window is represented as: ; The global matching confidence matrix is represented by concatenating all the windows: ; In the formula, R represents the real number field; Step 4-3: Using lightweight MLP Predict and filter high-quality matching point pairs; Step 4-3-1: Extract mutual nearest neighbor matching information from global matching confidence matrix P Obtain matching set M, where is a mutual nearest neighbor matching threshold and The specific calculation method is as follows:​​ ; Step 4-3-2: Permutation of each pair of matches Obtaining stitching features from permutations , using MLP the confidence of the prediction , the calculation method is as follows: ; wherein , are learnable weight values, retaining matching pairs, obtaining a high-quality matching set wherein is a confidence threshold for high-quality matching pairs; Step 4-4: Fuse the cross-attention feature map with the backbone network extracted features; the specific steps are as follows: Step 4-4-1: First, cross attention feature maps are obtained by using the cross attention module , , ;​ Step 4-4-2: Continue to fuse features with , GhostELA-DWC2f-x module, and then with fine features , Channel splicing fusion, get fused fine features , ; Step 4-5: mapping the points in the high quality match set screened in step 4-3 to the fused fine feature map , , the pixel correlation of the corresponding window and the correlation probability , and calculating the expected coordinates as the sub-pixel position value of the mapping area of the fused fine feature map corresponding to the matching point pair ;​​ Step 4-6: Use the correlation probability as the weight to further calculate the more accurate sub-pixel displacement: ; ; wherein is a weighting function generated based on Gaussian weights, is a control center weighting strength, taken as ; Step 4-7: Obtain the optimal weight parameter set of the registration model by updating the offset of the effective matching points in the registration model training process in this way .

6. The method according to claim 1, wherein the method further comprises: The step 5 needs to construct a convolutional neural network model for change area extraction , including lightweight backbone network, efficient transform neck network, boundary perception decoding network three parts; input the position and category label of all segmentation graph targets for training neural network to obtain , Change detection model The optimal weight parameter set, and the specific steps are as follows: Step 5-1: input the pre- and post-phase images of the change detection sample set into the lightweight backbone network, which uses the same backbone network as the image registration model to obtain semantic features and input into the lightweight backbone network, which uses the same backbone network as the image registration model to obtain semantic features 、 ; Step 5-2: Perform channel feature exchange and spatial feature exchange on the semantic features sequentially; , The input is fed into the channel switching module to obtain switching features, and then the GhostELA-DWC2f module is used to extract deep switching features. , Furthermore, the half-channel features are input into the spatial exchange module for exchange at corresponding positions in the spatial dimension. Then, GhostELA-DWC2f is used to extract deep exchange features to obtain the interaction information features. , ; Step 5-3: input the interaction feature into the feature fusion network to obtain a fusion feature with deep semantic information and spatial information ; using convolution and 2 times upsampling structure to extract dimension reduction feature , , using splicing fusion and convolution to obtain initial fusion feature , the feature will be channel splicing fusion with the third layer network of the feature extraction network, and then GhostELA-DWC2f-x module is adopted to obtain the first layer enhanced fusion feature , ; in this way, the fusion is continued twice to obtain the final enhanced fusion feature ; Step 5-4: The fused features are input into the decoding network to predict the binary mask of changed and unchanged; first, align the input image space scale using bilinear interpolation, and then map the fused features to the target class number using 1x1 Conv N wherein N= 2 including both changed and unchanged categories; Then use 3x3Conv to enhance local features and improve spatial details, and then use the activation function SiLU to output the change pixel probability graph; Step 5-5: During model training, an effective loss function is constructed; Step 5 - 5-1: Binary cross-entropy loss is used for pixel-wise classification to measure the difference between the predicted change mask and the ground truth mask The pixel classification loss is calculated as follows: ; wherein, represents the true label, 0 represents no change, and 1 represents change; represents the probability predicted by the model, the value is predicted based on the SiLU function of step 5-4; Step 5-5-2: A boundary loss is constructed to enhance edge accuracy; the edges of the changed regions are explicitly optimized by a Sobel operator; the boundary loss is defined as: ; wherein represents the confidence of the true edge generated using the Sobel operator; represents the Sobel output value of the model predicted edge map; Step 5-5-3: Obtain the total loss, which is a weighted combination of the boundary loss and the pixel classification loss, used for model training weight parameter optimization, represented as: ; wherein is a weight for the pixel classification loss, is a weight for the boundary loss; and the optimal weight parameter set of the change detection model is obtained after training .

7. The method according to claim 1, wherein the method further comprises: When step 6 is performed, for the unaligned before and after phase test images Input it into the registration model Using a set of weight parameters Extracting images Compared to The registration point information is obtained, and the aligned image after registration is acquired; the specific steps are as follows: Step 6-1: Assume the image representing the first phase image is , the second phase image is , input to the configuration model , output a set of matching point coordinate pairs in the non-aligned images; Step 6-2: Map the set of matching point coordinate pairs to the non-aligned pre- and post- temporal test images in the grayscale images to compute the homography transformation matrix , selecting 4 matching point pairs , each point pair satisfying the following homogeneous coordinate relationship: ; Transforming into an unknown parameter vector and solving for the matrix by least squares Step 6-3: Homography transformation matrix based perspective transformation to map the image to the image coordinate system; obtaining a registered image with irregular black borders ; Step 6-4: Extracting irregular black border area and filling it to the image to obtain the current period image corresponding registration image of the overlapping registration image ; wherein the irregular black border area is extracted by detecting the position of the pixel value of 0 in the image.

8. The method according to claim 1, wherein the method is characterized in that: The step 7 needs to align the image , to the change detection model , using the change detection model optimal weight parameter set , extracting the binary change region information, and marking it to the original image .

Citation Information

Patent Citations

  • Remote sensing image binary change detection method based on feature deviation alignment

    CN113378727A

  • Unmanned aerial vehicle change detection method and system based on multi-task learning

    CN115546671A

  • Remote sensing image change detection method, electronic equipment and storage medium

    CN118411620A

  • Double-branch multi-mode remote sensing image registration method based on multi-scale and multi-direction features

    CN120355758A

  • Remote sensing image change detection method based on semantic guidance and SAM optimization

    CN120451664A

Cited By

  • Railway multi-temporal change detection method and system based on cross-domain generated image

    CN121708515A

  • Geological disaster change area rapid screening method based on multi-temporal unmanned aerial vehicle image

    CN122244738A

  • A rapid method for screening geological hazard change areas based on multi-temporal UAV imagery

    CN122244738B