Method and System for Wire Registration between UAV Aerial Images Based on Deep Learning
The wire registration of drone aerial image by deep learning methods solves the accuracy and robustness of wire registration in complex scenarios, and realizes high-precision automatic wire registration, which is suitable for intelligent inspection of power lines in complex scenarios.
Patent Information
- Application Number
- CN202411775919.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-05
AI Technical Summary
The prior art wire registration method in aerial images of drones is difficult to cope with complex light changes, viewing angle differences and wire deformation, resulting in insufficient or incorrect extraction of feature points. The traditional method ignores the slender structure and coherence of the wires, making it difficult to achieve accurate matching.
Using a deep learning-based method, a multi-constraint matching matrix is constructed through adaptive dynamic range compression, nonlinear multi-scale feature pyramid, spatially perceived self-attention processing, multi-layer mutual attention network and bidirectional consistency inspection to achieve high-quality extraction and accurate registration of wire features.
Effectively deal with lighting changes and viewing angle differences in complex scenarios, improves the accuracy of wire detection and registration, reduces the computational complexity, and achieves high-precision automatic registration.
Smart Images

Figure CN119672078B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and artificial intelligence, and in particular, to a wire registration method and system for UAV aerial images based on deep learning. Background Art
[0002] In recent years, with the rapid economic development and the acceleration of the urbanization process, the demand for electricity has increased sharply, posing higher requirements for the construction and maintenance of the power grid. In the construction of the power grid, wires are the hubs connecting the power industry and users, and the safety of power lines is of crucial importance. However, power lines span vast areas and often pass through complex terrains and environments, which brings huge challenges to the construction, monitoring, and maintenance of the lines. The traditional manual inspection method is not only time-consuming and laborious but also has safety hazards in bad weather or remote areas, and can no longer meet the needs of modern power grid management. The safe and stable operation of the power system is directly related to the development of the national economy and the stability of social life, and the intelligent inspection of transmission lines is an important means to ensure the safe operation of the power grid.
[0003] Under this background, UAV technology has emerged and demonstrated great potential in the power grid industry with its flexibility and efficiency. UAVs can quickly fly over complex terrains, conduct high-precision aerial photography of power lines, and collect a large amount of image data, providing an unprecedented perspective for the risk assessment of potential threats posed by elements around the wires to the safe operation of power lines. This method not only greatly improves the inspection efficiency and reduces the labor cost but also can obtain more comprehensive and clear line status information. However, since UAVs take pictures of the same transmission line segment at different times and from different angles, how to accurately achieve the automatic registration of wires between different images and then realize the dynamic monitoring and change analysis of wire status has become a key scientific problem to be solved urgently. Accurate wire registration is not only the basis for realizing the intelligent inspection of transmission lines but also an important prerequisite for subsequent fault diagnosis such as insulator shedding, fitting corrosion, and line icing.
[0004] Currently, the research on wire registration mainly focuses on methods based on traditional image processing and machine learning. Traditional methods usually use feature operators such as SIFT and SURF to extract local feature points, and combine algorithms such as RANSAC for feature matching and geometric verification. Some researchers have also proposed using the Hough transform to detect line features or using local binary patterns (LBP) to describe wire texture features. In terms of machine learning, some methods use support vector machines (SVM) or random forests to classify the extracted features to improve the accuracy of matching. At the same time, some researchers have also tried to use graph matching methods, modeling wires as graph structures and realizing registration through graph matching algorithms. These methods can achieve certain effects under ideal conditions, but there are still many limitations in actual complex scenarios.
[0005] Specifically, the existing methods mainly have the following technical problems: First, traditional feature extraction methods are difficult to cope with the complex illumination changes in UAV aerial images. Especially under backlight and strong light conditions, problems such as insufficient or mis-extraction of feature points are likely to occur. Second, the existing feature description methods do not fully consider the slender structural characteristics of wires, resulting in a lack of effective description of the wire geometry in the extracted features. Especially when the wire twists or deforms, the discriminability of the feature description significantly decreases. Third, the information fusion and analysis of wires among multiple images, especially when the wires perform differently under different angles and illumination conditions, how to accurately identify and match the wires has become a key problem to be solved urgently. Finally, traditional image processing methods, such as feature point-based matching algorithms, when dealing with slender and continuous objects like wires, mostly focus on the matching of single key points, and often the effect is not good. These methods ignore the importance of the coherence of the wire as a whole and the spatial layout. Currently, the matching of wires between images is essentially a matching of a set of points, which pays more attention to the similarity and difference between different groups, and is more difficult than the matching of single points. These technical problems severely restrict the application effect of the wire registration algorithm in practical engineering. Summary of the Invention
[0006] Object of the Invention: To propose a method and system for wire registration between UAV aerial images based on deep learning to solve the above problems existing in the prior art.
[0007] Technical Solution: A method for wire registration between UAV aerial images based on deep learning includes the following steps:
[0008] S1. Obtain at least two original UAV aerial images containing overlapping regions, and perform adaptive dynamic range compression processing on them to obtain preprocessed images; based on the preprocessed images, construct a non-linear multi-scale feature pyramid, extract wire features of each scale layer to obtain an enhanced wire feature set; based on the enhanced wire feature set, construct a fusion feature matrix through dynamic feature fusion;
[0009] S2. Based on the fusion feature matrix, through spatial-aware self-attention processing, obtain a position-aware feature matrix; based on the position-aware feature matrix, calculate enhanced self-attention scores to obtain fusion feature vectors between different wires;
[0010] S3. Based on the fusion feature vectors, construct a multi-layer mutual attention network to obtain a multi-layer attention weight set; based on the multi-layer attention weight set, perform feature aggregation to obtain cross-image fusion features; based on the cross-image fusion features, through bidirectional consistency checking, obtain consistency-enhanced features;
[0011] S4. Based on the consistency-enhanced features, construct a multi-constraint matching matrix to obtain an initial matching matrix; optimize the initial matching matrix to obtain an optimized matching matrix.
[0012] S5. Based on the optimized matching matrix, calculate the matching reliability score to obtain a reliability scoring matrix; based on the reliability scoring matrix and a preset threshold, screen to obtain the final matching result.
[0013] A wire registration system between UAV aerial images based on deep learning, comprising:
[0014] At least one processor; and,
[0015] A memory communicatively connected to at least one of the processors; wherein,
[0016] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned wire registration method between UAV aerial images based on deep learning.
[0017] Advantageous effects: The present invention can effectively handle various complex situations in the aerial photography scene, such as light changes, perspective differences, local deformations, etc. At the same time, it also improves the accuracy of wire detection and registration, reduces the computational complexity, improves the processing efficiency, and realizes the high-precision automatic registration of wires between UAV aerial images. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of the method of the present invention.
[0019] Figure 2 It is a flowchart of step S1 of the present invention.
[0020] Figure 3 It is a flowchart of step S2 of the present invention.
[0021] Figure 4 It is a flowchart of step S3 of the present invention.
[0022] Figure 5 It is a flowchart of step S4 of the present invention.
[0023] Figure 6 It is a flowchart of step S5 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] As Figure 1 shown, the present application proposes a wire registration method between UAV aerial images based on deep learning, including the following steps:
[0025] S1. Obtain at least two original UAV aerial images containing overlapping regions, and perform adaptive dynamic range compression processing on them to obtain preprocessed images; based on the preprocessed images, construct a non-linear multi-scale feature pyramid, extract wire features at each scale layer to obtain an enhanced wire feature set; based on the enhanced wire feature set, construct a fusion feature matrix through dynamic feature fusion;
[0026] S2. Based on the fusion feature matrix, through spatial-aware self-attention processing, obtain a position-aware feature matrix; based on the position-aware feature matrix, calculate enhanced self-attention scores to obtain fusion feature vectors between different wires;
[0027] S3. Based on the fusion feature vectors, construct a multi-layer mutual attention network to obtain a multi-layer attention weight set; based on the multi-layer attention weight set, perform feature aggregation to obtain cross-image fusion features; based on the cross-image fusion features, through bidirectional consistency checking, obtain consistency-enhanced features;
[0028] S4. Based on the consistency-enhanced features, construct a multi-constraint matching matrix to obtain an initial matching matrix; optimize the initial matching matrix to obtain an optimized matching matrix;
[0029] S5. Based on the optimized matching matrix, calculate matching reliability scores to obtain a reliability scoring matrix; based on the reliability scoring matrix and a preset threshold, screen to obtain the final matching result.
[0030] As Figure 2 shown, according to one aspect of the present application, step S1 is further as follows:
[0031] S11. Obtain at least two original UAV aerial images containing overlapping regions; based on the image quality of the original UAV aerial images, evaluate to obtain an image quality feature vector; based on the image quality feature vector, perform adaptive enhancement to obtain an enhanced image; based on the enhanced image, perform illumination equalization and noise suppression to obtain a balanced image; based on the balanced image, perform edge feature enhancement to obtain an edge-enhanced image; based on the edge-enhanced image, perform adaptive dynamic range compression to obtain a preprocessed image;
[0032] S12. Based on the preprocessed image, construct an improved Laplacian pyramid; based on the Laplacian pyramid, perform compression operations on each layer to obtain compressed feature maps; based on the compressed feature maps, perform pyramid decomposition operations to obtain Laplacian feature maps; based on the Laplacian feature maps, generate a multi-scale feature pyramid;
[0033] S13. Based on the multi-scale feature pyramid, extract local structural features to obtain a structural feature matrix; based on the structural feature matrix, calculate the direction consistency feature to obtain a direction feature set; based on the direction feature set, extract and fuse context information to obtain context-enhanced features; based on the context-enhanced features, integrate to obtain a multi-dimensional feature tensor, and finally generate an enhanced wire feature set;
[0034] S14. Based on the enhanced wire feature set, calculate the dynamic fusion weight; based on the dynamic fusion weight, perform a weighted summation operation to obtain a fusion feature matrix.
[0035] In an embodiment of the present application, two original UAV aerial images IA and IB containing an overlapping area are obtained and enhanced to obtain edge-enhanced images, and an adaptive dynamic range compression operation is performed on each pixel value x: y = α·log(1 + βx) / log(1 + β), where α is the compression coefficient, β is the dynamic adjustment factor, and y is the pixel value after the adaptive dynamic range compression operation, to obtain preprocessed images IA' and IB'.
[0036] Obtain the preprocessed images IA' and IB', construct an improved Laplacian pyramid, and perform a compression operation on each layer k to obtain Gk+1 = x·tanh(γx), where γ is a learnable parameter, and Gk+1 is the feature map after the compression operation on the (k + 1)-th layer of the Laplacian pyramid; perform a pyramid decomposition operation to obtain Lk = Gk - Expand(Gk+1), where Expand is an upsampling expansion operation, and finally obtain multi-scale feature pyramids PA and PB.
[0037] Obtain the multi-scale feature pyramids PA and PB, extract wire features at each scale layer. For M wires LA in image IA and N wires LB in image IB, each wire is represented as an ordered point set Li = {pi1, pi2,..., pin}, and each point contains spatial coordinates (x, y), scale response s, and direction feature o, to obtain enhanced wire feature sets FA and FB.
[0038] Obtain the enhanced wire feature sets FA and FB, calculate the dynamic fusion weight wi = softmax(νi·tanh(μFi)), where νi and μ are learnable parameters, and Fi represents the feature vector extracted from the enhanced wire feature set; perform a weighted summation operation Z = Σ(wi·Fi) to obtain a fusion feature matrix Z with dimensions M×N×D.
[0039] In this embodiment, through multi-stage adaptive image preprocessing and non-linear multi-scale feature extraction, high-quality extraction and enhancement of wire features in UAV aerial images are achieved. Specifically, the adaptive adjustment mechanism of the compression coefficient and the dynamic adjustment factor in the adaptive dynamic range compression operation enables the image to obtain the best dynamic range display under different lighting and contrast conditions; the improved Laplacian pyramid enhances the feature retention ability for slender targets such as wires by introducing a compression operation with learnable parameters, avoiding the problem of easy detail loss in traditional pyramid decomposition; during the wire feature extraction process, by representing each wire as an ordered point set containing spatial coordinates, scale response, and direction features, both the geometric features of the wire and its local morphological information are retained; finally, through the design of dynamic fusion weights, adaptive integration of different features is achieved. This embodiment can accurately extract and enhance wire features in complex aerial photography scenes, laying a solid foundation for subsequent feature matching and improving the accuracy and robustness of subsequent registration.
[0040] According to one aspect of the present application, step S11 is further as follows:
[0041] S111. Obtain at least two original UAV aerial images containing overlapping regions, and calculate the image sharpness feature, illumination uniformity feature, and signal-to-noise ratio feature; combine the image sharpness feature, illumination uniformity feature, and signal-to-noise ratio feature to obtain an image quality feature vector;
[0042] S112. Based on the original UAV aerial image and its corresponding image quality feature vector, construct an adaptive enhancement function; based on the adaptive enhancement function, perform an enhancement operation on each pixel to obtain an enhanced image;
[0043] S113. Based on the enhanced image, construct an adaptive weight matrix, perform a weighted smoothing operation to obtain a smoothed image; based on the smoothed image, perform local contrast correction to obtain a balanced image;
[0044] S114. Based on the balanced image, calculate the multi-directional gradient response, construct a direction enhancement kernel; based on the direction enhancement kernel, perform a direction selective convolution operation to obtain an edge enhanced image;
[0045] S115. Based on the edge enhanced image, calculate the local statistical feature, construct an adaptive compression coefficient; based on the adaptive compression coefficient, calculate the dynamic adjustment factor; based on the edge enhanced image and the dynamic adjustment factor, perform a pixel-level adaptive compression operation to obtain a preprocessed image.
[0046] In an embodiment of the present application, two original UAV aerial images IA and IB containing an overlapping area are obtained, and the image sharpness feature fc = Σ|▽I(x, y)| / N is calculated, where ▽I(x, y) is the image gradient and N is the total number of pixels; the illumination uniformity feature fl = std(I) / mean(I) is calculated, where std is the standard deviation operator and mean is the mean operator; the signal-to-noise ratio feature fn = 10·log10(P_signal / P_noise) is calculated, where P_signal is the signal power and P_noise is the noise power; the three features are combined to obtain the image quality feature vector Q = [fc, fl, fn].
[0047] The original UAV aerial images IA and IB and their corresponding image quality feature vectors Q are obtained, and an adaptive enhancement function E(x) = x +λ·(1-Q)·sinh(γx) is constructed, where x is the input pixel value, λ is the enhancement intensity coefficient dynamically adjusted by the sharpness feature fc, and γ is the non-linear coefficient adjusted by the illumination uniformity feature fl; the enhancement operation is performed on each pixel to obtain the enhanced images IE_A and IE_B.
[0048] The enhanced images IE_A and IE_B are obtained, and an adaptive weight matrix W(x, y) = exp(-||I(x, y)-μ||2 / 2σ 2 ) is constructed, where I(x, y) is the pixel value, μ is the local mean, σ is adjusted by the signal-to-noise ratio feature fn, and || ||2 represents the Euclidean norm; the weighted smoothing operation Is(x, y) =ΣW(i, j)·I(i, j) / ΣW(i, j) is performed, where (i, j) are the local window coordinates; the local contrast correction Ic(x, y) = Is(x, y)·(μg / μl)η is applied, where μg is the global mean, μl is the local mean, and η is the adaptive adjustment factor, to obtain the balanced images IB_A and IB_B.
[0049] The balanced images IB_A and IB_B are obtained, and the multi-directional gradient response Gθ(x, y)=cos(θ)·ΨI / Ψx+ sin(θ)·ΨI / Ψy is calculated, where θ is the gradient direction angle and Ψ is the partial derivative; the direction enhancement kernel Kθ=exp(-d 2 / 2ρ 2 )·cos(θ) is constructed, where d is the spatial distance and ρ is the spatial scale parameter, and exp is the exponential function; the direction selective convolution operation Ie(x, y) = max_θ{I(x, y) * Kθ} is performed, where * represents the convolution operation, to obtain the edge enhanced images IEE_A and IEE_B.
[0050] Obtain the edge-enhanced images IEE_A and IEE_B, calculate the local statistical features μl(x, y) and σl(x, y); construct the adaptive compression coefficient α(x, y) = α0·(1 + k·σl(x, y) / σg), where α0 is the base compression coefficient, k is the adjustment coefficient, and σg is the global standard deviation; calculate the dynamic adjustment factor β(x, y) = β0·exp(-|μl(x, y) - μg| / σg), where β0 is the base adjustment factor; perform pixel-level adaptive compression operation y(x, y) = α(x, y)·log(1 + β(x, y)·x(x, y)) / log(1 + β(x, y)) to obtain the preprocessed images IA' and IB'.
[0051] In this embodiment, through a multi-level adaptive image enhancement and preprocessing strategy, high-quality optimization of aerial images in complex scenarios is achieved. By calculating image sharpness features, illumination uniformity features, and signal-to-noise ratio features, a complete image quality evaluation system is constructed; based on these features, an adaptive enhancement function is designed to achieve targeted image enhancement; local weighted smoothing and contrast correction are performed through an adaptive weight matrix, effectively suppressing noise while maintaining edge details; finally, through pixel-level adaptive compression operation, optimal adjustment of the dynamic range is achieved. This embodiment improves the quality of aerial images, especially showing excellent adaptability and stability when dealing with complex scenarios such as uneven illumination, low contrast, and noise interference.
[0052] According to one aspect of the present application, step S13 is further as follows:
[0053] S131. Based on the multi-scale feature pyramid, construct a structure tensor at each scale layer, calculate the eigenvalues and eigenvectors; based on the eigenvalues and eigenvectors, construct a local structure descriptor; based on the local structure descriptor, perform feature aggregation on each detection point to obtain a structure feature matrix;
[0054] S132. Based on the structure feature matrix, calculate the direction vector field, construct a direction consistency metric; based on the direction consistency metric, perform non-maximum suppression to obtain the local principal direction; based on the local principal direction, adopt a tensor voting mechanism to enhance the direction coherence to obtain a direction feature set;
[0055] S133. Based on the direction feature set, construct an adaptive sampling window, extract local context features; based on the local context features, apply non-local mean filtering to obtain context-enhanced features;
[0056] S134. Based on the context-enhanced features, construct a multi-dimensional feature descriptor, calculate the feature importance weights; based on the feature importance weights, perform weighted feature fusion to obtain a multi-dimensional feature tensor;
[0057] S135. Based on the multi-dimensional feature tensors, extract a predetermined number of key points for each wire, and calculate the feature representation of each key point; based on the feature representation of each key point, perform feature normalization and dimensionality reduction to obtain an enhanced wire feature set.
[0058] In an embodiment of the present application, obtain multi-scale feature pyramids PA and PB, and construct a structure tensor T(x, y) = [Ix 2 , IxIy; IxIy, Iy 2 at each scale layer k, where Ix and Iy are image gradients; calculate the eigenvalues λ1, λ2 and eigenvectors v1, v2; construct a local structure descriptor s(x, y) = [λ1, λ2, v1·v2, trace(T)], where trace is the trace of the matrix; perform feature aggregation on each detection point to obtain structure feature matrices SMA and SMB.
[0059] Obtain structure feature matrices SMA and SMB, calculate the direction vector field ν(x, y) = λ1v1, where λ1 is the main eigenvalue and v1 is the main eigenvector; construct a direction consistency measure c(x, y) = Σw(i, j)·|ν(x, y)·ν(i, j)| / ||ν(x, y)||·||ν(i, j)||, where w(i, j) is the spatial weight; perform non-maximum suppression to obtain the local main direction; use the tensor voting mechanism to enhance the direction coherence to obtain direction feature sets DA and DB.
[0060] Obtain direction feature sets DA and DB, and construct an adaptive sampling window W(x, y) = exp(-d 2 / 2σ 2 )·(1 + κ·c(x, y)), where d is the spatial distance, σ is the scale parameter, and κ is the direction consistency adjustment factor; extract the local context feature h(x, y) = Σ(W(i, j)·f(i, j)), where f(i, j) is the original feature; apply non-local mean filtering to enhance the feature consistency to obtain context-enhanced features CA and CB.
[0061] Obtain context-enhanced features CA and CB, and construct a multi-dimensional feature descriptor m(x, y) = [s(x, y), d(x, y), h(x, y)], where s is the structure feature, d is the direction feature, and h is the context feature; calculate the feature importance weight w = softmax(φ(m)), where φ is the feature mapping function; perform weighted feature fusion to obtain multi-dimensional feature tensors MA and MB.
[0062] Obtain multi-dimensional feature tensors MA and MB, and extract n key points {pi1, pi2, ..., pin} for each wire Li; calculate the feature representation fij = [xij, yij, sij, oij, mij] for each point, where xij and yij are spatial coordinates, sij is the scale response, oij is the orientation feature, and mij is the multi-dimensional feature; perform feature normalization and dimensionality reduction to obtain enhanced wire feature sets FA and FB.
[0063] This embodiment realizes the all-round characterization and enhancement of wire features based on multi-dimensional feature extraction and fusion. Local structure features are extracted through structure tensor analysis, eigenvalues and eigenvectors are calculated, and a complete structure descriptor is constructed; the direction coherence is enhanced through the direction vector field and tensor voting mechanism, effectively capturing the continuity features of the wires; context features are extracted using an adaptive sampling window, and the feature consistency is enhanced through non-local mean filtering; finally, through multi-dimensional feature fusion and normalization processing, a comprehensive feature representation containing spatial position, scale response, orientation feature, and structure information is generated. This embodiment not only improves the discriminability and robustness of the features but also shows strong adaptability when dealing with abnormal situations such as wire breaks and deformations.
[0064] As Figure 3 shown, according to one aspect of the present application, step S2 is further as follows:
[0065] S21. Obtain the wire feature set of the same image in the fusion feature matrix, and calculate the position encoding; based on the position encoding, perform a feature transformation operation to obtain a position-aware feature matrix;
[0066] S22. Based on the position-aware feature matrix, construct a spatial constraint matrix, calculate the attention score, and obtain a fusion feature vector.
[0067] In an embodiment of the present application, obtain the wire feature set Li of the same image in the fusion feature matrix Z, and calculate the position encoding PE(Li) = [sin(ωpdi), cos(ωpdi)], where di is the spatial distribution feature of the wire and ωp is the position encoding frequency parameter; perform feature transformation operations K = WKf(Li) + PE(Li), Q = WQf(Li) + PE(Li), and V = WVf(Li), where WK, WQ, and WV are learnable transformation matrices and f(Li) is the feature representation of wire Li, to obtain a position-aware feature matrix F.
[0068] Obtain the position-aware feature matrix F, and construct a spatial constraint matrix Mij = exp(-||pi - pj||2 / σ 2), where pi and pj are the coordinates of the wire endpoints, and σ is the spatial scale parameter; calculate the attention score A(Q, K, V) = softmax(QKT / sqrt(d) + M)·V, where d is the feature dimension and M is the spatial constraint matrix, to obtain the fused feature vector f'.
[0069] In this embodiment, through the spatial-aware self-attention mechanism, the position perception and context relationship modeling of wire features are realized. First, by calculating the position encoding, the spatial position information of the wire is effectively encoded into the feature representation; then, the features are mapped through a learnable transformation matrix, and the position-aware features are generated in combination with the position encoding; on this basis, a spatial constraint matrix is introduced to modulate the attention calculation, enabling the attention mechanism to better capture the spatial relationship between wires. This embodiment enables the system to accurately grasp the spatial distribution and relative position relationship of wires, effectively solves the difficulties of traditional methods in dealing with large-range deformations and perspective changes, and at the same time, through the adaptive feature fusion mechanism, improves the discriminability of feature expression, enabling the system to better handle the matching problem of wires between different images.
[0070] As Figure 4 shown, according to one aspect of the present application, step S3 is further as follows:
[0071] S31. Based on the fused feature vectors of different images, construct a multi-scale feature representation to obtain a scale feature set; based on the scale feature set, calculate the spatial dependence relationship to obtain a spatial correlation matrix; based on the spatial correlation matrix, perform an adaptive feature transformation to obtain a transformed feature matrix; based on the transformed feature matrix, calculate the cross-scale attention to obtain an inter-layer attention tensor; based on the inter-layer attention tensor, generate a multi-layer attention weight set;
[0072] S32. Based on the multi-layer attention weight set, perform a dynamic weight fusion operation to obtain a fusion weight vector; based on the fusion weight vector, calculate the weighted sum to obtain a cross-image fused feature;
[0073] S33. Based on the cross-image fused feature, calculate the forward matching probability and the reverse matching probability; based on the forward matching probability and the reverse matching probability, perform a two-way consistency check to obtain a consistency-enhanced feature.
[0074] In an embodiment of the present application, the fused feature vectors f' of different images are obtained, denoted as fA and fB respectively, and the mutual attention matrix Cl = tanh((WlfA)T(WlfB)) is calculated in each layer l, where Wl is the learnable weight matrix of the l-th layer; after adding the relative position constraint Rl, perform normalization Al = softmax(Cl + Rl), where Rl is the relative position constraint matrix, to obtain the multi-layer attention weight set A.
[0075] Obtain the multi-layer attention weight set A, and perform the dynamic weight fusion operation w = sigmoid(FC([A1;...; AL])), where FC is the fully connected layer and L is the number of attention layers; calculate the weighted sum fAB = Σ(wi·Ai), where wi is the fusion weight of the i-th layer, to obtain the cross-image fusion feature fAB.
[0076] Obtain the cross-image fusion feature fAB, calculate the forward matching probability CA→B = softmax(fAB) and the reverse matching probability CB→A = softmax(fBAT); perform the bidirectional consistency check C = min(CA→B, CB→AT) to obtain the consistency enhanced feature C.
[0077] In this embodiment, by constructing a multi-layer mutual attention network and a bidirectional consistency check mechanism, high-precision matching of cross-image wire features is achieved. First, calculate the mutual attention matrix at each layer, and introduce relative position constraints for modulation, and capture the feature correlations at different scales and abstraction levels through the multi-layer structure; then, realize the adaptive integration of attentions at different layers through the dynamic weight fusion operation; finally, by calculating the forward matching probability and the reverse matching probability, and performing the bidirectional consistency check, the reliability of the matching is effectively improved. This embodiment improves the accuracy of wire registration, especially in dealing with challenging scenarios such as complex backgrounds and partial occlusions, showing strong robustness, and at the same time reducing the probability of false matching.
[0078] According to one aspect of the present application, step S31 is further as follows:
[0079] S311. Based on the fusion feature vector, construct a multi-scale decomposition function; based on the multi-scale decomposition function, perform feature decomposition to obtain the feature representation at each scale level; perform normalization processing on the feature representation at each scale level to obtain the scale feature set;
[0080] S312. Based on the scale feature set, calculate the spatial displacement vector between feature point pairs, and construct a spatial correlation function; based on the spatial correlation function, construct a spatial constraint matrix to generate a spatial association matrix;
[0081] S313. Based on the scale feature set and the spatial association matrix, construct an adaptive transformation function, and calculate the feature response matrix; based on the feature response matrix, perform feature enhancement operations to obtain the transformed feature matrix.
[0082] S314. Based on the transformed feature matrix, construct a cross-scale attention calculation function, and calculate the inter-layer attention score; based on the inter-layer attention score and the transformed feature matrix, construct a multi-layer attention response to obtain the inter-layer attention tensor;
[0083] S315. Based on the inter-layer attention tensor, construct an attention aggregation function to calculate the final attention score; based on the final attention score, perform attention normalization on each layer to obtain a set of multi-layer attention weights.
[0084] In an embodiment of the present application, obtain the fused feature vectors fA and fB, and construct a multi-scale decomposition function ψl(f)= f * Gl + ηl·▽ 2 f, where Gl is a Gaussian kernel, and ▽ 2 is the Laplace operator, and ηl is a scale coefficient; perform feature decomposition at L scale levels: sl = ψl(f) + ρl·s(l-1), where ρl is an inter-layer coupling factor, and sl and s(l-1) are feature representations at different scale levels; perform normalization processing on the features of each layer to obtain the scale feature sets SA and SB.
[0085] Obtain the scale feature sets SA and SB, and calculate the spatial displacement vector d(i, j) = pi - pj between feature points, where pi and pj are the spatial coordinates of the feature points; construct a spatial correlation function φ(d) = exp(-||d||2 / 2σ 2 )·(1 + κ·cos(θij)), where σ is a spatial scale parameter, κ is a direction modulation coefficient, and θij is the direction angle; generate a spatial constraint matrix R(i, j) = φ(d(i, j)) to obtain the spatial correlation matrix R.
[0086] Obtain the scale feature sets SA and SB and the spatial correlation matrix R, and construct an adaptive transformation function τ(f, R) = f + λ·tanh(μ·(R * f)), where λ is a transformation intensity coefficient and μ is a non-linear modulation factor; calculate the feature response matrix E(i, j) = τ(fi, Rij)·τ(fj, Rji)T; perform a feature enhancement operation T = norm(E + β·E 2 ), where β is a second-order enhancement coefficient, to obtain the transformed feature matrices TA and TB.
[0087] Obtain the transformed feature matrices TA and TB, and construct a cross-scale attention calculation function ω(Tl, Tk) = softmax(Tl·WlkT·Tk), where Wlk is an inter-layer mapping matrix; calculate the inter-layer attention score αlk = ω(Tl, Tk)·(1 + γ·|l-k|)-1, where γ is an inter-layer attenuation factor; integrate the multi-layer attention responses L(l, k) = αlk·(Tl + Tk) / 2 to obtain the inter-layer attention tensor L.
[0088] Obtain the inter-layer attention tensor L, and construct the attention aggregation function a(L) = Σ(πl·Ll), where πl is the layer weight coefficient; calculate the final attention score A = sigmoid(ξ·a(L)), where ξ is the scaling factor; perform attention normalization on each layer to obtain the multi-layer attention weight set A.
[0089] In this embodiment, through multi-scale feature representation and inter-layer attention mechanism, accurate matching of cross-image wire features is achieved. First, a hierarchical feature representation is constructed through a multi-scale decomposition function, and an inter-layer coupling factor is introduced to achieve effective transmission of features; a spatial constraint relationship between feature points is established through a spatial correlation function to ensure the spatial consistency of the matching; an adaptive transformation function is used to enhance the feature response, and the discriminability of the features is improved through second-order enhancement; finally, through cross-scale attention calculation and inter-layer attention fusion, effective integration of multi-layer features is achieved. This embodiment improves the accuracy and robustness of wire registration, especially when dealing with large-scale changes and complex deformations, showing excellent performance.
[0090] As Figure 5 shown, according to one aspect of the present application, step S4 is further as follows:
[0091] S41. Based on the consistency-enhanced features, construct multi-dimensional similarity features to obtain a similarity feature tensor; based on the similarity feature tensor, extract local structure constraints to obtain a structure constraint matrix; based on the structure constraint matrix, calculate topological relationship features to obtain a topological feature set; based on the similarity feature tensor, structure constraint matrix, and topological feature set, construct a multi-constraint fusion function to obtain a comprehensive constraint tensor; based on the comprehensive constraint tensor, generate an initial matching matrix;
[0092] S42. Construct topological consistency constraints, direction consistency constraints, and distance consistency constraints to optimize the initial matching matrix to obtain an optimized matching matrix.
[0093] In an embodiment of the present application, obtain the consistency-enhanced feature C, calculate the feature similarity sim(fi, fj) and geometric similarity sim(gi, gj); perform a weighted fusion operation Sij = βf·sim(fi, fj) + βg·sim(gi, gj), where βf and βg are adaptive weight coefficients, to obtain the initial matching matrix S.
[0094] Obtain the initial matching matrix S, and optimize it based on the topological consistency constraint T(Xi)≈T(Xj), direction consistency constraint |θi - θj| < ε, and distance consistency constraint D(pi, pj) < τ, where T(·) is the topological feature extraction function, θi, θj are direction angles, D(·) is the distance metric function, and ε and τ are threshold parameters, to obtain the optimized matching matrix X.
[0095] This embodiment adopts a multi-constraint matching matrix generation and optimization strategy, achieving high-precision output of wire matching results. By fusing feature similarity and geometric similarity to construct an initial matching matrix and introducing an adaptive weight coefficient for dynamic balance; in the optimization stage, by imposing topological consistency constraints, direction consistency constraints, and distance consistency constraints, multi-dimensional constraints and optimization of the matching results are achieved. This embodiment not only ensures the accuracy of local matching but also maintains the consistency of the global geometric structure, effectively solving complex problems such as deformation and scale change in wire registration, and improving the reliability and precision of the registration results.
[0096] According to one aspect of the present application, step S41 is further as follows:
[0097] S411. Obtain the consistency-enhanced feature C, construct the feature distance metric function df(fi, fj) = ||fi - fj||2·(1 - cos<fi, fj>), where cos<·, ·> is the cosine similarity; calculate the local shape descriptor ds(pi, pj) = Σwk·||▽kpi - ▽kpj||2, where ▽k is the k-order difference operator and wk is the weight coefficient; fuse the multi-dimensional similarity D(i, j) = [df(fi, fj), ds(pi, pj), C(i, j)] to obtain the similarity feature tensor D.
[0098] S412. Obtain the similarity feature tensor D, construct the local structure tensor S(p) =Σwij·(pi - p)(pj -p) T , where wij is the spatial weight; calculate the structural similarity metric gs(Si, Sj) = ||Si - Sj|| F / max(||Si|| F , ||Sj|| F ), where ||·|| F is the Frobenius norm; perform non-local structure comparison g(i, j) = exp(-gs(Si, Sj) 2 / 2σ 2 ), to obtain the structure constraint matrix G.
[0099] S413. Obtain the structure constraint matrix G, construct the local neighborhood graph Ne(i) = {j | ||pi - pj|| <ε}; calculate the topological feature vector t(i) = [deg(i), ecc(i), cen(i)], where deg is the degree centrality, ecc is the eccentricity, and cen is the betweenness centrality; perform the topological similarity metric τ(i, j) = exp(-||t(i) - t(j)||2 / 2ν 2 ), to obtain the topological feature set T.
[0100] S414. Obtain the similarity feature tensor D, the structure constraint matrix G, and the topological feature set T, and construct the multi-constraint fusion function h(i, j) = [D(i, j), G(i, j), T(i, j)]; calculate the constraint weight w = softmax(φ(h)), where φ is the feature mapping function; perform weighted fusion H(i, j) = Σwk·hk(i, j) to obtain the comprehensive constraint tensor H.
[0101] S415. Obtain the comprehensive constraint tensor H, and construct the matching score function s(i, j) = ω·H(i, j) + (1 - ω)·∏ k∈N(i) max(H(k, l)), where ω is the global-local balance factor, N(i) is the neighborhood set of i, and ∏ is the product operation; perform matching score normalization: S(i, j) = s(i, j) / max(s(i, :), s(:, j)) to obtain the initial matching matrix S.
[0102] In this embodiment, through multi-dimensional similarity feature and constraint fusion, high-reliability wire matching is achieved. By constructing the feature distance metric function and the local shape descriptor, multi-dimensional similarity measurement is realized; through local structure tensor analysis and non-local structure comparison, a complete structure constraint system is established; the topological feature vector is introduced to capture the global structure information, and the constraints are unified and integrated through the multi-constraint fusion function; finally, the global and local information is balanced through the matching score function to generate a high-quality initial matching result. This embodiment effectively improves the accuracy and reliability of the matching, and especially shows strong robustness when dealing with complex scenarios and abnormal situations.
[0103] As Figure 6 shown, according to one aspect of the present application, step S5 is further as follows:
[0104] S51. Based on the optimized matching matrix, construct multi-dimensional evaluation features to obtain an evaluation feature set; based on the evaluation feature set, calculate the local consistency metric to obtain a consistency feature matrix; based on the consistency feature matrix, extract the global stability feature to obtain a stability feature tensor; fuse the evaluation feature set, the consistency feature matrix, and the stability feature tensor to generate a comprehensive evaluation tensor; based on the comprehensive evaluation tensor, construct a reliability scoring function to generate a reliability scoring matrix;
[0105] S52. Based on the reliability scoring matrix and a preset adaptive threshold, perform a conditional judgment on each position to obtain the final matching result.
[0106] In an embodiment of the present application, an optimized matching matrix X is obtained, and a comprehensive reliability score R(i, j) = α1·Sij + α2·Cij + α3·Gij is calculated, where Sij is the feature similarity score, Cij is the consistency score, Gij is the geometric consistency score, and α1, α2, and α3 are adaptive weight coefficients, to obtain a reliability scoring matrix R.
[0107] The reliability scoring matrix R is obtained, an adaptive threshold η is set, and a conditional judgment is performed for each position (i, j): when R(i, j) > η and R(i, j) is simultaneously the maximum value in its row and column, the corresponding position is set as a matching point; otherwise, it is set as a non-matching point, to obtain the final matching result M.
[0108] In this embodiment, by constructing multi-dimensional evaluation features and calculating local consistency metrics, each element of the matching matrix can be evaluated in detail, capturing more detailed information, thereby improving the accuracy of the matching; the local consistency metric ensures the consistency of the matching results within a local range, reducing the possibility of false matches. Extracting global stability features based on the consistency feature matrix can effectively evaluate the stability of the matching results within a global range; by comprehensively considering local and global features, it is ensured that the matching results are not only consistent within a local range but also stable within a global range, improving the reliability of the matching results. Fusing the evaluation feature set, the consistency feature matrix, and the stability feature tensor to generate a comprehensive evaluation tensor can comprehensively evaluate all aspects of the matching results, generate a reliability scoring matrix for each matching result, and provide an important reference for subsequent decision-making. Based on the reliability scoring matrix and a preset adaptive threshold, a conditional judgment is performed for each position, which can dynamically adjust the screening criteria for the matching results. This adaptive threshold judgment mechanism can be flexibly adjusted according to the actual situation, improving the accuracy and adaptability of the matching results. This embodiment can generate a final matching result with high precision and high reliability, providing strong guarantees both in terms of local consistency and global stability, ensuring the accuracy and reliability of the matching results.
[0109] According to one aspect of the present application, step S51 is further as follows:
[0110] S511. Obtain the optimized matching matrix X, and calculate the feature similarity evaluation score sf(i, j) = (1 + ||fi - fj||2 / δf) -1 , where δf is the feature distance threshold; construct the spatial transformation consistency evaluation sc(i, j) = exp(-||Ti(pi) - pj||2 / 2σs 2), where \(T_i\) is a local affine transformation and \(\sigma_s\) is a spatial scale parameter; calculate the directional continuity evaluation \(sd(i, j)=\cos(\theta_i - \theta_j)\), where \(\theta_i\) and \(\theta_j\) are local directional angles; combine the multi-dimensional evaluation features \(E(i, j)=[sf(i, j), sc(i, j), sd(i, j)]\) to obtain the evaluation feature set \(E\).
[0111] S512. Obtain the evaluation feature set \(E\) and construct the local neighborhood consistency function \(\psi(i, j)=\sum_{}\) k∈N(i) \(\sum_{}\) l∈N(j) \(X(k, l)\cdot\exp(-||d_{ik}-d_{jl}||^2 / 2\sigma_l)\) 2 ), where \(d_{ik}\) and \(d_{jl}\) are relative displacement vectors and \(\sigma_l\) is a local scale parameter; calculate the structure preservation metric \(\mu(i, j)=|\log(r_i / r_j)|\), where \(r_i\) and \(r_j\) are local radii; fuse the local evaluation indicators \(L(i, j)=[\psi(i, j), \mu(i, j), E(i, j)]\) to obtain the consistency feature matrix \(L\).
[0112] S513. Obtain the consistency feature matrix \(L\) and construct the global shape descriptor \(h(P)=[\kappa(P), \tau(P), \rho(P)]\), where \(\kappa\) is the curvature feature, \(\tau\) is the torsion feature, and \(\rho\) is the density feature; calculate the shape similarity metric \(ds(h_i, h_j)=||h_i - h_j||^2\cdot(1 - \cos\langle h_i, h_j\rangle)\); perform the stability evaluation \(q(i, j)=\exp(-ds(h_i, h_j)\) 2 / 2\sigma_q\) 2 )\cdot(1 + \eta\cdot|Ne(i)-Ne(j)|)^{-1}\), where \(Ne\) is the number of effective matches and \(\sigma_q\) is the stability scale parameter, to obtain the stability feature tensor \(Q\).
[0113] S514. Obtain the evaluation feature set \(E\), the consistency feature matrix \(L\), and the stability feature tensor \(Q\), and construct the feature fusion network \(v(i, j)=[E(i, j), L(i, j), Q(i, j)]\); calculate the adaptive weight coefficient \(w = \text{softmax}(\varphi(v))\), where \(\varphi\) is a non-linear mapping function; perform weighted fusion \(V(i, j)=\sum_{}w_k\cdot v_k(i, j)\) to obtain the comprehensive evaluation tensor \(V\).
[0114] S515. Obtain the comprehensive evaluation tensor \(V\) and construct the reliability scoring function \(r(i, j)=\lambda_g\cdot V(i, j)+\lambda_l\cdot\sum_{}\) k∈N(i)max(V(k, l)), where λg and λl are global-local balance factors; apply an adaptive threshold to adjust r'(i, j) = sigmoid(ζ·(r(i, j) – r*)), where r* is the average reliability score and ζ is a scaling factor; perform regional consistency correction R(i, j) = r'(i, j)·∏ k∈Ω(i,j) (1 + ε·r'(k, l)), where Ω is the local region and ε is the association strength parameter, to obtain the reliability score matrix R.
[0115] In this embodiment, through multi-dimensional evaluation features and comprehensive reliability metrics, high-precision screening of matching results is achieved. First, a complete evaluation system is constructed through feature similarity evaluation, spatial transformation consistency evaluation, and direction continuity evaluation; the local consistency of the match is evaluated through the local neighborhood consistency function and structure-preserving metrics; the global stability is evaluated using global shape descriptors and shape similarity metrics; and finally, the final matching screening is achieved through the reliability scoring function and regional consistency correction. This embodiment effectively improves the reliability of the registration results, especially when dealing with noise interference and abnormal matches, showing excellent discrimination ability.
[0116] In an embodiment of the present application, a method for wire registration between UAV aerial images based on deep learning includes the following steps:
[0117] S1. Construct a mathematical model;
[0118] Obtain two UAV aerial images I_A and I_B with an overlapping area. M wires L_A := {L_A 1 , L_A 2 , …, L_A M} have been detected in image I_A, and their indices are A := {1, 2, …, M}. N wires L_B := {L_B 1 , L_B 2 , …, L_B N} have been detected in image I_B, and their indices are B := {1, 2, …, N}. Each wire detected in the two images is given in the form of a set of points. After matching, the final output is the matching matrix P ∈ [0, 1] M×N .
[0119] S2. Extract the features of different wires in the same image based on the self-attention mechanism;
[0120] Use the feature extraction part of the ResNet18 model to extract features from the image, and then sample on the feature map according to the positions of the points on the wire. Each point on the wire can obtain 512-dimensional features. Assume that there are n points on each wire, then each wire corresponds to a feature f_i ∈ R n×512. To integrate the features of all points on each wire for subsequent feature matching, a self-attention module is used to integrate the global information of the wire.
[0121] Among them, the self-attention module associates different positions of a single sequence, allowing interaction between the input and the input, and keeping the dimensions of the output and the input the same. In the self-attention module, first, a linear transformation is performed on the input f_i to generate three matrices K, Q, and V: K = W K f_i; Q = W Q f_i; V = W V f_i; where W K , W Q , W V represents the linear transformation layer, and the final output is: f_o = Attention(Q, K, V) = Softmax((QK T ) / sqrt(d_k))V; where d_k is the number of columns of the K matrix. Since the self-attention mechanism can utilize global information, the first dimension of f_o is used as the extracted wire feature, that is, f_o ∈ R 1×512 .
[0122] Through the wire feature extraction module, each wire L_i ∈ {L_A, L_B} in the image I ∈ {I_A, I_B} can obtain a 512-dimensional feature f_o I , and it is used as the initial state s_i of each wire I ← f_o I . Subsequently, these wire features are fed into the feature matching module, which consists of L (L = 7) identical layers that jointly process two feature sets. Each layer consists of a self-attention module and a cross-attention module, which are used to update the representation of each wire. In each module, a multi-layer perceptron (MLP) is used to update the state s_i I .
[0123] For the self-attention module, the state s_i I is updated as follows: s_i I ← s_i I + MLP([s_i I | y_i I←I ), where [∙|∙] represents concatenating two vectors dimensionally; y_i I←I = ∑ J∈{I_A,I_B} Softmax(α_ik IJ ) j Ws_j {I_A,I_B} ; where Softmax(x_i) = e x_i / (∑ je x_j ), where \(W\) is the mapping matrix and \(k\in\{I_A, I_B\}\). \(\alpha_{ij}\) IJ is the attention score between lines \(L_i\) and \(L_j\). In the self-attention module, each wire focuses on all the wires in the same image. For each line \(L_i\in\{L_A, L_B\}\), the current state \(x_i\) I is transformed into a key vector \(k_i\) and a query vector \(q_i\) through a linear transformation. The attention score \(\alpha_{ij}\) between lines \(L_i\) and \(L_j\) is defined as: \(\alpha_{ij} = q_i\) T k_j\).
[0124] S3. Extract the features of the same wire in different images based on the cross-attention mechanism;
[0125] For the cross-attention module, each wire in the image focuses on all the wires in the other image, and the state \(x_i\) I is updated as follows: \(s_i\) I ← \(s_i\) I + MLP([\(s_i\) I | \(y_i\) I←{I_A,I_B}\I ); \(y_i\) I←{I_A,I_B}\I = ∑ J∈{I_A,I_B} Softmax(\(\alpha_{ik}\) IJ ) j Ws_j {I_A,I_B} ; \(\alpha_{ij} = k_i\) T k_j. After the self-attention module and the cross-attention module form a multi-layer network, a lightweight output head is used to predict the feature matching situation based on the updated state.
[0126] S4. Obtain the correspondence between wires through feature matching;
[0127] Calculate the similarity matrix \(S\in\mathbb{R}\) M×N between the wires in the two images: \(S_{ij} = Linear(s_i\) A ) T Linear(s_i\) B ); for all \((i, j)\in A\times B\); where the similarity can be considered as the probability that this pair of wires is projected from the same 3D wire in space. In the formula, \(Linear(\cdot)\) represents the linear transformation layer, and the data dimension remains unchanged. \(s_i\) A and \(s_i\) B represent the results after the initial states of the wires in images \(I_A\) and \(I_B\) are updated through the multi-layer network. In addition to calculating the similarity matrix, the matchability score \(\sigma_i\in[0, 1]\) of each wire is also considered: \(\sigma_i = Sigmoid(Linear(s_i))\), where \(Sigmoid(x)=\frac{1}{1 + e}\) -x) is a common activation function that can map a variable between 0 and 1. The matchability score represents the likelihood that the current wire has the corresponding wire. The closer the value is to 1, the higher the likelihood. Combining the similarity matrix and the matchability score, we get the wire matching matrix P: P_ij = σ_i A σ_j B Softmax(S_(k_1 j) ) i Softmax(S_(ik_2 )) j , k_1 ∈ A, k_2 ∈ B.
[0128] The loss function in the model training process consists of three parts, corresponding to the correctly matched and unmatched wires respectively. Among them, the first part in the loss function is for the correctly matched wires, calculating the log-likelihood loss of the predicted correspondence P_ij, where P_ij is an element in the predicted assignment matrix; the second part is for the unmatched wires in image A, calculating the log-likelihood loss of the predicted non-matchability score 1 - σ_i A where σ_i A is the matchability score of wire i in image A; the third part is for the unmatched wires in image B, calculating the log-likelihood loss of the predicted non-matchability score 1 - σ_j B where σ_j B is the matchability score of wire j in image B. As follows:
[0129] loss = -1 / L ∑ l (1 / |M| ∑ (i,j)∈M logP_ij + 1 / 2|A*| ∑ i∈A* log(1 - σ_i A ) + 1 / 2|B*| ∑ j∈B* log(1 - σ_j B ));
[0130] where M represents the correctly matched wires, A* represents the unmatched wires in image A, and B* represents the unmatched wires in image B.
[0131] S5. Given the matching results based on the transmission line samples, effectively realizing the correspondence and information fusion between wires;
[0132] The matching matrix P takes into account the similarity of the corresponding wires and the matching scores. In the prediction stage, for a pair of wires, when both wires are matchable and the similarity between them is higher than that of other wires, then this pair of wires is related. When the value of P_ij is greater than the threshold (set to 0.1) and greater than other values in its column and row, then this pair of wires is confirmed as a match.
[0133] In this embodiment, the self-attention mechanism helps the model understand and focus on the internal features of the wire, while the cross-attention mechanism can focus on the features of the same wire in different images, improving the accuracy of matching. By accurately matching the wires, data from different images can be integrated to achieve comprehensive fusion and in-depth analysis of information, providing data support for the safety assessment and maintenance strategies of power lines. Based on the accurate wire matching results, power companies can make more scientific and reasonable maintenance plans and resource allocation decisions, which helps to build a more intelligent, efficient, and reliable power system.
[0134] According to one aspect of the present application, a wire registration system between UAV aerial images based on deep learning includes:
[0135] At least one processor; and,
[0136] A memory communicatively connected to the at least one processor; wherein,
[0137] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method for registering wires between UAV aerial images based on deep learning described in any one of the above embodiments.
[0138] The present invention realizes high-precision automatic registration of wires between UAV aerial images through a multi-stage adaptive processing strategy. Through the organic combination of an adaptive image preprocessing mechanism, non-linear multi-scale feature extraction, position-aware self-attention feature fusion, a multi-layer cross-attention matching network, and a multi-constraint optimization strategy, the system can effectively handle various complex situations in the aerial photography scene, such as light changes, perspective differences, local deformations, etc. In practical applications, the present invention improves the accuracy of wire detection and registration, reduces the computational complexity, and improves the processing efficiency. Especially when dealing with large-scale aerial photography data, it shows excellent robustness and scalability. The end-to-end design of the system not only simplifies the actual deployment process but also provides good real-time performance, providing reliable technical support for intelligent inspection of power systems.
[0139] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A method for wire registration between UAV aerial images based on deep learning, characterized in that It includes the following steps: S1. Obtain at least two original UAV aerial images containing overlapping regions, and perform adaptive dynamic range compression processing on them to obtain preprocessed images; based on the preprocessed images, construct a non-linear multi-scale feature pyramid, extract wire features of each scale layer, and obtain an enhanced wire feature set; Based on the enhanced wire feature set, construct a fusion feature matrix through dynamic feature fusion; S2. Based on the fusion feature matrix, perform spatial-aware self-attention processing to obtain a position-aware feature matrix; Based on the position-aware feature matrix, calculate enhanced self-attention scores to obtain fusion feature vectors between different wires; S3. Based on the fusion feature vectors, construct a multi-layer mutual attention network to obtain a multi-layer attention weight set; based on the multi-layer attention weight set, perform feature aggregation to obtain cross-image fusion features; Based on the cross-image fusion features, perform bidirectional consistency checking to obtain consistency-enhanced features; S4. Based on the consistency-enhanced features, construct a multi-constraint matching matrix to obtain an initial matching matrix; Optimize the initial matching matrix to obtain an optimized matching matrix; S5. Based on the optimized matching matrix, calculate matching reliability scores to obtain a reliability scoring matrix; based on the reliability scoring matrix and a preset threshold, screen to obtain the final matching result; Step S2 is further as follows: S21. Obtain the wire feature set of the same image in the fusion feature matrix and calculate position encoding; Based on the position encoding, perform a feature transformation operation to obtain a position-aware feature matrix; S22. Based on the position-aware feature matrix, construct a spatial constraint matrix, calculate attention scores, and obtain fusion feature vectors.
2. The method for wire registration between UAV aerial images based on deep learning according to claim 1, wherein Step S1 is further as follows: S11. Obtain at least two original UAV aerial images containing overlapping regions; based on the image quality of the original UAV aerial images, evaluate to obtain an image quality feature vector; Based on the image quality feature vector, perform adaptive enhancement to obtain enhanced images; Based on the enhanced images, perform illumination equalization and noise suppression to obtain balanced images; Based on the balanced images, perform edge feature enhancement to obtain edge-enhanced images; Based on the edge-enhanced images, perform adaptive dynamic range compression to obtain preprocessed images; S12. Based on the preprocessed images, construct an improved Laplacian pyramid; based on the Laplacian pyramid, perform a compression operation on each layer to obtain compressed feature maps; Based on the compressed feature maps, perform pyramid decomposition operations to obtain Laplacian feature maps; based on the Laplacian feature maps, generate a multi-scale feature pyramid; S13. Based on the multi-scale feature pyramid, extract local structure features to obtain a structure feature matrix; based on the structure feature matrix, calculate direction consistency features to obtain a direction feature set; based on the direction feature set, extract and fuse context information to obtain context-enhanced features; Based on the context-enhanced features, integrate to obtain a multi-dimensional feature tensor, and finally generate an enhanced wire feature set; S14. Based on the enhanced wire feature set, calculate dynamic fusion weights; Based on the dynamic fusion weights, perform a weighted summation operation to obtain a fusion feature matrix.
3. The method for wire registration between UAV aerial images based on deep learning according to claim 2, wherein Step S3 is further as follows: S31. Based on the fusion feature vectors of different images, construct multi-scale feature representations to obtain a scale feature set; Based on the scale feature set, calculate the spatial dependence relationship to obtain a spatial correlation matrix; Based on the spatial correlation matrix, perform adaptive feature transformation to obtain a transformed feature matrix; Based on the transformed feature matrix, calculate cross-scale attention to obtain an inter-layer attention tensor; Based on the inter-layer attention tensor, generate a multi-layer attention weight set; S32. Based on the multi-layer attention weight set, perform dynamic weight fusion operation to obtain a fusion weight vector; based on the fusion weight vector, calculate the weighted sum to obtain a cross-image fusion feature; S33. Based on the cross-image fusion feature, calculate the forward matching probability and the reverse matching probability; Based on the forward matching probability and the reverse matching probability, perform two-way consistency check to obtain a consistency-enhanced feature.
4. The method for wire registration between UAV aerial images based on deep learning according to claim 3, wherein Step S4 is further as follows: S41. Based on the consistency-enhanced feature, construct multi-dimensional similarity features to obtain a similarity feature tensor; based on the similarity feature tensor, extract local structure constraints to obtain a structure constraint matrix; Based on the structure constraint matrix, calculate topological relationship features to obtain a topological feature set; based on the similarity feature tensor, the structure constraint matrix and the topological feature set, construct a multi-constraint fusion function to obtain a comprehensive constraint tensor; Based on the comprehensive constraint tensor, generate an initial matching matrix; S42. Construct topological consistency constraints, direction consistency constraints and distance consistency constraints to optimize the initial matching matrix to obtain an optimized matching matrix.
5. The method for wire registration between UAV aerial images based on deep learning according to claim 4, characterized in that, Step S5 is further as follows: S51. Based on the optimized matching matrix, construct multi-dimensional evaluation features to obtain an evaluation feature set; based on the evaluation feature set, calculate local consistency metrics to obtain a consistency feature matrix; Based on the consistency feature matrix, extract global stability features to obtain a stability feature tensor; Fuse the evaluation feature set, the consistency feature matrix and the stability feature tensor to generate a comprehensive evaluation tensor; based on the comprehensive evaluation tensor, construct a reliability scoring function to generate a reliability scoring matrix; S52. Based on the reliability scoring matrix and a preset adaptive threshold, perform a conditional judgment on each position to obtain a final matching result.
6. The method for wire registration between UAV aerial images based on deep learning according to claim 5, wherein Step S11 is further as follows: S111. Obtain at least two original UAV aerial images containing overlapping regions, calculate image sharpness features, illumination uniformity features and signal-to-noise ratio features; combine the image sharpness features, illumination uniformity features and signal-to-noise ratio features to obtain an image quality feature vector; S112. Based on the original UAV aerial image and its corresponding image quality feature vector, construct an adaptive enhancement function; Based on the adaptive enhancement function, perform an enhancement operation on each pixel to obtain an enhanced image; S113. Based on the enhanced image, construct an adaptive weight matrix and perform a weighted smoothing operation to obtain a smoothed image; based on the smoothed image, perform local contrast correction to obtain a balanced image; S114. Based on the balanced image, calculate multi-directional gradient responses and construct a direction enhancement kernel; Based on the direction enhancement kernel, perform a direction-selective convolution operation to obtain an edge-enhanced image; S115. Calculate local statistical features based on the edge-enhanced image, and construct an adaptive compression coefficient; Calculate a dynamic adjustment factor based on the adaptive compression coefficient; Perform pixel-level adaptive compression operation based on the edge-enhanced image and the dynamic adjustment factor to obtain a preprocessed image.
7. The method for wire registration between UAV aerial images based on deep learning according to claim 5, wherein Step S13 is further as follows: S131. Based on the multi-scale feature pyramid, construct a structure tensor at each scale layer, calculate eigenvalues and eigenvectors; based on the eigenvalues and eigenvectors, construct a local structure descriptor; Perform feature aggregation on each detection point based on the local structure descriptor to obtain a structure feature matrix; S132. Calculate a direction vector field based on the structure feature matrix, and construct a direction consistency measure; perform non-maximum suppression based on the direction consistency measure to obtain a local principal direction; enhance the direction coherence by using a tensor voting mechanism based on the local principal direction to obtain a direction feature set; S133. Construct an adaptive sampling window based on the direction feature set, and extract local context features; Apply non-local mean filtering based on the local context features to obtain context-enhanced features; S134. Construct a multi-dimensional feature descriptor based on the context-enhanced features, and calculate feature importance weights; Perform weighted feature fusion based on the feature importance weights to obtain a multi-dimensional feature tensor; S135. Based on the multi-dimensional feature tensor, extract a predetermined number of key points for each wire, and calculate the feature representation of each key point; Perform feature normalization and dimensionality reduction based on the feature representation of each key point to obtain an enhanced wire feature set.
8. The method for wire registration between UAV aerial images based on deep learning according to claim 5, characterized in that, Step S31 is further as follows: S311. Construct a multi-scale decomposition function based on the fusion feature vector; perform feature decomposition based on the multi-scale decomposition function to obtain the feature representation at each scale level; perform normalization processing on the feature representation at each scale level to obtain a scale feature set; S312. Calculate the spatial displacement vector between feature point pairs based on the scale feature set, and construct a spatial correlation function; Construct a spatial constraint matrix based on the spatial correlation function, and generate a spatial association matrix; S313. Construct an adaptive transformation function based on the scale feature set and the spatial association matrix, and calculate a feature response matrix; Perform feature enhancement operation based on the feature response matrix to obtain a transformed feature matrix; S314. Construct a cross-scale attention calculation function based on the transformed feature matrix, and calculate the inter-layer attention score; Construct a multi-layer attention response based on the inter-layer attention score and the transformed feature matrix to obtain an inter-layer attention tensor; S315. Construct an attention aggregation function based on the inter-layer attention tensor, and calculate the final attention score; Perform attention normalization on each layer based on the final attention score to obtain a multi-layer attention weight set.
9. The wire registration system between UAV aerial images based on deep learning is characterized in that Including: At least one processor; And, A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method for wire registration between UAV aerial images based on deep learning according to any one of claims 1 to 8.
Citation Information
Patent Citations
A defect detection method for transmission equipment based on multi-source image feature matching of unmanned aerial vehicle (UAV)
CN109544501A
Power transmission line inspection image detection method based on deep convolutional neural network
CN117541535A