Wire Instance Detection Method and System Based on Key Point Matching
The keypoint matching method enhances feature extraction and alignment to address challenges in power line detection in aerial images, ensuring accurate and efficient detection across varying conditions.
Patent Information
- Application Number
- CN202411775989.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-12-05
AI Technical Summary
The existing wire detection methods are difficult to effectively distinguish the conductor from the background under complex backgrounds. The consistency of feature descriptions when connected across blocks is poor. The changes in light and viewing angles affect atmospheric disturbances, resulting in obvious differences in apparent features, and it is difficult to accurately distinguish conductor examples in cross-sections and dense areas. The effect of the existing methods is limited in practical applications.
The wire instance detection method based on key point matching is adopted, and a four-dimensional feature matrix is generated through dynamic feature mapping and enhancement processing, projection transformation and multi-scale feature fusion are carried out, enhanced feature maps are constructed, feature reorganization and recursive decomposition, key points are screened, thermal maps and spatial correlation matrix are generated, endpoint matching and topological optimization are performed, and accurate positioning and global reconstruction of wire instances are realized.
Effectively capture multi-scale features of wires, ensure the continuity and integrity of wire instances, improve the accuracy and reliability of wire detection, and be able to deal with interfering factors such as complex backgrounds and lighting, viewing angle changes, and ensure the consistency of computing efficiency and detection results.
Smart Images

Figure CN119672004B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target detection, and in particular, to a wire instance detection method and system based on key point matching. Background Art
[0002] In the operation and maintenance of power systems, the safety status of transmission lines is directly related to the reliability and stability of power supply. With the rapid development of ultra-high voltage transmission networks and the continuous advancement of smart grid construction, traditional manual inspection methods are increasingly unable to meet the inspection requirements of large scale, high frequency, and high precision. The introduction of aerial inspection technology has brought new development opportunities for power line detection. It can not only quickly obtain high-resolution image data over a large range but also effectively avoid the safety risks of manual inspection in complex terrains and harsh environments. Especially in complex terrain areas such as mountains and jungles, aerial inspection can obtain detailed information in areas that are difficult to reach by traditional manual inspection, providing important data support for the preventive maintenance of power lines.
[0003] Currently, the research on wire detection in aerial images mainly focuses on methods such as edge detection, Hough transform, and deep learning. Edge detection-based methods extract image edge features through operators such as Canny and combine morphological processing to achieve preliminary wire recognition. However, they are prone to generating a large number of false edges when dealing with images with complex backgrounds and uneven lighting. Hough transform-based methods utilize the fact that wires appear as straight or approximately straight lines in images for detection, but they have poor detection effects for scenarios with curved wires and multi-wire intersections. Deep learning methods achieve detection by constructing convolutional neural network models to learn high-level semantic features of wires. However, existing models often require a large amount of labeled data for training and have problems with receptive field limitations when dealing with ultra-long wires.
[0004] In practical applications, the existing wire detection methods still have the following technical difficulties: First, wires in aerial images exhibit the characteristics of "ultra-thin and long, low contrast". Especially in complex backgrounds, it is difficult to effectively distinguish wire pixels from the background, and traditional feature extraction methods are difficult to capture such subtle target features. Second, in high-resolution aerial images, wires span multiple image blocks. Existing feature matching algorithms are difficult to ensure the consistency of feature descriptions when dealing with cross-block wire connections, resulting in break points and false connections. Third, due to the influence of camera perspective, lighting changes, and atmospheric disturbances, the apparent features of the same wire in different images are significantly different, and existing feature description methods are difficult to establish stable feature expressions. Fourth, in wire intersection and dense areas, the projections of multiple wires overlap, and coupled with the influence of shadows and occlusions, existing instance segmentation methods are difficult to accurately distinguish different wire instances, easily causing confusion in wire connection relationships. These technical difficulties severely restrict the application effect of wire detection algorithms in practical engineering. Summary of the Invention
[0005] Objective of the invention: To propose a wire instance detection method and system based on key point matching to solve the above problems existing in the prior art.
[0006] Technical solution: The wire instance detection method based on key point matching includes the following steps:
[0007] S1. Obtain the original aerial image data, generate a four-dimensional feature matrix through dynamic feature mapping and enhancement processing; perform a projection transformation on the four-dimensional feature matrix to obtain three-dimensional projection feature data; based on the three-dimensional projection feature data, perform feature enhancement and multi-scale feature fusion to generate fused feature data; based on the fused feature data, generate enhanced feature map data;
[0008] S2. Based on the enhanced feature map data, obtain feature matrix data through feature recombination; based on the feature matrix data, perform recursive decomposition and mapping to generate response matrix data; based on the response matrix data, calculate adaptive threshold data and screen key points; perform feature enhancement on the screened key points, and output a feature description set and a key point set;
[0009] S3. Based on the key point set and the feature description set, generate heat map data through a multi-layer heat map; based on the heat map data and the spatial positions of the key points, construct a spatial association matrix and multi-dimensional constraints, perform end point matching, and obtain matching score data; based on the matching score data, construct and verify wire instances, and finally output a wire instance set;
[0010] S4. Based on the wire instance set, perform boundary feature analysis to obtain a boundary feature matrix; based on the boundary feature matrix, perform cross-block instance matching to generate cross-block matching data; based on the cross-block matching data, perform topological optimization to generate topological credibility data; based on the topological credibility data, perform instance fusion and quality assessment to generate a confidence map and a global wire map.
[0011] The wire instance detection system based on key point matching includes:
[0012] At least one processor; and,
[0013] A memory communicatively connected to at least one of the processors; wherein,
[0014] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned wire instance detection method based on key point matching.
[0015] Beneficial effects: The present invention effectively captures the multi-scale features of wires, realizes the accurate positioning of key points, ensures the continuity and integrity of wire instances, and achieves the precise reconstruction of a large-scale wire system; it can effectively handle interference factors such as complex backgrounds, lighting changes, and perspective changes in aerial images, improving the accuracy and reliability of wire detection; at the same time, it not only ensures the computational efficiency but also guarantees the global consistency of the detection results. Description of the Drawings
[0016] Figure 1 This is a flowchart of the method of the present invention.
[0017] Figure 2 This is a flowchart of step S1 of the present invention.
[0018] Figure 3 This is a flowchart of step S2 of the present invention.
[0019] Figure 4 This is a flowchart of step S3 of the present invention.
[0020] Figure 5 This is a flowchart of step S4 of the present invention. Detailed Embodiments
[0021] As Figure 1 shown, the present application proposes a wire instance detection method based on key point matching, including the following steps:
[0022] S1. Obtain the original aerial image data, generate a four-dimensional feature matrix through dynamic feature mapping and enhancement processing; perform a projection transformation on the four-dimensional feature matrix to obtain three-dimensional projection feature data; based on the three-dimensional projection feature data, perform feature enhancement and multi-scale feature fusion to generate fused feature data; based on the fused feature data, generate enhanced feature map data;
[0023] S2. Based on the enhanced feature map data, obtain feature matrix data through feature recombination; based on the feature matrix data, perform recursive decomposition and mapping to generate response matrix data; based on the response matrix data, calculate adaptive threshold data and screen key points; perform feature enhancement on the screened key points, and output a feature description set and a key point set;
[0024] S3. Based on the key point set and the feature description set, generate heat map data through a multi-layer heat map; based on the heat map data and the spatial positions of the key points, construct a spatial association matrix and multi-dimensional constraints, perform endpoint matching to obtain matching score data; based on the matching score data, construct and verify wire instances, and finally output a wire instance set;
[0025] S4. Based on the wire instance set, perform boundary feature analysis to obtain a boundary feature matrix; based on the boundary feature matrix, perform cross-block instance matching to generate cross-block matching data; based on the cross-block matching data, perform topological optimization to generate topological credibility data; based on the topological credibility data, perform instance fusion and quality assessment to generate a confidence map and a global wire map.
[0026] As Figure 2 shown, according to one aspect of the present application, step S1 is further:
[0027] S11. Obtain the original aerial image data, construct a four-dimensional feature matrix including spatial coordinate dimension, time dimension, and image channel dimension; calculate the first standard deviation and the first mean of each slice in the four-dimensional feature matrix, and generate an adaptive weight coefficient based on the ratio relationship between the first standard deviation and the first mean; perform weighted combination on the four-dimensional feature matrix and the adaptive weight coefficient, and output three-dimensional projection feature data;
[0028] S12. Based on the three-dimensional projection feature data, calculate the feature variance value and the gradient amplitude; based on the feature variance value, determine the enhancement coefficient and the balance coefficient; based on the median of the gradient amplitude, determine the adjustment coefficient; perform a linear combination on the three-dimensional projection feature data and its second-order gradient, and perform weighting using the enhancement coefficient, the balance coefficient, and the adjustment coefficient, and output enhanced feature data; where the balance coefficient is obtained by subtracting the enhancement coefficient from 1;
[0029] S13. Divide the enhanced feature data into different scales to generate a multi-scale feature sequence; calculate the correlation matrix between the features of each scale in the multi-scale feature sequence to obtain feature correlation degree data; based on the feature correlation degree data, construct a weight matrix; based on the weight matrix, perform weighted fusion on the multi-scale feature sequence, and output fused feature data;
[0030] S14. Based on the fused feature data, calculate its maximum value, minimum value, and standard deviation; based on the maximum value and the minimum value, perform normalization processing on the fused feature data to obtain normalized feature data; based on the standard deviation, calculate the dynamic scaling factor; multiply the normalized feature data by the dynamic scaling factor, and finally output the enhanced feature map data.
[0031] In an embodiment of the present application, obtain the original aerial image I(x, y, t), where (x, y) represents the spatial coordinates and t represents the time dimension, construct a fourth-order tensor T(x, y, t, c), where c represents the image channel dimension, and define a projection operator P: T → R 3, project the fourth-order tensor into a three-dimensional feature space; the projection rule is: P(T) = ∑(w_i * T_i), where w_i is an adaptive weight coefficient, calculated by the following formula: w_i = softmax(σ(T_i) / μ(T_i)), where σ(T_i) represents the standard deviation of the tensor slice, and μ(T_i) represents the mean value.
[0032] Construct an enhancement function E(f) = α * f + β * G(f), where f is the projected feature and G(f) is the gradient enhancement term: G(f) = ▽ 2 f * (1 + γ|▽f|), where α, β, and γ are adaptive parameters, dynamically adjusted by feature statistics: α = log(1 + var(f)); β = 1 - α; γ = median(|▽f|). Construct a pyramid feature sequence {F_k}, k ∈ [1, K], using a cross-scale attention mechanism: A(F_k) = ∑(θ_ij * F_i * F_j), i, j ∈ [1, K]; θ_ij = normalize(F_i T * W * F_j), where W is a learnable weight matrix. The final output is the enhanced feature map F_out: F_out = M(E(P(T))) * A({F_k}), where M is an adaptive normalization operator: M(x) = (x - min(x)) / (max(x) - min(x)) * λ, and λ is a dynamic scaling factor: λ = log(1 + std(x)).
[0033] In another embodiment of the present application, the adaptive weight calculation of the four-dimensional feature matrix is specifically: W(k) = η(R(k))·exp(-λ·σ(k) / μ(k)); where R(k) = Σ(|F(i, j, k) - F(i, j, k - 1)|) / N, and R(k) is the temporal change rate of the k-th dimensional feature; σ(k) = sqrt(Σ(F(i, j, k) - μ(k)) 2 / M), where σ(k) is the standard deviation of the k-th dimensional feature; μ(k) = Σ(F(i, j, k) / M), and μ(k) is the mean value of the k-th dimensional feature; η(x) = 1 / (1 + exp(-ax)) is a sigmoid normalization function; F(i, j, k) represents the value of the k-th dimension at the position (i, j) of the four-dimensional feature matrix; N is the number of temporal samples; M is the number of spatial sampling points; λ is a balance coefficient used to adjust the influence degree of variance on the weight; a is the scaling factor of the sigmoid function.
[0034] The process of generating the dynamically enhanced coefficient is specifically as follows: E(x, y) = α(x, y)·F(x, y) + β(x, y)·L(x, y); where α(x, y) = (1 - exp(-v(x, y) / v_max))·g(x, y), and α(x, y) is the local enhancement coefficient; β(x, y) = 1 - α(x, y), and β(x, y) is the balance coefficient; v(x, y) = Σ(|▽F(x, y)| 2 ), v(x, y) is the local gradient intensity; L(x, y) = F(x, y) - γ·▽ 2 F(x, y), L(x, y) is the Laplacian enhancement term; g(x, y) = exp(-|I(x, y) - μ| 2 / 2σ 2 ), g(x, y) is the Gaussian weight; v_max is the maximum value of the gradient intensity; γ is the weight coefficient of the Laplacian operator; ▽ represents the gradient operator; ▽ 2 represents the Laplacian operator; I(x, y) is the original image intensity; μ is the local mean; σ is the Gaussian kernel parameter.
[0035] In this embodiment, a feature extraction and conversion framework from four dimensions to three dimensions is constructed through dynamic feature mapping and enhancement processing, realizing multi-dimensional capture and enhancement of wire features in aerial images. By constructing a four-dimensional feature matrix, spatial position, time series, and image channel information are encoded simultaneously, providing a rich information basis for subsequent feature extraction. During the feature conversion process, an adaptive weight coefficient is generated by calculating the ratio relationship between the standard deviation and the mean, which can automatically adjust the importance of features in different dimensions and effectively suppress noise and interference in aerial images. In the feature enhancement stage, joint constraints of the feature variance value and the gradient amplitude are adopted, and dynamic enhancement of features is achieved through adaptive adjustment coefficients, which not only preserves the detailed features of the wire but also enhances the contrast between it and the background. Through multi-scale feature fusion, while maintaining the integrity of the wire, the discriminability of the features is improved, making subsequent key point detection and matching more accurate and reliable. This embodiment improves the detection ability of slender wire structures in complex backgrounds, especially showing strong robustness when dealing with interference factors such as common illumination changes and perspective changes in aerial images.
[0036] According to one aspect of the present application, step S11 is further as follows:
[0037] S111. Obtain and based on the original aerial image data, extract spatial coordinate information to generate a position matrix; extract time series information to obtain time series data; separate image channels to obtain channel feature data; combine the position matrix, time series data, and channel feature data to construct a four-dimensional feature matrix;
[0038] S112. Slice the four-dimensional feature matrix. Obtain spatial slice data by slicing along the spatial dimension, temporal slice data by slicing along the temporal dimension, and channel slice data by slicing along the channel dimension; calculate the local statistics of the spatial slice data, temporal slice data, and channel slice data, and output the slice statistical data;
[0039] S113. Based on the slice statistical data, calculate the first standard deviation of each slice to obtain a standard deviation sequence; based on the slice statistical data, calculate the first mean of each slice to obtain a mean sequence; divide the standard deviation sequence by the mean sequence to generate a ratio sequence; perform normalization processing on the ratio sequence, and output the adaptive weight coefficient;
[0040] S114. Based on the four-dimensional feature matrix and the adaptive weight coefficient, calculate the projection of the feature matrix in each dimension to obtain dimension projection data; perform weighted combination on the dimension projection data and the adaptive weight coefficient to generate weighted projection data; perform dimensionality reduction transformation on the weighted projection data, and output three-dimensional projection feature data.
[0041] In this embodiment, multi-dimensional feature encoding of aerial image data is achieved by constructing a four-dimensional feature matrix. By extracting spatial coordinate information, time series information, and channel feature information, a four-dimensional matrix containing complete feature information is constructed, which can comprehensively capture the spatial structure and temporal change characteristics of the wire. During the feature processing, by performing multi-dimensional slicing on the matrix, spatial slice data, temporal slice data, and channel slice data are obtained respectively, and the local statistics of each slice are calculated, which can effectively suppress noise interference. By calculating the ratio relationship between the standard deviation sequence and the mean sequence, an adaptive weight coefficient is generated, which can automatically adjust the contribution of each dimension feature according to the importance of the feature. Through feature projection and weighted combination, feature dimensionality reduction from four dimensions to three dimensions is achieved while maintaining the key feature information of the wire. This embodiment provides a rich and reliable feature basis for subsequent feature processing through multi-dimensional feature encoding and adaptive weight adjustment.
[0042] According to one aspect of the present application, step S12 is further as follows:
[0043] S121. Receive the three-dimensional projection feature data, calculate the feature distribution within the local window to obtain distribution feature data; statistically calculate the local variance of the features to generate a variance matrix; perform normalization processing on the variance matrix, and output the feature variance data;
[0044] S122. Read the three-dimensional projection feature data, calculate the gradients in the horizontal and vertical directions to obtain direction gradient data; calculate the gradient magnitude based on the direction gradient data to generate a gradient magnitude map; extract the main direction feature of the gradient, and output the gradient feature data;
[0045] S123. Obtain the feature variance data, calculate the maximum and minimum values of the local area to obtain the extreme value data; construct an adaptive mapping function based on the extreme value data to generate the mapping coefficient data; convert the mapping coefficient data into the enhancement coefficient α and output the Alpha parameter data.
[0046] S124. Receive the gradient feature data, calculate the distribution histogram of the gradient amplitude to obtain the amplitude distribution data; extract the median feature of the distribution to generate the statistical feature data; calculate the adjustment coefficient γ based on the statistical feature data and output the Gamma parameter data.
[0047] S125. Read the three-dimensional projection feature data, calculate the second derivative of the feature to obtain the second-order gradient data; extract the curvature information of the feature to generate the curvature feature data; combine the second-order gradient data and the curvature feature data and output the enhancement basis data.
[0048] S126. Obtain the Alpha parameter data, Gamma parameter data and enhancement basis data, construct a feature enhancement function to obtain the enhancement function data; calculate the enhancement weight distribution to generate the weight distribution data; perform adaptive enhancement processing on the feature and output the enhanced feature data.
[0049] In this embodiment, effective enhancement of the projection feature data is achieved through feature variance calculation and gradient analysis. By calculating the feature distribution and variance matrix within the local window, the local statistical characteristics of the feature are obtained, which can accurately reflect the local changes of the wire structure. During the gradient feature extraction process, by calculating the gradient information in the horizontal and vertical directions, a complete gradient feature description is constructed, and the adjustment coefficient is determined through amplitude distribution analysis, which can effectively enhance the edge features of the wire. By combining the feature variance and gradient information, an adaptive feature enhancement function is constructed, which can dynamically adjust the enhancement degree according to the importance of the local feature, not only maintaining the detailed features of the wire but also improving its contrast with the background. Through feature statistical analysis and adaptive enhancement in this embodiment, the discriminability of the wire features is improved, providing a reliable feature basis for subsequent key point detection.
[0050] As Figure 3 shown, according to one aspect of the present application, step S2 is further as follows:
[0051] S21. Reorganize the enhanced feature map data into feature matrix data including spatial position and feature dimension; calculate the eigenvectors of the feature matrix data to obtain the left eigenvector matrix and the right eigenvector matrix; based on the left eigenvector matrix and the right eigenvector matrix, calculate the eigenvalue diagonal matrix, and process the residual term through a non-linear mapping function to output the decomposed feature data.
[0052] S22. Calculate the spatial attention mapping value and the feature attention mapping value based on the left eigenvector matrix and the right eigenvector matrix in the decomposed feature data; perform matrix multiplication on the spatial attention mapping value and the feature attention mapping value to generate a key point response matrix; construct a learnable weight matrix to adjust the key point response matrix and output response matrix data;
[0053] S23. Calculate the second mean, the second standard deviation and the information entropy based on the response matrix data; calculate the dynamic coefficient based on the information entropy; combine the second mean, the second standard deviation and the dynamic coefficient to generate adaptive threshold data; filter the response matrix data based on the adaptive threshold data and output candidate point set data;
[0054] S24. Extract local neighborhood features for each candidate point based on the candidate point set data; calculate the gradient and the second derivative of the local neighborhood features to generate adaptive convolution kernel data; obtain the image size and calculate the position encoding value; combine the position encoding value, the adaptive convolution kernel data and the local neighborhood features to generate a feature description set;
[0055] S25. Calculate the local saliency score and the spatial distribution score for each point based on the candidate point set data and the feature description set; calculate the balance factor based on the variance value of the feature description set; perform weighted combination on the local saliency score and the spatial distribution score based on the balance factor to obtain a weighted combination score; finally filter the points through a dynamically screened preset threshold based on the weighted combination score and output the key point set.
[0056] In an embodiment of the present application, the feature map F_out is reorganized into a feature matrix M(p, q), where p represents the spatial position and q represents the feature dimension; use Bi-directional Recursive Feature Decomposition (BRFD): M = U∑V T + R, where U is the left eigenvector matrix, ∑ is the eigenvalue diagonal matrix, V is the right eigenvector matrix, and R is the residual matrix: R = H(M - U∑V T ), H is the non-linear mapping function: H(x) = tanh(ωx) / (1 + e -φx ), ω, φ are adaptive parameters.
[0057] Construct the key point response matrix K: K = Ψ(U) * Ω(V); where Ψ(U) is the spatial attention mapping: Ψ(U)= softmax(U T * W_u * U), Ω(V) is the feature attention mapping: Ω(V) = softmax(V T* W_v * V) W_u, where W_u and W_v are learnable weight matrices. Introduce the Dynamic Threshold Optimization Network (DTON) to calculate the adaptive threshold τ: τ = μ(K) + Δ * σ(K), where Δ = log(1 + entropy(K)), and entropy(K) is the information entropy; generate the candidate point set C: C = {(x, y) | K(x, y) > τ}.
[0058] For each candidate point c ∈ C, calculate its feature descriptor D(c): D(c) = Φ(N(c)) * Γ(c), where N(c) is the local neighborhood of point c, Φ is the local feature extraction operator: Φ(x) = conv(x, g(x)), g(x) is the adaptive convolution kernel: g(x) = normalize(▽ 2 x + |▽x|), Γ(c) is the position encoding: Γ(c) = [sin(πx / W), cos(πx / W), sin(πy / H), cos(πy / H)], where W and H represent the width and height of the feature map respectively. Finally, output the optimized key point set P and the corresponding feature description set D: P = {p | p ∈ C, Q(p) > η}, where Q(p) is the point quality score: Q(p) = ρ * L(p) + (1 - ρ) * S(p), L(p) is the local significance score, S(p) is the spatial distribution score, ρ is the balance factor: ρ = sigmoid(var(D(p))), D(p) is the feature descriptor of point p, var represents variance, and η is the dynamic screening threshold: η = median(Q) + std(Q), where median represents the median and std represents the standard deviation.
[0059] In another embodiment of the present application, the process of feature recursive decomposition calculation is specifically: D(F) = U·Σ·V T + R(F); where R(F) = ψ(F - U·Σ·V T ), R(F) is the non - linear residual term; ψ(x) = sign(x)·log(1 + |x|) is the non - linear mapping function; U is the left eigenvector matrix; V is the right eigenvector matrix; Σ is the eigenvalue diagonal matrix; F is the input feature matrix; sign(x) is the sign function; |x| represents the absolute value.
[0060] In this embodiment, an efficient feature recombination and key point extraction framework is constructed by designing a recursive decomposition and mapping strategy. First, the enhanced feature map is recombined into a feature matrix containing spatial positions and feature dimensions. The effective decomposition of features is achieved by calculating the left and right eigenvector matrices, which not only reduces the computational complexity but also retains the main structural information of the wire. In the feature mapping stage, a dual mapping mechanism of spatial attention and feature attention is adopted, and the adaptive enhancement of key regions is realized through a learnable weight matrix, which is particularly suitable for processing the feature extraction of slender wire structures in aerial images. In the key point screening stage, an adaptive threshold mechanism based on information entropy is introduced, and the intelligent screening of potential key points is realized by comprehensively considering the mean, standard deviation, and information entropy, improving the positioning accuracy of key points. In the final feature enhancement stage, a robust feature descriptor is constructed by extracting local neighborhood features and combining position encoding. This descriptor can not only accurately depict the local geometric features of the wire but also has strong rotation and scale invariance. Through the synergistic effect of feature decomposition, dual attention mechanism, and adaptive threshold, the detection accuracy of wire key points and the discriminability of feature description in complex scenes are improved in this embodiment.
[0061] According to one aspect of the present application, step S22 is further as follows:
[0062] S221. Read the decomposed feature data, extract the left eigenvector matrix and the right eigenvector matrix; perform scale normalization processing on the two matrices to obtain the normalized left matrix and the normalized right matrix; calculate the statistical features of the matrices, and output the feature statistical data;
[0063] S222. Obtain the normalized left matrix, calculate the relative relationship of spatial positions to obtain the position relationship data; construct a local connection map, generate the connection map data; calculate the spatial attention distribution based on the position relationship data and the connection map data, and output the spatial attention mapping value;
[0064] S223. Receive the normalized right matrix, calculate the correlation between feature channels to obtain the channel correlation matrix; extract the main patterns of the features, generate the pattern feature data; calculate the feature attention distribution based on the channel correlation matrix and the pattern feature data, and output the feature attention mapping value;
[0065] S224. Read the spatial attention mapping value and the feature attention mapping value, construct the corresponding relationship of the mapping values to obtain the corresponding relationship data; calculate the matching degree of the mapping values, generate the matching degree data; perform matrix multiplication on the corresponding relationship data and the matching degree data, and output the initial key point response matrix;
[0066] S225. Obtain the initial key point response matrix, construct a dynamic weight learning network to obtain learning network data; calculate the response distribution of the network to generate distribution feature data; construct a learnable weight matrix based on the learning network data and the distribution feature data, and output weight matrix data;
[0067] S226. Receive the initial key point response matrix and the weight matrix data, calculate the fitness between the matrices to obtain fitness data; perform matrix adjustment operations to generate adjusted matrix data; perform normalization processing on the adjusted matrix data, and output response matrix data.
[0068] In this embodiment, by constructing a dual attention mechanism, efficient mapping and enhancement of the decomposed features are realized. First, by performing scale normalization processing on the left and right feature vector matrices, a stable feature basis is established, which can reduce the influence of feature scale changes. During the spatial attention calculation process, by constructing a local spatial relationship graph and calculating the relative distance, an adaptive evaluation of the importance of spatial positions is realized, which can highlight the key structural features of the wire. In the feature attention calculation stage, by analyzing the correlation and main direction information between feature channels, an attention mapping at the feature level is constructed, which can simultaneously focus on the importance of spatial structure and feature expression. Finally, the response matrix is dynamically adjusted through a learnable weight matrix to achieve adaptive enhancement of the features. Through the synergistic effect of spatial attention and feature attention in this embodiment, the expression ability and discriminability of the wire features are improved.
[0069] According to one aspect of the present application, step S23 is further as follows:
[0070] S231. Receive the response matrix data, divide it according to spatial regions to obtain a sequence of region matrices; calculate the pixel distribution of each region to obtain distribution feature data; perform statistical analysis on the distribution feature data, and output statistical basic data;
[0071] S232. Read the statistical basic data, calculate the mean of the local region to obtain a local mean matrix; extract the global mean feature to generate global mean data; perform weighted combination of the local mean matrix and the global mean data, and output comprehensive mean data;
[0072] S233. Obtain the statistical basic data, calculate the local standard deviation to obtain a local variance matrix; extract the global dispersion feature to generate dispersion data; perform normalization processing based on the local variance matrix and the dispersion data, and output standard deviation data;
[0073] S234. Receive the statistical basic data, calculate the probability distribution of the local region to obtain a probability distribution matrix; calculate the information entropy based on the probability distribution to generate local entropy value data; perform spatial aggregation on the local entropy value data, and output information entropy data;
[0074] S235. Read the information entropy data, construct a dynamic mapping function to obtain mapping function data; calculate the change trend of the entropy value to generate change trend data; calculate coefficients based on the mapping function data and the change trend data, and output dynamic coefficient data;
[0075] S236. Obtain the comprehensive mean data, standard deviation data, and dynamic coefficient data, construct an adaptive combination function to obtain combination function data; calculate the weight distribution of the parameters to generate weight distribution data; perform weighted combination on the three types of data, and output adaptive threshold data;
[0076] S237. Receive the response matrix data and the adaptive threshold data, calculate the difference between the response value and the threshold to obtain difference matrix data; extract the positions that meet the threshold conditions to generate position index data; extract the corresponding response values based on the position index data, and output the candidate point set data.
[0077] In this embodiment, through the construction of the adaptive threshold, the accurate screening of the response matrix is realized. First, through the regional division and statistical analysis of the response matrix, detailed distribution characteristic information is obtained, which can accurately capture the change characteristics of the local response. In the process of calculating the threshold parameters, by comprehensively considering the local mean, standard deviation, and information entropy, a multi-dimensional evaluation system is established, which can adaptively determine the screening criteria. Through the introduction of the dynamic coefficient, the adaptive adjustment of the threshold is realized, and the strictness of the screening criteria can be automatically adjusted according to the complexity of the image content. In the screening process, by calculating the difference between the response value and the threshold, the candidate key points are accurately extracted, which not only ensures the reliability of the key points but also avoids excessive false detections. In this embodiment, through multi-dimensional statistical feature analysis and dynamic threshold adjustment, the accurate screening of the key points is realized.
[0078] According to one aspect of the present application, step S24 is further:
[0079] S241. Based on the candidate point set data, extract multi-scale neighborhood blocks centered on each candidate point to generate a neighborhood block sequence; perform direction normalization processing on the neighborhood block sequence to obtain normalized neighborhood data; calculate the gray distribution characteristics of the normalized neighborhood data, and output neighborhood feature data;
[0080] S242. Based on the neighborhood feature data, calculate the first-order gradients in the horizontal and vertical directions to obtain a gradient vector field; based on the gradient vector field, calculate the Laplacian operator response to generate a second-order derivative map; perform weighted combination on the gradient vector field and the second-order derivative map according to the intensity distribution, and output gradient feature data;
[0081] S243. Calculate the local response intensity and the local gradient direction histogram based on the gradient feature data to obtain the direction distribution data; construct an adaptive Gaussian filter based on the direction distribution data to generate filter coefficients; modulate the filter coefficients with the local response intensity and output the adaptive convolution kernel data;
[0082] S244. Obtain the image size parameters, generate the coordinate grids in the horizontal and vertical directions to obtain the position coordinate data; perform sine and cosine transforms on the position coordinate data to obtain the periodic encoding data; extract the global scale information based on the image size parameters; normalize the periodic encoding data based on the global scale information and output the position encoding values;
[0083] S245. Filter the neighborhood feature data based on the adaptive convolution kernel data to obtain the filtered feature data; concatenate the filtered feature data and the position encoding values in the channel dimension to generate the combined feature data; perform dimensional adjustment and normalization processing on the combined feature data and output the feature description set.
[0084] In an embodiment of the present application, the process of generating the adaptive convolution kernel is specifically: K(x, y) = λ(x, y)·D(x, y) + (1 - λ(x, y))·G(x, y); where D(x, y) = Σ(h(θ)·Δ(x·cos(θ)+y·sin(θ))), D(x, y) is the direction response kernel; G(x, y) = exp(-(x 2 +y 2 ) / (2σ_g 2 ))·(1 - x 2 / σ_x 2 -y 2 / σ_y 2 ), G(x, y) is the Gaussian-Laplacian kernel; λ(x, y) = sigmoid(||▽I(x, y)||) is the adaptive weight; h(θ) is the direction histogram; Δ(·) is the Dirac function; σ_g, σ_x, σ_y are the kernel function parameters; ▽I(x, y) is the image gradient.
[0085] In this embodiment, an accurate description of candidate points is achieved by constructing a multi-scale local feature descriptor. First, a stable feature extraction foundation is established by extracting multi-scale neighborhood blocks and performing orientation normalization, which can effectively handle the scale changes of the wire. During the gradient feature extraction process, a complete local structure description is constructed by calculating the first-order gradient and second-order derivative responses. Through the design of an adaptive Gaussian filter, the self-adaptive enhancement of features is realized. By introducing a position encoding mechanism, the global position information is organically integrated with the local features, which can better express the spatial structure relationship of the wire. Finally, through feature combination and normalization processing, a feature description set with strong discriminability is generated. In this embodiment, a feature descriptor that maintains both local details and contains global information is constructed through the combination of multi-scale feature extraction and position encoding.
[0086] In another embodiment of the present application, step S22 may also be:
[0087] S22a. Read the decomposed feature data, separate and obtain the left eigenvector matrix and the right eigenvector matrix; calculate the spatial distribution characteristics of the eigenvectors to obtain distribution feature data; perform normalization processing on the distribution feature data, and output the normalized feature data;
[0088] S22b. Obtain the left eigenvector matrix, construct a local spatial relationship graph to obtain spatial relationship data; calculate the relative distance of spatial positions to generate distance matrix data; based on the spatial relationship data and the distance matrix data, output spatial association data;
[0089] S22c. Receive the right eigenvector matrix, calculate the similarity between feature channels to obtain a channel similarity matrix; extract the main direction information of the features to generate direction feature data; combine the channel similarity matrix with the direction feature data, and output feature association data;
[0090] S22d. Read the spatial association data, construct a self-attention mapping network to obtain attention network data; calculate the attention weight distribution to generate weight distribution data; apply the attention weight to the spatial association data, and output a spatial attention mapping;
[0091] S22e. Obtain the feature association data, extract the local response pattern to obtain response pattern data; calculate the response intensity distribution to generate intensity distribution data; based on the response pattern data and the intensity distribution data, output a feature attention mapping;
[0092] S22f. Receive the spatial attention mapping and the feature attention mapping, calculate an interaction response matrix to obtain interaction matrix data; extract the main components of the response to generate main component data; perform feature enhancement on the interaction matrix data, and output enhanced response data;
[0093] S22g. Read the enhanced response data, construct a learnable weight matrix to obtain learning weight data; apply the weight matrix for feature transformation to generate transformed feature data; perform normalization processing on the transformed feature data to output a key point response matrix; adjust the key point response matrix through the learnable weight matrix to output response matrix data.
[0094] As Figure 4 shown, according to one aspect of the present application, step S3 is further as follows:
[0095] S31. Based on the key point set, calculate the feature entropy value and local variance value of each key point; based on the feature entropy value and local variance value, calculate the adaptive variance parameter, and construct an improved Gaussian distribution model; based on the Gaussian distribution model, generate and fuse local heat maps for each key point to output heat map data;
[0096] S32. Based on the heat map data and the feature description set, calculate the cosine similarity of the feature descriptions to obtain feature similarity data; obtain and calculate the distance matrix based on the spatial positions of the key points to generate position-related data; perform weighted combination on the feature similarity data and the position-related data to output a spatial association matrix;
[0097] S33. Based on the spatial association matrix, calculate the feature similarity, geometric constraint value, and linear consistency value between candidate endpoint pairs; based on the feature similarity, geometric constraint value, and linear consistency value, calculate the constraint mean data to generate the first adaptive weight; based on the first adaptive weight, perform weighted combination on the feature similarity, geometric constraint value, and linear consistency value to output matching score data;
[0098] S34. Based on the matching score data, calculate the topological consistency score for each pair of matching endpoints; based on the variance of the matching score data, calculate the dynamic balance factor; based on the dynamic balance factor, perform weighted combination on the matching score data and the topological consistency score to generate wire credibility data; based on the wire credibility data and the coordinates of the matching endpoints, construct a wire path, and use a path correction algorithm to generate wire path data;
[0099] S35. Based on the wire path data and the wire credibility data, calculate the visualization consistency score, region support score, and topological uniqueness score for each wire instance; based on the change degree of the visualization consistency score, region support score, and topological uniqueness score, determine the second adaptive weight; based on the second adaptive weight, perform weighted combination on the visualization consistency score, region support score, and topological uniqueness score to generate instance quality data; based on the instance quality data, screen and optimize the wire instances to output a wire instance set.
[0100] In an embodiment of the present application, based on the key point set P and the feature description set D, an Adaptive Multi-layer Heatmap Network (AMHN) is constructed. For each key point p∈P, an improved Gaussian heatmap H is generated: H(x, y, p) = exp(-((x - x_p) 2 + (y - y_p) 2 ) / (2σ_p 2 ))), where σ_p is the adaptive variance: σ_p = α*E(p) + β*V(p), E(p) is the feature entropy of the point: E(p) = -∑(D(p)_i * log(D(p)_i)), V(p) is the local variance: V(p) = var(N(p)), α and β are dynamic weights: α = sigmoid(E(p)), β = 1 - α, and N(p) is the local neighborhood of point p.
[0101] A Hierarchical Spatial Attention (HSA) mechanism is introduced to construct a spatial correlation matrix S: S(i, j) = ψ(D(p_i), D(p_j)) * φ(pos(p_i), pos(p_j)), where ψ is the feature similarity function: ψ(a, b) = (a·b) / (|a|·|b|); φ is the position correlation function: φ(a, b) = exp(-||a - b|| 2 / λ); λ is the adaptive scale parameter: λ = median(||pos(p_i) - pos(p_j)|| 2 ). A Multi-constraint Endpoint Matching Network (MEMN) is constructed. For each pair of candidate endpoints (p_i, p_j), a matching score M is calculated: M(i, j) = w_1·F(i, j) + w_2·G(i, j) + w_3·L(i, j); where F(i, j) is the feature similarity: F = S(i, j); G(i, j) is the geometric constraint term: G = exp(-|d(i, j) - μ_d| / σ_d); L(i, j) is the linear consistency: L = cos(θ(i, j)); d(i, j) is the distance between endpoints; θ(i, j) is the angle with the main direction; μ_d and σ_d are distance statistics; w_1, w_2, and w_3 are adaptive weights: [w_1, w_2, w_3] = softmax([F_avg, G_avg, L_avg]).
[0102] Build an Adaptive Wire Instance Generator (AWIG). For each matching pair (p_i, p_j): Calculate the wire credibility C: C(i, j) = ρ·M(i, j) + (1 - ρ)·T(i, j), where T is the topological consistency score and ρ is the dynamic balance factor: ρ = sigmoid(var(M)); Build the wire path W: W(i, j) = {x(t)| t∈[0, 1]} x(t) = p_i + t(p_j - p_i) + B(t), where B(t) is the path correction term: B(t) = κ·sin(πt)·n, κ is the bending coefficient and n is the normal vector.
[0103] Apply the Multi-verification Optimization Network (MVON). For each wire instance W: Calculate the instance quality score Q: Q(W) = α·V(W) + β·R(W) + γ·U(W), where V is the visualization consistency score, R is the region support score, and U is the topological uniqueness score; α, β, γ are adaptive weights.
[0104] In another embodiment of the present application, the process of generating a multi-layer heat map calculation is specifically: H(x, y) = Σ(w_k·G_k(x, y)); where G_k(x, y) = exp(-(|x - x_k| 2 +|y - y_k| 2 ) / (2σ_k 2 ))·s_k; w_k = exp(-E_k / E_max) / (Σexp(-E_i / E_max)), w_k is the adaptive weight of the k-th key point; G_k(x, y) is a Gaussian heat map centered on the key point (x_k, y_k); σ_k = β·sqrt(V_k) is the adaptive variance parameter; s_k is the confidence score of the key point; E_k is the feature entropy value of the key point, E_k = -Σ(p_i·log(p_i)); V_k is the local variance of the key point neighborhood; E_max is the maximum entropy value; β is the scale adjustment coefficient; p_i is the normalized probability of the feature histogram.
[0105] The specific process of outputting the spatial correlation matrix is as follows: The weighted fusion calculation formula for feature similarity and spatial distance is F(i, j) = α(t)·S_feat(i, j) + β(t)·S_pos(i, j) + γ(t)·S_dir(i, j); where S_feat(i, j) = cos(V(i), V(j)) = V(i)·V(j) / (||V(i)||·||V(j)||), and S_feat(i, j) is the cosine similarity of the feature descriptor; S_pos(i, j) = exp(-||P(i)-P(j)|| 2 / σ 2 ), S_pos(i, j) is the Gaussian similarity of the spatial position; S_dir(i, j) = |cos(θ(i)-θ(j))|, and S_dir(i, j) is the consistency measure of the direction vector; α(t), β(t), γ(t) are adaptive weight coefficients, dynamically adjusted through the statistical characteristics of the local region; V(i), V(j) are the feature description vectors of key points i, j; P(i), P(j) are the spatial coordinates of key points i, j; θ(i), θ(j) are the main direction angles at key points i, j; σ is the scale parameter of the Gaussian kernel, adaptively set according to the image resolution; ||·|| represents the Euclidean norm; t is the current processing time step.
[0106] The specific process of calculating the matching score with multi-dimensional constraints is as follows: S(p, q) = ω1·D(p, q) + ω2·G(p, q) + ω3·C(p, q); where D(p, q) = exp(-||f(p) - f(q)|| 2 / σ_f 2 ), D(p, q) is the similarity measure of the feature descriptor; G(p, q) = exp(-θ(p, q) 2 / σ_θ 2 )·exp(-d(p, q) 2 / σ_d 2 ), G(p, q) is the geometric constraint term; C(p, q) = exp(-|κ(p) - κ(q)| / σ_κ), and C(p, q) is the curvature consistency constraint; f(p), f(q) are the feature descriptors of endpoints p, q; θ(p, q) is the angle between the connection direction of the endpoints and the main direction; d(p, q) is the normalized distance between the endpoints; κ(p), κ(q) are the local curvatures at the endpoints; ω1, ω2, ω3 are the weights of the constraint terms, and Σωi = 1; σ_f, σ_θ, σ_d, σ_κ are the scale parameters of each constraint term.
[0107] In this embodiment, by constructing a multi-layer heat map and a spatial correlation matrix, accurate detection and matching of wire instances are achieved. First, a multi-layer heat map is generated based on the key point set. By calculating the feature entropy value and the local variance value, an improved Gaussian distribution model is constructed, which can better express the continuity and integrity of the wire structure. In the spatial feature enhancement stage, an effective spatial correlation mechanism is established by calculating the weighted combination of the cosine similarity of feature descriptions and the distance matrix, improving the connection accuracy of wire instances. In the multi-dimensional constraint matching stage, multiple constraints such as feature similarity, geometric constraints, and linear consistency are introduced. Through the dynamic adjustment of the adaptive weight, accurate matching of wire instances is achieved, effectively solving complex situations such as wire crossing and overlapping in aerial images. In the wire instance verification stage, a comprehensive quality evaluation system is constructed by calculating the topological consistency score and the visualization consistency score, which can effectively filter out incorrect matches and improve the reliability of wire detection. Through the synergistic effect of the heat map, spatial correlation, and multi-dimensional constraints, this embodiment realizes accurate detection and reliable matching of complex wire structures in aerial images.
[0108] According to one aspect of the present application, step S32 is further as follows:
[0109] S321. Receive heat map data and a set of feature descriptions, extract the direction information of the feature descriptions to obtain direction feature data; calculate the amplitude distribution of the features to generate amplitude distribution data; perform normalization processing on the direction feature data and the amplitude distribution data, and output normalized feature data;
[0110] S322. Read the normalized feature data, construct feature vector pairs to obtain a set of feature pairs; calculate the cosine distance between the feature pairs to generate distance mapping data; perform scale adjustment on the distance mapping data, and output feature similarity data;
[0111] S323. Obtain the spatial coordinate information of the key points, calculate the Euclidean distance between point pairs to obtain the original distance matrix data; construct a Gaussian kernel function for spatial distance to generate kernel function data; apply kernel function transformation to the original distance matrix data, and output position-related data;
[0112] S324. Receive the feature similarity data, calculate the similarity distribution in the local area to obtain local distribution data; extract the global similarity pattern to generate pattern feature data; construct an adaptive weight based on the local distribution data and the pattern feature data, and output similarity weight data;
[0113] S325. Read the position-related data, analyze the clustering characteristics of the spatial positions to obtain clustering feature data; calculate the continuity of the spatial distribution to generate continuity data; combine the clustering feature data and the continuity data to construct a spatial weight, and output position weight data;
[0114] S326. Obtain feature similarity data, position-related data, similarity weight data, and position weight data, construct a multi-dimensional association function to obtain association function data; calculate the fusion coefficients for each dimension to generate fusion coefficient data; perform weighted combination on the multi-dimensional features, and output a spatial association matrix.
[0115] In this embodiment, through feature similarity calculation and spatial association analysis, accurate matching of wire instances is achieved. First, by extracting the heat map data and the direction information of the feature description, the basis for feature matching is established, which can effectively handle the direction change of the wire. During the similarity calculation process, by constructing feature vector pairs and calculating the cosine distance, the similarity measurement at the feature level is realized. At the same time, the Gaussian kernel function transformation of the spatial distance is introduced, effectively combining feature similarity and spatial position constraints. By analyzing the similarity distribution in the local area and the clustering characteristics of the spatial position, an adaptive weight allocation mechanism is constructed, which can automatically adjust the importance of features in the matching process according to the reliability of the features. Finally, the feature similarity and position correlation are fused through a multi-dimensional association function to generate a comprehensive spatial association matrix. This embodiment improves the accuracy and reliability of wire instance matching through the organic combination of feature similarity and spatial constraints.
[0116] According to one aspect of the present application, step S33 is further as follows:
[0117] S331. Based on the spatial association matrix, extract the feature vectors of the candidate endpoint pairs, calculate the cosine distance of the feature vectors to obtain a similarity matrix; based on the similarity matrix, construct a feature distribution histogram to generate distribution feature data; perform normalization processing on the distribution feature data, and output feature similarity data;
[0118] S332. Based on the spatial association matrix, obtain the spatial coordinates of the candidate endpoint pairs, calculate the Euclidean distance between the endpoints to obtain a distance matrix; based on the distance matrix, statistically analyze the distance distribution characteristics in the local area to generate a distance statistic; compare the distance statistic with a preset threshold, and output distance constraint data;
[0119] S333. Based on the spatial association matrix, obtain the position information of the endpoint pairs, calculate the connection direction vector to obtain a direction matrix; based on the gradient magnitude in step S12, extract the main direction feature of the image to generate main direction data; calculate the angle between the direction matrix and the main direction data, and output direction constraint data;
[0120] S334. Based on the distance constraint data and the direction constraint data, extract the structural features of the local area to obtain structural feature data; based on the structural feature data, calculate the structural consistency score to generate structural constraint data; combine the distance constraint data, the direction constraint data, and the structural constraint data, and output a geometric constraint value;
[0121] S335. Based on the spatial correlation matrix, obtain the connection information of the endpoint pairs, calculate the curvature characteristics of the connections, and obtain curvature data; based on the connection information, extract the linear measurement characteristics of the connections, and generate linear feature data; based on the curvature data and the linear feature data, calculate the linear consistency score and output the linear constraint value;
[0122] S336. Based on the feature similarity data, geometric constraint value, and linear constraint value, calculate the mean of each constraint to obtain the constraint mean data; based on the constraint mean data, generate the second weight coefficient data; perform weighted combination on the feature similarity data, geometric constraint value, linear constraint value, and the second weight coefficient data, and output the matching score data.
[0123] In this embodiment, through multi-dimensional constraint evaluation, precise screening of candidate matching pairs is achieved. First, by analyzing the distribution characteristics of feature vectors and distance statistics, a basic matching evaluation system is established, which can comprehensively measure the reliability of the matching. In the process of geometric constraint calculation, by analyzing the direction relationship and spatial structure characteristics of endpoint pairs, a complete geometric constraint system is constructed, which can effectively filter out unreasonable matches. In the linear constraint evaluation stage, by analyzing the curvature characteristics and linear measurement characteristics of the connections, the rationality of the matching result is further ensured. Finally, through the dynamic allocation of adaptive weights, effective integration of each constraint is achieved, and a comprehensive matching score is generated. In this embodiment, a reliable matching evaluation system is constructed through the synergistic effect of feature similarity, geometric constraint, and linear constraint.
[0124] As Figure 5 shown, according to one aspect of the present application, step S4 is further as follows:
[0125] S41. Obtain the boundary region data of adjacent picture blocks in the wire instance set, extract the boundary features, and generate a boundary feature matrix; based on the boundary feature matrix, calculate the inner product and norm of the boundary feature vectors of adjacent blocks to obtain feature vector data; perform normalization processing on the feature vector data and output the boundary alignment score data;
[0126] S42. Based on the wire instance set, extract the wire endpoint information at the boundary; based on the boundary alignment score data and the wire endpoint information, calculate the included angle of the direction vectors between endpoint pairs to obtain the direction consistency data; based on the neighborhood region of the wire endpoints, calculate the spatial overlap degree and generate the overlap degree data; based on the wire endpoint information, extract the curvature characteristics of the endpoints and calculate the difference value, and output the curvature feature data; based on the mean of the direction consistency data, overlap degree data, and curvature feature data, calculate the adaptive weight and then combine to generate cross-block matching data;
[0127] S43. Based on the cross-block matching data, construct an initial graph structure containing all endpoints and connection relationships; calculate the structural support, path smoothness, and global consistency values of each edge in the initial graph structure to obtain edge attribute data; calculate the dynamic weight coefficients based on the variances of the attributes in the edge attribute data; based on the dynamic weight coefficients, perform weighted combination on the edge attribute data, and output the topological credibility data.
[0128] S44. Based on the topological credibility data and the wire instance set, calculate the output values of a predetermined number of fusion feature evaluation functions to obtain the evaluation results; generate fusion parameter data based on the evaluation results; based on the fusion parameter data, apply the fusion operator family to each pair of wire instances for feature fusion, and output the fusion instance data.
[0129] S45. Based on the fusion instance data, calculate the output values of a predetermined number of quality evaluation functions to generate quality index data; calculate the information entropy of each quality index based on the quality index data to obtain the first weight coefficient data; perform weighted combination on the quality index data and the first weight coefficient data to generate a confidence map; based on the confidence map, optimize the fusion instance data, and output the global wire graph.
[0130] In an embodiment of the present application, based on the wire instance set {W_k} in each picture block, where k represents the block index, construct a Boundary Feature Alignment Network (BFAN), and for the boundary region of adjacent picture blocks (k, l): construct a boundary feature matrix B: B(k, l) = Ω(F_k) ∩ Ω(F_l), where Ω(F) is the boundary feature extraction operator, and F_k, F_l are the feature maps of adjacent blocks; calculate the boundary alignment score A: A(k, l) = ∑(μ_i * ν_i) / (|μ_i|·|ν_i|); μ_i, ν_i are the feature vectors of the corresponding boundary points. Construct a Cross-block Instance Association Network (CIAN), and for the wire endpoint pair (p_i, p_j) at the boundary: calculate the cross-block matching score M_c: M_c(i, j) = w_1·D(i, j) + w_2·O(i, j) + w_3·T(i, j), where D is the direction consistency: D = cos(θ_i - θ_j); O is the spatial overlap degree: O = IoU(R_i, R_j); T is the topological continuity: T = exp(-|c_i - c_j| / σ); θ is the wire direction vector, R is the endpoint neighborhood region, c is the curvature feature; w_1, w_2, w_3 are the adaptive weights: [w_1, w_2, w_3] = softmax([D_avg, O_avg, T_avg]).
[0131] Construct a Dynamic Topology Optimization Network (DTON) to generate a global wire graph G(V, E), where V is the set of all endpoints and E is the wire connection relationship. For each edge e ∈ E, calculate the topology credibility C_t: C_t(e) = α·S(e) + β·P(e) + γ·Q(e), where S is the structure support, P is the path smoothness, and Q is the global consistency; α, β, and γ are dynamic weights: [α, β, γ] = normalize([var(S), var(P), var(Q)]). Construct an Adaptive Instance Fusion Network (AIFN). For each pair of instances to be fused (W_i, W_j), calculate the fusion parameters: F(i, j) = {λ_k | k ∈ [1, K]}; λ_k = softmax(ψ_k(W_i, W_j)); ψ_k is the k-th fusion feature evaluation function. Generate the fused instance W_f: W_f = ∑(λ_k * Φ_k(W_i, W_j)), where Φ_k is a family of fusion operators. Construct a Multi-dimensional Quality Assessment Network (MQAN) to evaluate the global result. Calculate the quality metric set Q: Q = {q_1, q_2,..., q_n}; q_i = Ψ_i(G, {W_f}), where Ψ_i is the i-th quality evaluation function. Generate the final confidence map C: C(x, y) = ∑(w_i * q_i(x, y)), w_i = softmax(entropy(q_i)), where entropy is the entropy.
[0132] In another embodiment of the present application, the process of calculating the boundary alignment score is specifically: B(i, j) = ρ·A(i, j) + (1 - ρ)·M(i, j); where A(i, j) = (f_i·f_j T ) / (||f_i||·||f_j||), A(i, j) is the inner product similarity of the boundary feature vectors; M(i, j) = exp(-d(i, j) 2 / σ_m 2 )·exp(-|θ_i - θ_j| / σ_θ), M(i, j) is the position-direction matching degree; f_i and f_j are boundary feature vectors; d(i, j) is the distance between the boundary point pairs; θ_i and θ_j are the direction angles at the boundary; ρ is the balance coefficient; σ_m and σ_θ are scale parameters.
[0133] The process of optimizing the fusion parameters of cross-block instances is specifically as follows: M(i, j) = Δ(t)·H(i, j) + ε(t)·O(i, j) + ζ(t)·T(i, j); where H(i, j) = Σ(min(h_i(k), h_j(k))) / Σ(max(h_i(k), h_j(k))), and H(i, j) is the overlap measure of the heat map; O(i, j) = A_overlap / (A_i + A_j - A_overlap), and O(i, j) is the spatial overlap rate; T(i, j) = exp(-||τ_i - τ_j|| 2 / σ_τ 2 ), and T(i, j) is the topological similarity; h_i(k) and h_j(k) are the heat map values of instances i and j; A_overlap is the area of the overlapping region; A_i and A_j are the instance areas; τ_i and τ_j are the topological feature vectors; Δ(t), ε(t), and ζ(t) are time-varying weight coefficients; σ_τ is the scale parameter of the topological feature.
[0134] The process of topological credibility evaluation is specifically as follows: T(e) = φ(t)·S(e) + χ(t)·P(e) + ψ(t)·C(e); where S(e) = Σ(exp(-d(e, n) / σ_d)), and S(e) is the structural support of edge e; P(e)=exp(-|κ_max(e)| / σ_κ)·exp(-Δθ(e) / σ_θ), and P(e) is the path smoothness; C(e) = min(Σw_i·c_i(e), Σw_j·c_j(e)), and C(e) is the global consistency value; d(e, n) is the distance from edge e to the neighboring structure n; κ_max(e) is the maximum curvature of edge e; Δθ(e) is the direction change amount; c_i(e) and c_j(e) are the local consistency measures; φ(t), χ(t), and ψ(t) are time-varying weight coefficients; w_i and w_j are the local weights.
[0135] In this embodiment, through boundary feature analysis and cross-block instance matching, the accurate reconstruction of the global topological structure of the wire is achieved. First, feature extraction and alignment analysis are performed on the boundary regions of adjacent image blocks. By calculating the inner product and norm of the boundary feature vectors, an effective boundary alignment mechanism is established, which can accurately handle the splicing relationship between image blocks. In the cross-block instance matching stage, considering direction consistency, spatial overlap degree, and curvature features comprehensively, a multi-dimensional matching criterion is constructed. Through the dynamic adjustment of adaptive weights, the accurate matching of cross-block wire instances is achieved, effectively solving the problem of wire cross-region continuity in large-scale aerial images. In the topological optimization stage, by calculating structural support, path smoothness, and global consistency, a complete graph structure optimization framework is established, which can effectively eliminate incorrect connections and improve the accuracy of the wire topological structure. In the final instance fusion stage, by constructing multiple fusion feature evaluation functions, the intelligent fusion of wire instances in the overlapping region is realized, and the reliability of the fusion result is ensured through quality evaluation. Through the organic combination of boundary analysis, cross-block matching, and topological optimization, this embodiment realizes the global detection and reconstruction of wires in large-scale aerial images.
[0136] According to one aspect of the present application, step S42 is further as follows:
[0137] S421. Receive the boundary alignment score data and the wire endpoint information at the boundary, extract the spatial coordinates of the endpoints to obtain endpoint coordinate data; calculate the connection vectors of adjacent endpoints to generate connection vector data; perform unit normalization processing on the connection vector data and output normalized direction data;
[0138] S422. Read the normalized direction data, construct a direction matrix for the endpoint pairs to obtain direction matrix data; calculate the angles between the direction vectors to generate angle data; calculate the direction consistency score based on the angle data and output direction consistency data;
[0139] S423. Obtain the spatial coordinates of the endpoints, extract the local neighborhood centered on the endpoints to obtain neighborhood block data; calculate the feature distribution of the neighborhood blocks to generate distribution feature data; perform feature matching on adjacent neighborhood blocks and output neighborhood matching data;
[0140] S424. Receive the neighborhood matching data, calculate the overlapping regions between the neighborhood blocks to obtain overlapping region data; count the feature consistency of the overlapping regions to generate consistency data; calculate the overlap degree score based on the overlapping region data and the consistency data and output overlap degree data;
[0141] S425. Read the path information near the endpoints, extract the local path segments to obtain path segment data; calculate the discrete curvature of the path to generate discrete curvature data; perform smoothing processing on the discrete curvature data and output curvature feature data;
[0142] S426. Obtain direction consistency data, overlap degree data, and curvature feature data, calculate the statistical distribution of various types of data to obtain statistical feature data; extract the mean feature of the data to generate mean feature data; construct a weight coefficient based on the statistical feature data and the mean feature data, and output the weight coefficient data.
[0143] S427. Receive the three types of feature data and the weight coefficient data, construct a feature fusion function to obtain fusion function data; calculate the weighted combination of the features to generate combined feature data; perform normalization processing on the combined feature data, and output cross-block matching data.
[0144] In this embodiment, through boundary feature analysis and cross-block matching, the continuous reconstruction of the wire between adjacent image blocks is realized. First, by extracting the spatial coordinates and direction information of the endpoints, the basis for cross-block matching is established, and the connection relationship between image blocks can be accurately processed. During the direction consistency analysis process, by calculating the direction angle between endpoint pairs and local neighborhood features, a complete direction constraint system is constructed, which can effectively identify reasonable cross-block connections. By analyzing the overlap features of neighborhood blocks and the curvature features of paths, the continuity and smoothness of cross-block matching are further ensured. Finally, through the dynamic adjustment of the adaptive weight, the effective fusion of multi-dimensional features is realized, and a reliable cross-block matching result is generated. This embodiment improves the accuracy of cross-block wire reconstruction through the comprehensive analysis of direction consistency, overlap degree, and curvature features.
[0145] According to one aspect of the present application, step S43 is further as follows:
[0146] S431. Based on the cross-block matching data, extract the spatial coordinates of the endpoints to generate a node coordinate set; based on the cross-block matching data and the node coordinate set, extract the connection relationship between the endpoints, construct an adjacency matrix, and obtain adjacency relationship data; combine the node coordinate set and the adjacency relationship data to construct an initial graph structure, and output the initial graph data.
[0147] S432. Based on the initial graph data, calculate the degree distribution feature of each node to obtain degree distribution data; based on the initial graph data and the degree distribution data, count the local connection pattern to generate connection pattern data; based on the degree distribution data and the connection pattern data, calculate the structural importance of each node, and output the node weight data.
[0148] S433. Based on the initial graph data and the node weight data, calculate the weight of the node pair connected by each edge to obtain the initial value of the edge weight; based on the initial graph data and the node weight data, extract the direction vector feature of the edge to generate edge direction data; combine the initial value of the edge weight and the edge direction data, and output the basic attribute data of the edge.
[0149] S434. Calculate the local support degree for each edge based on the edge basic attribute data to obtain support degree data; statistically analyze the local distribution characteristics of the edges based on the edge basic attribute data to generate distribution characteristic data; combine the support degree data and the distribution characteristic data and output the structure support degree value;
[0150] S435. Calculate the curvature change of the path based on the edge basic attribute data to obtain curvature change data; extract the turning point characteristics of the path based on the edge basic attribute data and the curvature change data to generate turning point characteristic data; evaluate the path smoothness based on the curvature change data and the turning point characteristic data and output the smoothness value;
[0151] S436. Construct a global consistency evaluation matrix based on the structure support degree value and the smoothness value to obtain a consistency matrix; calculate the eigenvalue distribution of the consistency matrix to generate eigenvalue data; calculate the global consistency score based on the consistency matrix and the eigenvalue data and output the consistency value;
[0152] S437. Calculate the variance of each attribute based on the structure support degree value, the smoothness value and the consistency value to obtain variance statistical data; generate the third weight coefficient data based on the variance statistical data; perform weighted combination on the structure support degree value, the smoothness value, the consistency value and the third weight coefficient data and output the topology credibility data.
[0153] In this embodiment, through constructing a global topology graph structure, the overall optimization of the wire network is realized. First, an initial graph structure is constructed by extracting the endpoint coordinates and adjacency relationships, providing a theoretical basis for the global optimization of the wire network. During the node importance evaluation process, a node weight system is constructed by analyzing the degree distribution characteristics and local connection patterns, which can highlight the importance of key nodes. In the edge attribute calculation stage, a complete edge evaluation system is established by comprehensively considering the structure support degree, path smoothness and global consistency, which can accurately reflect the reliability of the edges. Finally, through the calculation of dynamic weight coefficients, the adaptive fusion of various attributes is realized, and a reliable topology credibility evaluation result is generated. Through the synergistic effect of node weight analysis and edge attribute evaluation in this embodiment, the global consistency and reliability of the wire network reconstruction are improved. Especially when dealing with complex wire crossing and branching situations, the integrity and continuity of the network structure can be effectively maintained.
[0154] According to one aspect of the present application, step S44 is further as follows:
[0155] S441. Obtain the topology credibility data and the wire instance pairs to be fused, extract the spatial coordinate sequences of the instances to obtain coordinate sequence data; calculate the spatial overlapping regions of the instance pairs to generate overlapping region data; construct the feature description of the instance pairs based on the coordinate sequence data and the overlapping region data and output the instance feature data;
[0156] S442. Read the instance feature data, calculate the local shape features of the instance to obtain the shape feature matrix; extract the direction change features of the instance to generate the direction change data; combine the shape feature matrix and the direction change data, and after normalization processing, output the morphological description data;
[0157] S443. Receive the instance feature data and the morphological description data, calculate the similarity score between instances to obtain the similarity matrix; generate the structural feature data based on the structural features of the local area; combine the similarity matrix and the structural feature data, and output the feature evaluation value;
[0158] S444. Obtain the feature evaluation value, construct a multi-layer evaluation network to obtain the evaluation network data; calculate the response values of each layer of the network to generate a response value sequence; perform weighted combination on the response value sequence, and output the fusion parameter data;
[0159] S445. Read the instance feature data and the fusion parameter data of the instance pair to be fused, construct an adaptive fusion operator according to the parameters to obtain the fusion operator data; calculate the weight distribution of the fusion operator to generate the weight distribution data; combine the fusion operator data and the weight distribution data, and output the fusion operation data;
[0160] S446. Receive the instance feature data and the fusion operation data, perform feature reconstruction on the instance to obtain the reconstructed feature data; calculate the consistency constraint of the features to generate the constraint condition data; perform feature optimization based on the reconstructed feature data and the constraint condition data, and output the fused instance data.
[0161] In this embodiment, through feature fusion and instance optimization, the accurate reconstruction of wire instances is achieved. First, by extracting the spatial coordinate sequence and overlapping region information of the instances, the basis for instance fusion is established, and the overlapping and connection relationships between instances can be accurately processed. In the process of morphological feature analysis, by calculating the local shape features and direction change features, a complete morphological description system is constructed, which can accurately depict the geometric characteristics of the wire. In the fusion parameter calculation stage, by constructing a multi-layer evaluation network, the adaptive optimization of the fusion parameters is realized, and the fusion method can be automatically adjusted according to the instance features. Finally, through the design and application of the fusion operator, the accurate reconstruction of the instance features is achieved, and the rationality of the reconstruction result is ensured by introducing the constraint conditions. This embodiment improves the accuracy and reliability of wire instance reconstruction by combining feature reconstruction and constraint optimization. Especially when dealing with the overlapping and cross regions between instances, the continuity and geometric characteristics of the wire can be maintained, avoiding the problems of breakage and deformation that are prone to occur in traditional methods. By introducing an adaptive fusion operator and a multi-layer evaluation mechanism, this embodiment also has strong scene adaptability and can handle the reconstruction problems of different types and complexities of wire instances.
[0162] In an embodiment of the present application, a wire instance detection method based on key point matching is proposed. The resolution of the original captured image is 6k×8k. By analyzing the wire distribution block by block through the method of image tiling, there are two advantages: 1) For extremely slender and complexly distributed wires, tiling alleviates the difficulty of wire analysis; 2) In a local image block, it is approximately considered that the wire is a straight line. Therefore, only two endpoints of the wire need to be detected to complete wire instantiation. Feature extraction is performed. This module mainly includes high-resolution feature extraction by HRNet, global feature perception by the attention mechanism, and multi-resolution feature fusion. In the prediction stage, a key point regression heat map related to the wire position, a wire feature description map, and a wire position mask map are predicted simultaneously. The wire endpoints are matched through the predicted key points and the corresponding feature description maps to achieve wire instantiation. The specific steps are as follows:
[0163] S1. Use a convolutional neural network to construct a feature extraction architecture and incorporate a global perception mechanism to ensure accurate capture of multi-scale and multi-level features.
[0164] Considering the slender wires and the pixel-level width distribution characteristics, HRNet and the self-attention mechanism are used to achieve high-resolution feature extraction and global feature perception of features. HRNet is a network structure for image processing tasks (such as pose estimation, semantic segmentation, etc.). By maintaining a high-resolution feature map, it can better capture detailed information. Different from traditional convolutional neural networks, HRNet maintains a high-resolution feature representation throughout the network and enhances the feature expression ability through multi-scale fusion. In this embodiment, HRNet outputs four-scale feature maps, respectively achieving feature map extraction with 4-fold to 32-fold downsampling.
[0165] The wire distribution often spans the entire image. To enhance the global perception ability of the network, the self-attention mechanism in the Transformer model is referred to to achieve global perception of pixel points. Specifically, the 4-fold downsampled feature map is rearranged into a sequence. The rotational position encoding is used to enhance the model's perception ability of context positions. The self-attention mechanism is used to calculate the feature representation of each position in the model, while paying attention to the features of all positions in the input sequence, thereby capturing long-range dependencies.
[0166] Rotational position encoding is an encoding technique for the Transformer model. By injecting position information into the input sequence, it enhances the model's context understanding ability. Traditional absolute position encoding (such as sin / cos encoding) is directly added to the input embedding vector, while rotational position encoding performs element-wise rotation of the input embedding vector and the position encoding to represent the position relationship in the high-dimensional space.
[0167] For the input vector x and the position vector θ, the rotational position encoding formula is as follows:
[0168] ;
[0169] where x even represents the elements at even positions in the input vector; x odd represents the elements at odd positions in the input vector; θ is the rotation angle generated by the position. This way enables the attention mechanism of the model at different positions to better capture the relative position relationships.
[0170] The self-attention mechanism is the core component of the Transformer model. When calculating the representation of each word, the model simultaneously focuses on all words in the input sequence, thereby capturing long-range dependencies. By calculating the correlation between each word in the input sequence and other words (i.e., attention weights), the self-attention mechanism can flexibly focus on the most relevant information. The calculation process of the self-attention mechanism is as follows: Calculate the query Q (query), key K (key), and value V (value): Q = XW Q ; K = XW K ; V = XW V ; where X is the representation matrix of the input sequence, and W Q , W K , and W V are trainable matrix weights. Calculate the attention output: selfAtten(x) = softmax(QK T / sqrt(d k )) · V = softmax(XW Q · (XW K ) T / sqrt(d k )) · XW V ; where d k is the dimension of the key vector, used to scale the dot product result. Through the above steps, the self-attention mechanism aggregates the information of the input sequence, generates a new context representation for each word, enables the model to efficiently process sequence data, and captures the dependencies between global features of the entire graph.
[0171] Multi-level feature fusion is achieved through the Feature Pyramid Network. Specifically, for the feature map with 32x downsampling, through upsampling and lateral connections, high-level semantic features are gradually fused into low-level features. The fused feature map contains both high-level semantic information and low-level spatial detail information, providing a richer feature representation in the detection task. To control the video memory occupancy, the self-attention mechanism is only implemented on the feature map with 32x downsampling to enhance the global perception ability of pixel points. And through the Feature Pyramid Network to fuse multi-scale features, finally, a feature map with 4x downsampling is output for subsequent prediction and matching of key points.
[0172] S2. Use the heatmap regression technique to detect key points, and combine with spatial perception feature point matching to achieve precise positioning and connection of wire endpoints.
[0173] Key point detection based on heatmap regression is a technique for locating key points (such as joint points in human postures, facial feature points, etc.) in images. By generating a heatmap to represent the probability distribution of each key point, and determining the coordinates of the key point according to the position of the heatmap. The position of the true key point is represented by a Gaussian distribution H k (x, y), whose center is the true coordinate of the key point as (x k , y k ), k represents the key point index, and σ controls the diffusion degree of the Gaussian distribution. In this embodiment, only the L2 loss is used to supervise the training: H k (x, y) = exp(-((x - x k ) 2 +(y - y k ) 2 ) / 2σ 2 ); During model training, due to the wire being too slender, a larger σ needs to be set at the beginning of training to alleviate the optimization difficulty, and σ is gradually reduced in the later stage of model training to control the prediction diffusion degree. In the prediction stage, the wire endpoint position is determined by the maximum value within the local area.
[0174] The key point detection part can only obtain the wire endpoint positions, but cannot know which points form a wire. To complete the wire instance detection, a spatial perception feature point matching algorithm is constructed. At the model prediction end, a key point feature description feature map is predicted, and the corresponding position feature point description prediction x is extracted through the true coordinates of the key points k ∈R n , where R is the set of ordered arrays composed of n real numbers; then the similarity matrix Sim ∈ R m*m is calculated, where m is the number of feature points.
[0175] To strengthen the spatial position perception between points during matching, the rotation position encoding and self-attention mechanism are also used to enhance the perception of features in space and among each other. During training, the negative log-likelihood loss function is used to supervise the training of the matching matrix, and the loss function is expressed as: sim = selfAtten(x)·selfAtten(x) T ; loss match =(1 / m)∑ k=0 k=m -log(softmax(softmax(simk, dim = 0), dim = 1)), where k represents the kth pair of matching points, and there are m pairs of matching point pairs in total.
[0176] S3. Introduce an auxiliary perception branch based on wire mask prediction to optimize wire area recognition and enhance the coherence and integrity of detection results.
[0177] Due to the slender and extremely narrow characteristics of the wire, a wire segmentation mask prediction branch is added during training to enhance the model's perception ability of the overall wire. During training, cross-entropy loss is used to optimize the wire mask prediction, and the loss function is as follows: loss mask =(1 / HW) ∑ i=0 i=H-1 ∑ J=0 J=W-1 -log(sigmoid(x ij )); where x ij represents the mask prediction value at the i-th row and j-th column. Finally, the optimized loss function is expressed as: loss = λ1·loss heatmap +λ2·loss match +λ3·loss mask ; where λ1, λ2, and λ3 are used to adjust the loss ratio and are set to 0.1, 10, and 1 respectively, and loss heatmap is the loss function for supervising heatmap regression.
[0178] In this embodiment, by processing the large-resolution picture in blocks, the analysis difficulty of complex wire distributions is effectively reduced, making wire detection more accurate. In each smaller picture block, the wire can be regarded as a straight line, simplifying the detection process and improving the accuracy of wire endpoint detection; using the HRNet high-resolution feature extraction module, combined with the attention mechanism and multi-resolution feature fusion, can capture the fine-grained features and global context information of the wire, enhancing the model's adaptability to different scenarios and wire morphologies; by predicting the key point regression heatmap, wire feature description map, and wire position mask map, combined with the wire endpoint matching strategy enhanced by the self-attention mechanism, accurate extraction of wire instances is achieved, and good detection results can be maintained even in complex situations where wires are dense or cross; the picture block strategy reduces the size of the image processed by the model at one time, reduces the demand for computing resources, making the detection process more efficient. At the same time, through the merging of wire endpoints between picture blocks, the local wire instantiation results can be seamlessly extended to the entire picture, ensuring the integrity and coherence of wire detection within the entire picture range.
[0179] According to one aspect of the present application, a wire instance detection system based on key point matching includes:
[0180] At least one processor; and,
[0181] A memory communicatively connected to at least one of the processors; where
[0182] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the wire instance detection method based on key point matching described in any one of the above embodiments.
[0183] The present invention constructs a complete wire detection framework. Through the organic combination of four main links, namely feature extraction, key point matching, wire instance detection, and global topology reconstruction, the accurate recognition and reconstruction of complex wire systems in aerial images are realized. In the feature extraction stage, through the four-dimensional to three-dimensional feature transformation and enhancement, the multi-scale features of the wire are effectively captured; in the key point matching stage, through recursive decomposition and double attention mechanism, the accurate positioning of the key points is realized; in the wire instance detection stage, through multi-layer heat maps and multi-dimensional constraint matching, the continuity and integrity of the wire instances are ensured; in the global topology reconstruction stage, through boundary analysis and cross-block matching, the accurate reconstruction of the large-scale wire system is realized. The present invention has strong adaptability and robustness, can effectively handle interference factors such as complex backgrounds, illumination changes, and perspective changes in aerial images, and improves the accuracy and reliability of wire detection. Especially when processing large-scale aerial images, through the strategies of block processing and cross-block matching, both the calculation efficiency and the global consistency of the detection results are ensured.
[0184] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A wire instance detection method based on key point matching, characterized in that It includes the following steps: S1. Obtain the original aerial image data, generate a four-dimensional feature matrix through dynamic feature mapping and enhancement processing; perform projection transformation on the four-dimensional feature matrix to obtain three-dimensional projection feature data; Based on the three-dimensional projection feature data, perform feature enhancement and multi-scale feature fusion to generate fused feature data; Based on the fused feature data, generate enhanced feature map data; S2. Based on the enhanced feature map data, obtain feature matrix data through feature recombination; Based on the feature matrix data, perform recursive decomposition and mapping to generate response matrix data; Based on the response matrix data, calculate adaptive threshold data and screen key points; Perform feature enhancement on the screened key points, and output a feature description set and a key point set; S3. Based on the key point set and the feature description set, generate heat map data through a multi-layer heat map; based on the heat map data and the spatial positions of the key points, construct a spatial association matrix and multi-dimensional constraints, perform endpoint matching, and obtain matching score data; Based on the matching score data, construct wire instances and verify them, and finally output a wire instance set; S4. Based on the wire instance set, perform boundary feature analysis to obtain a boundary feature matrix; Based on the boundary feature matrix, perform cross-block instance matching to generate cross-block matching data; Based on the cross-block matching data, perform topological optimization to generate topological credibility data; Based on the topological credibility data, perform instance fusion and quality assessment to generate a confidence map and a global wire map; Step S1 is further as follows: S11. Obtain the original aerial image data, construct a four-dimensional feature matrix including a spatial coordinate dimension, a time dimension, and an image channel dimension; calculate the first standard deviation and the first mean of each slice in the four-dimensional feature matrix, and generate an adaptive weight coefficient based on the ratio relationship between the first standard deviation and the first mean; perform weighted combination of the four-dimensional feature matrix and the adaptive weight coefficient, and output three-dimensional projection feature data; S12. Based on the three-dimensional projection feature data, calculate the feature variance value and the gradient amplitude; Based on the feature variance value, determine the enhancement coefficient and the balance coefficient; based on the median of the gradient amplitude, determine the adjustment coefficient; perform linear combination of the three-dimensional projection feature data and its second-order gradient, and perform weighting using the enhancement coefficient, the balance coefficient, and the adjustment coefficient, and output enhanced feature data; S13. Divide the enhanced feature data into different scales to generate a multi-scale feature sequence; Calculate the correlation matrix between the features of each scale in the multi-scale feature sequence to obtain feature correlation degree data; Based on the feature correlation degree data, construct a weight matrix; Based on the weight matrix, perform weighted fusion on the multi-scale feature sequence, and output fused feature data; S14. Based on the fused feature data, calculate its maximum value, minimum value, and standard deviation; Based on the maximum value and the minimum value, perform normalization processing on the fused feature data to obtain normalized feature data; Based on the standard deviation, calculate the dynamic scaling factor; Multiply the normalized feature data by the dynamic scaling factor, and finally output the enhanced feature map data.
2. The method for detecting wire instances based on key point matching according to claim 1, wherein Step S2 is further as follows: S21. Recombine the enhanced feature map data into feature matrix data including spatial position and feature dimension; Calculate the eigenvectors of the characteristic matrix data to obtain the left eigenvector matrix and the right eigenvector matrix; Based on the left eigenvector matrix and the right eigenvector matrix, calculate the eigenvalue diagonal matrix, and process the residual term through a non-linear mapping function to output the decomposed feature data; S22. Based on the left eigenvector matrix and the right eigenvector matrix in the decomposed feature data, calculate the spatial attention mapping value and the feature attention mapping value; Perform matrix multiplication on the spatial attention mapping value and the feature attention mapping value to generate a key point response matrix; construct a learnable weight matrix to adjust the key point response matrix and output the response matrix data; S23. Based on the response matrix data, calculate the second mean, the second standard deviation, and the information entropy; Based on the information entropy, calculate the dynamic coefficient; Combine the second mean, the second standard deviation, and the dynamic coefficient to generate the adaptive threshold data; Based on the adaptive threshold data, screen the response matrix data and output the candidate point set data; S24. Based on the candidate point set data, extract the local neighborhood features for each candidate point; Calculate the gradient and the second derivative of the local neighborhood features to generate the adaptive convolution kernel data; obtain the image size and calculate the position encoding value; based on the position encoding value, the adaptive convolution kernel data, and the local neighborhood features, combine them to generate the feature description set; S25. Based on the candidate point set data and the feature description set, calculate the local saliency score and the spatial distribution score for each point; based on the variance value of the feature description set, calculate the balance factor; Based on the balance factor, perform a weighted combination of the local saliency score and the spatial distribution score to obtain the weighted combination score; based on the weighted combination score, finally screen the points through a dynamic screening preset threshold and output the key point set.
3. The method for detecting wire instances based on key point matching according to claim 2, characterized in that, Step S3 is further as follows: S31. Based on the key point set, calculate the feature entropy value and the local variance value for each key point; based on the feature entropy value and the local variance value, calculate the adaptive variance parameter and construct an improved Gaussian distribution model; based on the Gaussian distribution model, generate and fuse the local heat map for each key point and output the heat map data; S32. Based on the heat map data and the feature description set, calculate the cosine similarity of the feature descriptions to obtain the feature similarity data; Obtain and calculate the distance matrix based on the spatial position of the key points to generate the position-related data; Perform a weighted combination of the feature similarity data and the position-related data and output the spatial correlation matrix; S33. Based on the spatial correlation matrix, calculate the feature similarity, the geometric constraint value, and the linear consistency value between the candidate end point pairs; Based on the feature similarity, the geometric constraint value, and the linear consistency value, calculate the constraint mean data and generate the first adaptive weight; based on the first adaptive weight, perform a weighted combination of the feature similarity, the geometric constraint value, and the linear consistency value and output the matching score data; S34. Based on the matching score data, calculate the topological consistency score for each pair of matching end points; Based on the variance of the matching score data, calculate the dynamic balance factor; Based on the dynamic balance factor, perform a weighted combination of the matching score data and the topological consistency score to generate the wire credibility data; Based on the wire credibility data and the coordinates of the matching endpoints, construct a wire path, and use a path correction algorithm to generate wire path data; S35. Based on the wire path data and the wire credibility data, calculate the visualization consistency score, region support score, and topological uniqueness score for each wire instance; Determine the second adaptive weight based on the degree of change in the visualization consistency score, region support score, and topological uniqueness score; Based on the second adaptive weight, perform a weighted combination of the visualization consistency score, region support score, and topological uniqueness score to generate instance quality data; Based on the instance quality data, screen and optimize the wire instances, and output a set of wire instances.
4. The method for detecting wire instances based on key point matching according to claim 3, wherein, Step S4 is further as follows: S41. Obtain the boundary region data of adjacent picture blocks in the set of wire instances, extract boundary features, and generate a boundary feature matrix; based on the boundary feature matrix, calculate the inner product and norm of the boundary feature vectors of adjacent blocks to obtain feature vector data; Perform normalization processing on the feature vector data and output boundary alignment score data; S42. Based on the set of wire instances, extract the wire endpoint information at the boundary; Based on the boundary alignment score data and the wire endpoint information, calculate the included angle of the direction vectors between endpoint pairs to obtain direction consistency data; Based on the neighborhood region of the wire endpoints, calculate the spatial overlap degree and generate overlap degree data; Based on the wire endpoint information, extract the curvature features of the endpoints and calculate the difference value, and output curvature feature data; Based on the mean values of the direction consistency data, overlap degree data, and curvature feature data, calculate the adaptive weight and then combine them to generate cross-block matching data; S43. Based on the cross-block matching data, construct an initial graph structure including all endpoints and connection relationships; Calculate the structure support degree, path smoothness, and global consistency value of each edge in the initial graph structure to obtain edge attribute data; based on the variances of the attributes in the edge attribute data, calculate the dynamic weight coefficient; Based on the dynamic weight coefficient, perform a weighted combination of the edge attribute data and output topological credibility data; S44. Based on the topological credibility data and the set of wire instances, calculate the output values of a predetermined number of fusion feature evaluation functions to obtain evaluation results; based on the evaluation results, generate fusion parameter data; Based on the fusion parameter data, apply a family of fusion operators to each pair of wire instances for feature fusion and output fusion instance data; S45. Based on the fusion instance data, calculate the output values of a predetermined number of quality evaluation functions to generate quality index data; based on the quality index data, calculate the information entropy of each quality index to obtain the first weight coefficient data; perform a weighted combination of the quality index data and the first weight coefficient data to generate a confidence map; based on the confidence map, optimize the fusion instance data and output a global wire graph.
5. The method for detecting wire instances based on key point matching according to claim 4, characterized in that, Step S11 is further as follows: S111. Obtain and based on the original aerial image data, extract spatial coordinate information to generate a position matrix; extract time series information to obtain time series data; separate image channels to obtain channel feature data; Combine the position matrix, time series data, and channel feature data to construct a four-dimensional feature matrix; S112. Slice the four-dimensional feature matrix. Obtain spatial slice data by slicing along the spatial dimension, temporal slice data by slicing along the temporal dimension, and channel slice data by slicing along the channel dimension; Calculate the local statistics of the spatial slice data, temporal slice data, and channel slice data, and output the slice statistical data; S113. Based on the slice statistical data, calculate the first standard deviation of each slice to obtain a standard deviation sequence; Based on the slice statistical data, calculate the first mean of each slice to obtain a mean sequence; Divide the standard deviation sequence by the mean sequence to generate a ratio sequence; perform normalization processing on the ratio sequence and output the adaptive weight coefficient; S114. Based on the four-dimensional feature matrix and the adaptive weight coefficient, calculate the projections of the feature matrix in each dimension to obtain dimension projection data; perform weighted combination on the dimension projection data and the adaptive weight coefficient to generate weighted projection data; perform dimensionality reduction transformation on the weighted projection data and output three-dimensional projection feature data.
6. The method for detecting wire instances based on key point matching according to claim 4, characterized in that, Step S24 is further as follows: S241. Based on the candidate point set data, extract multi-scale neighborhood blocks centered on each candidate point to generate a neighborhood block sequence; perform direction normalization processing on the neighborhood block sequence to obtain normalized neighborhood data; Calculate the gray distribution characteristics of the normalized neighborhood data and output the neighborhood feature data; S242. Based on the neighborhood feature data, calculate the first-order gradients in the horizontal and vertical directions to obtain a gradient vector field; based on the gradient vector field, calculate the Laplacian operator response to generate a second-order derivative map; perform weighted combination on the gradient vector field and the second-order derivative map according to the intensity distribution and output the gradient feature data; S243. Based on the gradient feature data, calculate the local response intensity and the local gradient direction histogram to obtain the direction distribution data; based on the direction distribution data, construct an adaptive Gaussian filter to generate filter coefficients; modulate the filter coefficients with the local response intensity and output the adaptive convolution kernel data; S244. Obtain the image size parameters, generate coordinate grids in the horizontal and vertical directions to obtain position coordinate data; perform sine and cosine transformations on the position coordinate data to obtain periodic coding data; based on the image size parameters, extract the global scale information; Based on the global scale information, normalize the periodic coding data and output the position coding value; S245. Based on the adaptive convolution kernel data, filter the neighborhood feature data to obtain filtered feature data; splice the filtered feature data and the position coding value in the channel dimension to generate combined feature data; perform dimensionality adjustment and normalization processing on the combined feature data and output the feature description set.
7. The method for detecting wire instances based on key point matching according to claim 4, characterized in that, Step S33 is further as follows: S331. Based on the spatial correlation matrix, extract the feature vectors of the candidate endpoint pairs, calculate the cosine distance of the feature vectors to obtain a similarity matrix; based on the similarity matrix, construct a feature distribution histogram to generate distribution feature data; perform normalization processing on the distribution feature data and output the feature similarity data; S332. Based on the spatial correlation matrix, obtain the spatial coordinates of the candidate endpoint pairs, calculate the Euclidean distance between the endpoints to obtain a distance matrix; Based on the distance matrix, statistically analyze the distance distribution characteristics within the local area to generate distance statistics; Compare the distance statistics with a preset threshold and output distance constraint data; S333. Based on the spatial correlation matrix, obtain the position information of the endpoint pairs, calculate the connection direction vectors, and obtain the direction matrix; Based on the gradient magnitude in step S12, extract the main direction features of the image to generate main direction data; Calculate the angle between the direction matrix and the main direction data and output direction constraint data; S334. Based on the distance constraint data and the direction constraint data, extract the structural features of the local area to obtain structural feature data; Based on the structural feature data, calculate the structural consistency score to generate structural constraint data; Combine the distance constraint data, the direction constraint data, and the structural constraint data and output the geometric constraint value; S335. Based on the spatial correlation matrix, obtain the connection information of the endpoint pairs, calculate the curvature features of the connections to obtain curvature data; Based on the connection information, extract the linear measurement features of the connections to generate linear feature data; Based on the curvature data and the linear feature data, calculate the linear consistency score and output the linear constraint value; S336. Based on the feature similarity data, the geometric constraint value, and the linear constraint value, calculate the mean value of each constraint to obtain constraint mean data; Based on the constraint mean data, generate the second weight coefficient data; Perform weighted combination of the feature similarity data, the geometric constraint value, the linear constraint value, and the second weight coefficient data and output the matching score data.
8. The method for detecting wire instances based on key point matching according to claim 4, characterized in that Step S43 is further as follows: S431. Based on the cross-block matching data, extract the spatial coordinates of the endpoints to generate a set of node coordinates; Based on the cross-block matching data and the set of node coordinates, extract the connection relationships between the endpoints, construct an adjacency matrix, and obtain adjacency relationship data; Combine the set of node coordinates and the adjacency relationship data to construct an initial graph structure and output initial graph data; S432. Based on the initial graph data, calculate the degree distribution characteristics of each node to obtain degree distribution data; Based on the initial graph data and the degree distribution data, statistically analyze the local connection patterns to generate connection pattern data; Based on the degree distribution data and the connection pattern data, calculate the structural importance of each node and output node weight data; S433. Based on the initial graph data and the node weight data, calculate the weights of the node pairs connected by each edge to obtain the initial edge weight values; Based on the initial graph data and the node weight data, extract the direction vector features of the edges to generate edge direction data; Combine the initial edge weight values and the edge direction data and output the basic edge attribute data; S434. Based on the basic edge attribute data, calculate the local support degree for each edge to obtain support degree data; Based on the basic edge attribute data, statistically analyze the local distribution characteristics of the edges to generate distribution feature data; Combine the support degree data and the distribution feature data and output the structural support degree value; S435. Based on the basic edge attribute data, calculate the curvature change of the paths to obtain curvature change data; Based on the basic edge attribute data and the curvature change data, extract the turning point features of the paths to generate turning point feature data; Based on the curvature change data and the turning point feature data, evaluate the path smoothness and output the smoothness value; S436. Based on the structure support value and the smoothness value, construct a global consistency evaluation matrix to obtain a consistency matrix; Calculate the eigenvalue distribution of the consistency matrix to generate eigenvalue data; Based on the consistency matrix and the eigenvalue data, calculate the global consistency score and output the consistency value; S437. Based on the structure support value, the smoothness value and the consistency value, calculate the variance of each attribute to obtain variance statistical data; based on the variance statistical data, generate the third weight coefficient data; Perform a weighted combination of the structure support value, the smoothness value, the consistency value and the third weight coefficient data, and output the topology credibility data.
9. A wire instance detection system based on key point matching, characterized in that, It includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the wire instance detection method based on key point matching according to any one of claims 1 to 8.
Citation Information
Patent Citations
Aerial image target detection method, equipment and medium
CN114140683A
Complex road target detection method based on multi-modal fusion aerial view
CN117058646A