Deep learning feature matching method for low-light environment
Through the multi-layer self-attention and cross-attention mechanism of the LightGlue network, combined with the early stopping threshold and pruning threshold, the problem of insufficient robustness of traditional feature matching methods in low-light environments is solved, and efficient and accurate feature matching is achieved.
Patent Information
- Application Number
- CN202510787521.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-30
AI Technical Summary
Traditional feature matching methods are not robust enough in low-light environments, lack a self-evaluation mechanism, and are unable to dynamically adjust calculation strategies, resulting in poor image matching results in low-light environments, high computational complexity, and difficulty meeting real-time processing requirements.
The multi-layer self-attention and cross-attention mechanism of the LightGlue network is adopted, combined with the early stopping threshold and pruning threshold, to self-evaluate the confidence of feature points, dynamically adjust the calculation process, prune unreliable feature points, and improve matching accuracy and efficiency.
The accuracy and efficiency of feature matching are significantly improved in low-light environments, the amount of calculation is reduced, unnecessary calculation waste is avoided, and stable feature point detection and matching is ensured in low-light environments.
Smart Images

Figure CN120726355A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine vision technology, and in particular to a deep learning feature matching method, apparatus, medium and equipment for low-light environments. Background Art
[0002] Feature matching is a crucial technology in image processing, widely used in image stitching, 3D reconstruction, object recognition, and other applications. However, before the advent of LightGlue, traditional feature matching methods faced numerous challenges. First, early methods often relied on hand-crafted feature descriptors such as SIFT and SURF. While these features can capture key image information to a certain extent, they lack robustness in the face of complex situations such as illumination changes and perspective shifts. Second, with the increasing diversity of application scenarios, the demand for feature matching speed is also increasing. However, traditional methods, due to their high computational complexity and long processing time, cannot meet the needs of real-time processing. Furthermore, traditional methods lack self-evaluation mechanisms and cannot dynamically adjust their computational strategies based on actual conditions, resulting in wasted resources and low efficiency. Finally, traditional feature matching algorithms are often complex and difficult to integrate into existing software systems, which is undoubtedly a major obstacle for developers who want to quickly deploy solutions. Summary of the Invention
[0003] The main purpose of this application is to provide a deep learning feature matching method, device, medium and equipment for low-light environments, aiming to solve the technical problem that traditional methods lack a self-evaluation mechanism and cannot dynamically adjust the calculation strategy according to actual conditions.
[0004] To achieve the above-mentioned objectives, the present application provides a deep learning feature matching method for low-light environments, including: obtaining feature points of each image of a high-speed moving object in a power station / substation in a low-light environment; updating the state of the feature points of each image using the multi-layer self-attention and cross-attention mechanism of the LightGlue network; calculating the first confidence of each feature point after the state is updated based on an early stopping threshold classifier, and calculating the second confidence of each feature point after the state is updated based on a pruning threshold classifier; if the first confidence of at least one feature point among the feature points is less than a preset threshold, continuing to execute the multi-layer self-attention and cross-attention mechanism and pruning the feature points below the second confidence in each feature point after the state is updated, until the first confidence of each feature point remaining after pruning is greater than the preset threshold, then stopping the execution of the multi-layer self-attention and cross-attention mechanism to obtain each feature point remaining after pruning in each image; determining a point pair similarity matrix based on each feature point remaining after pruning in each image, determining a final allocation matrix based on the point pair similarity matrix, and determining the feature matching result of the high-speed moving object in the low-light environment based on the final allocation matrix.
[0005] Optionally, before updating the state of the feature points of each image using the multi-layer self-attention and cross-attention mechanism of the LightGlue network, the method further includes: normalizing the feature points of each image to obtain the coordinates of each feature point that are decoupled from the image resolution and the origin of the coordinate system is located at the exact center of the image.
[0006] Optionally, any two feature points of the feature points of each image are respectively the first feature point and the second feature point, and the multi-layer self-attention and cross-attention mechanism of the image includes a self-attention mechanism located at each layer; the self-attention mechanism includes: for each feature point in each image, using different linear transformation algorithms to process to obtain a query vector and a key-value vector; processing the query vector and key-value vector of the first feature point and the second feature point based on a preset self-attention calculation expression to obtain a self-attention score between the first feature point and the second feature point; and calculating the attention state between the first feature point and the second feature point based on the self-attention score calculation formula; performing weighted averaging on the attention state of the first feature point to obtain a message of the second feature point; updating the state of the second feature point based on the second feature point message to obtain the first state of any feature point.
[0007] Optionally, the feature points of any two pictures are respectively the third feature point and the fourth feature point; the multi-layer self-attention and cross-attention mechanism of the image includes a cross-attention mechanism located at each layer; the cross-attention mechanism includes: calculating the cross-attention score between each point of each image; calculating the attention state between the third feature point and the fourth feature point based on the cross-attention score between each point; performing weighted averaging on the attention state of the third feature point to obtain a message of the fourth feature point; updating the state of the fourth feature point based on the third feature point message to obtain the second state of any feature point in each picture.
[0008] In addition, to achieve the above-mentioned purpose, the present application also provides a deep learning feature matching device for low-light environments, including: a feature point acquisition module for acquiring feature points of each picture of a high-speed moving object in a power station / substation in a low-light environment; a confidence determination module for updating the state of the feature points of each picture using the multi-layer self-attention and cross-attention mechanism of the LightGlue network; a first confidence of each feature point after the state is updated is calculated based on an early stopping threshold classifier, and a second confidence of each feature point after the state is updated is calculated based on a pruning threshold classifier; a remaining feature point determination module for determining the remaining feature points if there are any remaining feature points in the feature points. If the first confidence of at least one feature point is less than a preset threshold, the multi-layer self-attention and cross-attention mechanisms continue to be executed and the feature points with a confidence lower than the second confidence among the feature points after the state update are pruned until the first confidence of each feature point remaining after the pruned is greater than the preset threshold, then the multi-layer self-attention and cross-attention mechanisms are stopped to obtain the feature points remaining after the pruned in each picture; a feature matching module is used to determine a point pair similarity matrix based on the feature points remaining after the pruned in each picture, determine a final allocation matrix based on the point pair similarity matrix, and determine the feature matching results of high-speed moving objects in a low-light environment based on the final allocation matrix.
[0009] To achieve the above objectives, the present application also provides a computer-readable storage medium, which includes instructions. When the instructions are run on a computer, the computer executes the deep learning feature matching method for low-light environments provided in the above embodiment.
[0010] To achieve the above-mentioned objectives, the present application also provides an electronic device, which includes: at least one processor, a memory and an input and output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the deep learning feature matching method for low-light environments provided in any of the aforementioned embodiments.
[0011] The embodiments of the present application propose a deep learning feature matching method, apparatus, medium and equipment for low-light environments, which obtain feature points of each picture of a high-speed moving object in a power station / substation in a low-light environment; update the state of the feature points of each picture using the multi-layer self-attention and cross-attention mechanism of the LightGlue network; calculate the first confidence of each feature point after the state is updated based on an early stopping threshold classifier, and calculate the second confidence of each feature point after the state is updated based on a pruning threshold classifier; if the first confidence of at least one feature point among the feature points is less than a preset threshold, continue to execute the multi-layer self-attention and cross-attention mechanism and prune the feature points below the second confidence among the feature points after the state is updated. Until the first confidence of each feature point remaining after pruning is greater than the preset threshold, the multi-layer self-attention and cross-attention mechanism is stopped to obtain the feature points remaining after pruning in each picture; the point pair similarity matrix is determined based on the feature points remaining after pruning in each picture, the final allocation matrix is determined based on the point pair similarity matrix, and the feature matching results of high-speed moving objects in low-light environments are determined based on the final allocation matrix. This application uses the LightGlue network, as well as a higher early stopping threshold and a higher pruning threshold to solve the problem that the texture and detail information of the image becomes blurred in low-light environments, making it difficult to extract stable and reliable feature points, resulting in premature exit of the system and poor detection robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A flowchart illustrating an embodiment of a deep learning feature matching method for low-light environments provided by this application;
[0013] Figure 2 A schematic diagram of an embodiment of a deep learning feature matching method for low-light environments provided by this application;
[0014] Figure 3 This is a structural block diagram provided for an embodiment of a deep learning feature matching device for low-light environments in this application.
[0015] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0016] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0017] Traditional feature matching methods include SIFT (Scale-Invariant Feature Transform) and SURF. While these traditional feature matching methods can capture key information in images to a certain extent, in low-light environments, the texture and detail information of the image becomes blurred. These traditional methods often have difficulty extracting stable and reliable feature points, resulting in poor matching results and poor robustness. Traditional methods lack a high confidence threshold and are unable to adjust the calculation strategy according to actual conditions. In low-light environments, image quality degrades, and traditional methods are often unable to effectively identify which feature points are reliable and which are unreliable, resulting in premature system exit due to the lack of a high confidence threshold.
[0018] In the existing technology, there is also a feature matching method based on SuperGlue. Problems with the feature matching method based on SuperGlue include low memory and computing efficiency, greater training difficulty, and, under certain conditions, lower matching accuracy than LightGlue. SuperGlue is not as efficient as LightGlue in terms of memory and computing efficiency. LightGlue has improved its design to make it more efficient in memory and computing, and can complete image matching tasks faster. SuperGlue performs relatively poorly in these aspects and usually requires a lot of computing resources. In particular, when large-scale image sets are processed in real time, its computational complexity is extremely high, resulting in slow processing speed, making it difficult to meet actual needs, large amount of computation and difficult to integrate.
[0019] In response to the problem that traditional feature matching is not robust enough in low-light environments, this application designs a deep learning feature matching method for low-light environments based on LightGlue. This method can solve the problem of low accuracy of traditional feature matching methods in low-light environments. It uses a higher early stopping threshold and a higher pruning threshold to implement a self-evaluation mechanism, which solves the problem that the computational complexity of large-scale image sets in low-light environments is extremely high when performing real-time processing, resulting in slow processing speed.
[0020] In an indoor environment with insufficient lighting, when a drone patrols a substation or power station, it only extracts a small number of feature points. The small and uneven distribution of feature points leads to problems such as large computational complexity and time consumption for feature matching, high integration difficulty and poor robustness, resulting in inaccurate positioning of the drone and a high risk of crashing.
[0021] In order to solve the above problems, the present invention proposes an improved LightGlue network, which is a graph neural network model designed for local feature matching. It introduces a higher early stopping threshold and a higher pruning threshold to maximize the accuracy of complex environments with low light levels and improve the matching accuracy. The most striking feature of LightGlue is that it proposes a major mechanism. Multi-layer self-attention and cross-attention mechanism: By stacking multiple layers of self-attention and cross-attention layers, LightGlue can capture the complex relationships between images, thereby improving matching accuracy. This mechanism enables LightGlue to more accurately match feature points between images. When LightGlue believes that all prediction results have reached a sufficiently high confidence level, it will choose to terminate the calculation process early to avoid unnecessary waste of calculations. At the same time, for those feature points that are judged to be unmatched, LightGlue will also decisively exclude them, further reducing unnecessary calculations, and has good accuracy in feature point detection and matching in low-light indoor environments, which can well solve the problem of poor feature extraction and matching effects and difficulty in integration in low-light indoor environments in this application. The overall framework is as follows Figure 1 As shown. Input two images A and B as a pair of local features (d, p), where p represents the position coordinates of the feature point and d represents the descriptor corresponding to the feature point. Visual representation is enhanced through position encoding and self- and cross-attention units. The introduced higher confidence classifier c helps decide whether to stop reasoning. If there are not many points that are credible, the reasoning process will continue to the next layer, but we will prune those points that are determined to be mismatched. Once a credible state is reached, LightGlue will predict a distribution based on the similarity and matching between point pairs.
[0022] Reference Figure 1 , Figure 1 The first embodiment of the present application provides a deep learning feature matching method for low-light environments. The method can be executed by a processor of any terminal or server. The deep learning feature matching method for low-light environments may include:
[0023] S10, obtaining feature points of each image of a high-speed moving object in a power station / substation under a low-light environment;
[0024] In one embodiment of the present application, before step S10, the deep learning feature matching method for low-light environments further includes:
[0025] The feature points of each image are normalized to obtain the coordinates of each feature point that are decoupled from the image resolution and the origin of the coordinate system is located at the exact center of the image.
[0026] Specifically, a dataset for training in low-light environments was constructed. Images in the dataset were normalized and preprocessed with data augmentation to increase the diversity of the training dataset and improve the model's generalization capabilities. An encoder was then used to extract features from the normalized data to obtain feature points.
[0027] When normalizing the extracted feature points, the processor first normalizes the x and y coordinates of the feature points. The purpose of normalization is to decouple the feature points from the image resolution. At the same time, when training the network structure proposed in this application, the values of [0, 1] are equivalent to the weight values, which will make the training more stable. The origin of the unnormalized image coordinate system is located in the upper left corner, that is, the pixel coordinate range is (0, 0) → (W, H), where W is the image width and H is the image height. The origin of the normalized coordinate system is located in the center of the image.
[0028] Traditional feature matching revolves around the position and visual descriptors of feature points. When actually matching, there are other aspects that can be referenced, such as the relative position relationship between feature points. For example, when we humans match two pictures with our naked eyes, it is like playing a "find the difference" game. We will look back and forth between the two photos, and after finding a prominent point (feature point) in picture A, we will look for the corresponding prominent point in picture B. Then, we will naturally find other prominent points around the prominent point in picture B, and then go to picture A to find out if there are any corresponding points. In this way, the possibility of false matching is checked back and forth. At the same time, humans will also look for some global information and additional related information to assist in judgment. For example, when matching, the relative relationship of feature points on the same object needs to be maintained. The LightGlue proposed in this application simulates humans to perform feature matching based on the attention mechanism of Transformer. This is explained in detail below.
[0029] S20, using the multi-layer self-attention and cross-attention mechanisms of the LightGlue network to update the state of the feature points of each image; calculating the first confidence of each feature point after the state is updated based on the early stopping threshold classifier, and calculating the second confidence of each feature point after the state is updated based on the pruning threshold classifier;
[0030] Wherein, any two feature points of the feature points of each image are respectively the first feature point and the second feature point, and the multi-layer self-attention and cross-attention mechanism of the image includes a self-attention mechanism located at each layer; the self-attention mechanism includes:
[0031] For each feature point in each image, different linear transformation algorithms are used to process it to obtain the query vector and key value vector;
[0032] Processing the query vector and the key-value vector of the first feature point and the second feature point based on a preset self-attention calculation expression to obtain a self-attention score between the first feature point and the second feature point;
[0033] and calculating the attention state between the first feature point and the second feature point based on the self-attention score calculation formula;
[0034] Perform weighted averaging on the attention state of the first feature point to obtain the message of the second feature point;
[0035] The state of the second feature point is updated based on the second feature point message to obtain the first state of any feature point.
[0036] In the actual execution process, the processor first calculates the current state x of the feature point i in the image. i Generate key vector k through different linear transformations i and query vector q i .
[0037] The self-attention score between feature point i and feature point j in the same image is defined as:
[0038]
[0039] in, To encode relative position rotation:
[0040]
[0041] Based on Scaled dot product attention, the attention is calculated as:
[0042]
[0043] Furthermore, by taking the weighted average of all states j, we get the message:
[0044]
[0045] It should be noted that the feature points of any two images are the third feature point and the fourth feature point respectively;
[0046] The multi-layer self-attention and criss-cross attention mechanisms of the image include criss-cross attention mechanisms at each layer;
[0047] The cross-attention mechanism includes:
[0048] Calculate the cross attention score between each point in each image;
[0049] Calculating the attention state between the third feature point and the fourth feature point based on the cross attention score between the points;
[0050] Perform weighted averaging on the attention state of the third feature point to obtain the message of the fourth feature point;
[0051] The state of the fourth feature point is updated based on the third feature point message to obtain the second state of any feature point in each picture.
[0052] After calculating the self-attention state of the two images, we use it as the input of the cross-attention unit to calculate the cross-attention state; each point in image I has a potential relationship with all points in image S. Here, we only calculate k for each point. i ,Correspondingly, the cross attention score is defined as:
[0053]
[0054] This way, the two images only need to be crossed once, eliminating the need to calculate I←J after the I→J calculation, thus saving computational power. Position encoding is not required between the two images, as relative position relationships are meaningless. Furthermore, state updates are performed similarly to self-attention, updating the states of the feature points in both images separately.
[0055] S30, if the first confidence of at least one feature point among the feature points is less than a preset threshold, continue to execute the multi-layer self-attention and cross-attention mechanisms and prune the feature points whose states are updated and whose confidence is lower than the second confidence, until the first confidence of each feature point remaining after pruning is greater than the preset threshold, then stop executing the multi-layer self-attention and cross-attention mechanisms, and obtain the feature points remaining after pruning in each image;
[0056] During the specific execution process, the processor can dynamically adjust the network depth according to the input image to reduce unnecessary calculations and reduce inference time. If the input image pair is easy to match, it means that the token confidence predicted by the previous network layer is very high and there is no difference with the subsequent network layer. At this time, the inference can be terminated early. After each layer of the network, LightGlue will calculate the first confidence c of each feature point. i .
[0057] c i =Sigmoid(MLP(x i ))∈[0,1]
[0058] The higher the Token confidence, the more reliable the representation of feature point i is, and the easier it is to be classified as matchable or unmatchable. On the contrary, the lower the Token confidence, the easier it is to be regarded as a pruning object. In this application, the Token confidence is set, that is, the first confidence, and the first confidence is a higher pruning threshold of 0.99. When the following conditions are met, that is, when the value of the second confidence exit is greater than the early stopping threshold of 0.96, the network layer reasoning will be terminated early. In this application, the second confidence is set to α, where the value of α is set to 0.96.
[0059]
[0060] in, Represents the first confidence of a feature point in image I. l It represents the confidence of network layer 1. The value of this execution degree is consistently the second confidence threshold in this application, that is, 0.96.
[0061] It should be noted that in this application, when the inference process does not exit early, some feature points will be pruned in advance to save computation. This is achieved by calculating the corresponding matching score for each feature point. The matching score represents the possibility that feature point i has a corresponding matching point, that is, the matching score σ i , for example, a point will not be detected in another image due to occlusion, then the corresponding σ i tends to 0:
[0062] σ i =Sigmoid(Linear(x i ))∈[0,1]
[0063] It should be noted that there is a difference between the second confidence level and the matching score. The processor will only match the points with the highest first confidence level. The matching quality is represented by the matching score. Therefore, a feature point is considered to be
[0064] unmatchable
[0065] Only when its first confidence score is higher than the threshold but the matching score is lower than the preset threshold.
[0066] S40: determining a point-pair similarity matrix based on the feature points remaining after pruning in each image, determining a final allocation matrix based on the point-pair similarity matrix, and determining a feature matching result of the high-speed moving object in the low-light environment based on the final allocation matrix.
[0067] In the specific execution process, the point pair similarity matrix is first calculated:
[0068]
[0069] Among them, Linear(*) is a linear transformation with bias.
[0070] The final allocation matrix P can be calculated as:
[0071]
[0072] For a pair of points, when both points are matchable and the similarity between them is higher than that between other points, then the pair of points is considered as related points. That is, when the element P of the assignment matrix ij If the value of is greater than the threshold and greater than other values in its column and row, then the pair of points is confirmed as matching points.
[0073] Experimental Conditions: To verify the pose estimation accuracy of our algorithm in low-light and fast-moving environments, we comprehensively evaluate the algorithm's performance from multiple perspectives, including feature point extraction, feature matching, estimated trajectory comparison, absolute trajectory error, and relative pose error. All comparative experiments were conducted on the same device and in the same environment. Absolute trajectory error was used as a key metric for evaluating algorithm performance, and its root mean square error (RMSE) was calculated to quantify this error for a more accurate assessment of the algorithm's performance. The hardware used in this experiment included an Intel Xeon Silver 4210 CPU, an RTX 6000 GPU with a 40GHz processing speed, 256GB of memory, and 2TB of hard drive storage. The system was Ubuntu 18.04.
[0074] Experimental process: By comparing the performance of feature extraction and matching of the traditional SIFT algorithm, the original Lightglue algorithm, and the improved Lightglue algorithm on the public dataset Euroc, the extraction and matching results are then integrated into the IMU for joint pose estimation evaluation.
[0075] The dataset used in the experiment is:
[0076] In this study, all data samples were derived from the EuRoC dataset, published by ETH Zurich. This dataset contains 11 synchronized and calibrated stereo vision sequences. Two representative sequences were selected: MH_04_difficult, a low-light sequence, and V2_03_difficult, a blurry sequence associated with fast motion. These sequences were used to verify the accuracy of feature extraction, matching, and pose estimation for this algorithm in complex low-light and fast-moving environments. The following table shows the comparison results of the experimental data. As can be seen from the table, the root mean square error (RMS) of this application is significantly lower than that of the control group. Therefore, this application's technical solution exhibits excellent performance in low-light and fast-moving scenarios.
[0077] Table 1 Root mean square error (m)
[0078]
[0079] Based on the above embodiments, the present application further provides a deep learning feature matching device for low-light environments. The deep learning feature matching device 100 for low-light environments includes:
[0080] The feature point acquisition module 1001 is used to acquire feature points of each image of a high-speed moving object in a power station / substation in a low-light environment;
[0081] The confidence determination module 1002 is used to update the state of the feature points of each image using the multi-layer self-attention and cross-attention mechanism of the LightGlue network; calculate the first confidence of each feature point after the state is updated based on the early stopping threshold classifier, and calculate the second confidence of each feature point after the state is updated based on the pruning threshold classifier;
[0082] The remaining feature point determination module 1003 is configured to continue executing the multi-layer self-attention and cross-attention mechanisms and prune the feature points whose states are updated and whose confidence levels are lower than the second confidence level if the first confidence level of at least one feature point among the feature points is lower than a preset threshold, until the first confidence levels of the remaining pruned feature points are all greater than the preset threshold, and then stop executing the multi-layer self-attention and cross-attention mechanisms to obtain the remaining pruned feature points in each image;
[0083] The feature matching module 1004 is used to determine a point pair similarity matrix based on the feature points remaining after trimming in each image, determine a final allocation matrix based on the point pair similarity matrix, and determine the feature matching results of high-speed moving objects in a low-light environment based on the final allocation matrix.
[0084] It is worth mentioning that all modules involved in this embodiment are logical modules. In actual applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovation of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application. However, this does not mean that other units do not exist in this embodiment.
[0085] It should be noted that the device embodiments of the present application correspond to the method embodiments and accordingly have the same technical effects as the method embodiments, which will not be repeated here.
[0086] Based on the above embodiments, the present application also provides a computer-readable storage medium, which includes instructions. When the instructions are run on a computer, the computer executes the deep learning feature matching method for low-light environments provided in the method claims.
[0087] Based on the above embodiments, the present application also provides an electronic device, which includes: at least one processor, a memory and an input and output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the deep learning feature matching method for low-light environments provided in the method claim.
[0088] The memory and processor are connected using a bus, which can include any number of interconnected buses and bridges. The bus connects various circuits of one or more processors and memories. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. These are all well known in the art and therefore will not be described further in this article. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. Data processed by the processor is transmitted on a wireless medium via an antenna. Furthermore, the antenna also receives data and transmits it to the processor.
[0089] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions, while the memory can be used to store data used by the processor when performing operations.
[0090] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0091] Each embodiment in this specification is described in a related manner. Similar portions between the embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are described briefly because they are generally similar to the method embodiments. For related portions, reference can be made to the description of the method embodiments.
[0092] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A deep learning feature matching method for low-light environments, characterized in that: include: Obtain feature points for each image of a high-speed moving object in a power station / substation under low-light conditions; Use the multi-layer self-attention and cross-attention mechanisms of the LightGlue network to update the status of the feature points of each image; Calculating a first confidence of each feature point after the state is updated based on the early stopping threshold classifier, and calculating a second confidence of each feature point after the state is updated based on the pruning threshold classifier; If the first confidence of at least one feature point among the feature points is less than the preset threshold, the multi-layer self-attention and cross-attention mechanisms are continued to be executed and the feature points below the second confidence level among the feature points after the state update are pruned until the first confidence levels of the remaining pruned feature points are all greater than the preset threshold. Then the multi-layer self-attention and cross-attention mechanisms are stopped to obtain the remaining pruned feature points in each image. A point-pair similarity matrix is determined based on the feature points remaining after pruning in each image, a final allocation matrix is determined based on the point-pair similarity matrix, and feature matching results of high-speed moving objects in low-light environments are determined based on the final allocation matrix.
2. The deep learning feature matching method for low-light environments according to claim 1, wherein: Before updating the state of the feature points of each image using the multi-layer self-attention and cross-attention mechanisms of the LightGlue network, the method further includes: The feature points of each image are normalized to obtain the coordinates of each feature point that are decoupled from the image resolution and the origin of the coordinate system is located at the exact center of the image.
3. The deep learning feature matching method for low-light environments according to claim 1, wherein: Any two feature points of the feature points of each image are respectively the first feature point and the second feature point, and the multi-layer self-attention and cross-attention mechanisms of the image include self-attention mechanisms at each layer; The self-attention mechanism includes: For each feature point in each image, different linear transformation algorithms are used to process it to obtain the query vector and key value vector; Processing the query vector and the key-value vector of the first feature point and the second feature point based on a preset self-attention calculation expression to obtain a self-attention score between the first feature point and the second feature point; and calculating the attention state between the first feature point and the second feature point based on the self-attention score calculation formula; Perform weighted averaging on the attention state of the first feature point to obtain the message of the second feature point; The state of the second feature point is updated based on the second feature point message to obtain the first state of any feature point.
4. The deep learning feature matching method for low-light environments according to claim 3, wherein: The feature points of any two images are the third feature point and the fourth feature point respectively; The multi-layer self-attention and criss-cross attention mechanisms of the image include criss-cross attention mechanisms at each layer; The cross-attention mechanism includes: Calculate the cross attention score between each point in each image; Calculating the attention state between the third feature point and the fourth feature point based on the cross attention score between the points; Perform weighted averaging on the attention state of the third feature point to obtain the message of the fourth feature point; The state of the fourth feature point is updated based on the third feature point message to obtain the second state of any feature point in each picture.
5. The deep learning feature matching method for low-light environments according to claim 4, wherein: The preset threshold of the first confidence level is 0.99, and the preset threshold of the second confidence level is 0.
96.
6. A deep learning feature matching device for low-light environments, characterized in that: include: Feature point acquisition module, used to obtain feature points of each image of high-speed moving objects in power stations / substations in low-light environments; The confidence determination module is used to update the state of the feature points of each image using the multi-layer self-attention and cross-attention mechanisms of the LightGlue network; Calculating a first confidence of each feature point after the state is updated based on the early stopping threshold classifier, and calculating a second confidence of each feature point after the state is updated based on the pruning threshold classifier; The remaining feature point determination module is configured to continue executing the multi-layer self-attention and cross-attention mechanisms and prune the feature points whose states are updated and whose confidence levels are lower than the second confidence level if the first confidence level of at least one feature point among the feature points is lower than a preset threshold, until the first confidence levels of the remaining pruned feature points are all greater than the preset threshold, and then stop executing the multi-layer self-attention and cross-attention mechanisms to obtain the remaining pruned feature points in each image; The feature matching module is used to determine the point-pair similarity matrix based on the feature points remaining after pruning in each image, determine the final allocation matrix based on the point-pair similarity matrix, and determine the feature matching results of high-speed moving objects in low-light environments based on the final allocation matrix.
7. A computer-readable storage medium, characterized in that It includes instructions, which, when run on a computer, enable the computer to execute the deep learning feature matching method for low-light environments as described in any one of claims 1 to 5.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor, memory, and input-output unit; The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the deep learning feature matching method for a low-light environment according to any one of claims 1 to 5.
Citation Information
Patent Citations
Visual SLAM method suitable for low-light dynamic environment
CN120031971A
Cited By
Unmanned aerial vehicle positioning method, device, equipment and medium
CN121962270A
A method, apparatus, device and medium for positioning a UAV
CN121962270B