Visible light-infrared modal target detection method based on deep learning

Through the visible-infrared modal object detection method based on deep learning, common key features are extracted and combined with the YOLO object detector, the problem of night infrared image processing and analysis is solved, and the accuracy and generalization ability of object detection are improved.

CN120070868AActive Publication Date: 2025-05-30UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510225641.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process and analyze night infrared images, resulting in a lack of accurate detection and recognition capabilities when facing an enemy plane for the first time.

Method used

The visible-infrared modal object detection method based on deep learning is adopted to extract common key features through convolutional neural networks and corner point and neighborhood attention modules, and combine with the YOLO object detector to realize object detection of images of different modal modes.

Benefits of technology

The detection effect of the target under different modes is improved, the generalization ability of the model is enhanced, the key feature areas of the target can be accurately found, and good detection ability is maintained after the loss of a specific mode.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070868A_ABST
    Figure CN120070868A_ABST
Patent Text Reader

Abstract

The invention provides a visible light-infrared modal target detection method based on deep learning, and belongs to the field of infrared detection. The method comprises the following steps: obtaining a visible light image or an infrared image needing to be detected; inputting the visible light image or the infrared image into a convolutional neural network to obtain an original visible light image feature or an original infrared image feature; inputting the original visible light image features or the original infrared image features into a corner point attention module to extract visible light corner point features or infrared corner point features, and inputting the original visible light image features or the original infrared image features into a neighborhood attention module to extract visible light edge features or infrared edge features; and inputting the extracted features into a trained YOLO target detector, outputting a target type, confidence and a bounding box position by the YOLO target detector, and outputting a final target detection result. The method can effectively improve the capability of finding the universal features of the target, can enable the model to learn more robust feature representation, and improves the generalization capability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of infrared detection, and particularly relates to a visible light-infrared modality target detection method based on deep learning. Background Art

[0002] An infrared image refers to an image formed by using an infrared detector to receive electromagnetic waves in the infrared band radiated by a target. It can reflect information such as the temperature, shape, and material of the target, and has the advantages of being observable day and night, not being affected by lighting conditions, and being able to penetrate smoke. It is widely used in military, industrial, medical, environmental monitoring and other fields. In the current military application environment, the visible light data and document materials of enemy aircraft are relatively rich. However, effective infrared image data can only be obtained when the enemy aircraft invades. In this case, we urgently need to be able to accurately detect the opponent and make up for the huge gap in this field by deeply studying night infrared image processing technology. Since infrared images can reflect information such as the temperature, shape, and material of the target, when facing the first encounter with an enemy aircraft, effectively processing and analyzing night infrared images will provide key advantages for military operations and intelligence collection. Therefore, it is of great significance to develop more accurate and reliable night infrared image processing and analysis technologies, which can not only help us better understand and utilize this valuable intelligence resource, but also make positive contributions to military defense and national security.

[0003] Deep neural networks based on visible light vision have achieved good performance in the field of target recognition and detection. Different from visible light images based on reflected light imaging, infrared images based on thermal radiation imaging lose a large number of visual representation features, such as image feature information such as color, texture, contour, and edge, resulting in difficulty for deep neural networks to better adapt to visual detection tasks based on infrared images.

[0004] The feature distribution of images has a crucial impact on infrared visual representation learning. Transfer learning theory shows that even if we can well mine the reproduction structure of a single image, when the feature distributions of images are inconsistent, it often leads to biases in the feature learning process, which in turn seriously affects the performance of target detection and recognition. Specifically, there are non-linear changes in pixel intensities between infrared images and visible light images, resulting in significant differences in the statistical characteristics of pixel intensities, gradient values, and gradient directions. Therefore, constructing a data feature distribution mapping relationship between visible light images and infrared images is crucial for extracting stable image feature representations. Summary of the Invention

[0005] The object of the present invention is to provide a visible light-infrared modality object detection method based on deep learning. Based on the common feature extraction algorithm between different modalities of visible light and infrared images, it can effectively improve the detection effect of objects in different modalities. To solve the technical problem of the lack of infrared modality samples in the existing object detection technology.

[0006] To solve the above technical problems, the specific technical solution of the present invention is as follows:

[0007] A visible light-infrared modality object detection method based on deep learning, the method includes the following steps:

[0008] Step S11: Obtain the visible light image or infrared image to be detected;

[0009] Step S12: Input the visible light image or infrared image into the convolutional neural network, and obtain the original visible light image features or original infrared image features through convolutional neural network feature extraction;

[0010] Step S13: Input the original visible light image features into the corner attention module to extract visible light corner features, and input the original visible light image features into the neighborhood attention module to extract visible light edge features; or input the original infrared image features into the corner attention module to extract infrared corner features, and input the original infrared image into the neighborhood attention module to extract infrared edge features;

[0011] Step S14: Input the extracted visible light corner features, visible light edge features or infrared corner features, infrared edge features into the trained YOLO object detector. The YOLO object detector outputs the object type, confidence level and bounding box position, and outputs the final object detection result.

[0012] Further, the YOLO object detector in step S14 is trained in the following manner:

[0013] Step S1: Obtain paired visible light images and infrared images;

[0014] Step S2: Input the visible light image and infrared image into the convolutional neural network respectively, and obtain the original visible light image features and original infrared image features through convolutional neural network feature extraction respectively;

[0015] Step S3: Input the original visible light image features into the corner attention module to extract visible light corner features, and input the original visible light image features into the neighborhood attention module to extract visible light edge features; input the original infrared image features into the corner attention module to extract infrared corner features, and input the original infrared image into the neighborhood attention module to extract infrared edge features; the extracted visible light corner features and visible light edge features form visible light image features; the extracted infrared corner features and infrared edge features form infrared image features;

[0016] Step S4: Process the visible light image features and the infrared image features to obtain representative features;

[0017] Step S5: Input the obtained representative features into the YOLO object detector;

[0018] Step S6: Optimize the model training.

[0019] Further, step S4 includes the following steps:

[0020] Step S41: Predict the center point of the target feature based on the visible light image features and the infrared image features;

[0021] Step S42: Map the visible light image features and the infrared image features to the feature distribution space and screen the core feature set;

[0022] Step S43: Screen the representative features through Laplacian feature constraints.

[0023] Further, step S41 includes the following steps:

[0024] Step S411: Construct corresponding feature sets according to the visible light image features and the infrared image features;

[0025] Step S412: Calculate the difference degree between the visible light image features and the infrared image features;

[0026] Step S413: Predict the center point of the target feature according to the difference degree between the visible light image features and the infrared image features.

[0027] Further, step S42 includes the following steps:

[0028] Step S421: Map the visible light image features and the infrared image features to the feature distribution space through linear features;

[0029] Step S422: Calculate the Mahalanobis distance between the visible light image features and the infrared image features and screen the core feature set;

[0030] Further, step S43 includes the following steps:

[0031] Step S431: Calculate the feature similarity between the core visible light image features and the core infrared image features in the core feature set;

[0032] Step S432: Construct a Laplacian matrix;

[0033] Step S433: Screen the representative features according to the Laplacian matrix.

[0034] Further, step S6 includes the following steps:

[0035] Step S61: Optimize the feature differences and commonalities using the comprehensive loss function;

[0036] Step S62: Update the weights through backpropagation.

[0037] Compared with the prior art, the present invention has the following beneficial technical effects:

[0038] 1) The present invention can effectively improve the ability to find target general features, and at the same time enable the model to learn more robust feature representations.

[0039] 2) Visible light images have rich visual representation features such as colors and textures, while infrared images lack such features. At the same time, the non-linear changes in pixel intensities of the two make the statistical characteristics of pixel intensities, gradient values, and gradient directions significantly different. This results in difficulty for deep neural networks to adapt to infrared image detection tasks. In view of these differences, the present invention uses corner and neighborhood attention modules to extract common key features, inputs them into the model for a series of operations, improves the ability to find general features, enhances the generalization of the model, and improves the target detection effect in different modalities. The present invention proposes, based on the feature differences between visible light images and infrared images, by focusing on the features that are common between the two modalities, to mine and learn visual features with strong expressive ability, improve the target detection task, and strengthen the generalization ability of the model.

[0040] 3) The present invention uses the feature commonalities of visible light and infrared images to enable the model to accurately find the key feature regions of the target.

[0041] 4) The present invention can be trained end-to-end and can be trained and inferred efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1 Schematic diagram of the visible light-infrared modality target detection framework based on deep learning of the present invention.

[0044] Figure 2 Schematic diagram of the visualization effect of the key features of the model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] A visible light-infrared modal object detection method based on deep learning proposed by the present invention first trains a model to obtain a trained YOLO object detector, as Figure 1 shown, and the training method includes the following steps:

[0047] Step S1: Obtain paired visible light images and infrared images.

[0048] Specifically, the paired visible light images and infrared images can be obtained from an open-source dataset; or paired simulated infrared images can be generated from the visible light images and used as training data.

[0049] Step S2: Input the visible light image and the infrared image into a convolutional neural network respectively, and obtain the original visible light image features and the original infrared image features through convolutional neural network feature extraction.

[0050] Step S3: Input the original visible light image features into a corner attention module to extract visible light corner features, and input the original visible light image features into a neighborhood attention module to extract visible light edge features; input the original infrared image features into a corner attention module to extract infrared corner features, and input the original infrared image into a neighborhood attention module to extract infrared edge features. The extracted visible light corner features and visible light edge features form visible light image features; the extracted infrared corner features and infrared edge features form infrared image features.

[0051] Specifically, the corner attention module applies Harris corner detection to find significant corners in the image, enhances the model's attention to these corners, and highlights the corner features through weighting. The neighborhood attention module uses the corner position information to extract dense neighborhood features around each corner to enhance the understanding of the object edges.

[0052] Step S4: Process the visible light image features and the infrared image features to obtain representative features.

[0053] Although the characteristics of visible light images and infrared images are different, the paired visible light images and infrared images capture the same target scene. The targets in the images have a corresponding relationship in terms of spatial position, and the visible light images and infrared images are correlated. For example, in a surveillance scenario, the positions of the same object in the visible light image and the infrared image are consistent, which provides a basis for target detection and recognition based on the two types of images.

[0054] The detailed features of visible light images and the temperature features of infrared images are complementary. The texture and shape information of visible light images helps to identify the object category; infrared images can detect the presence and approximate location of the target through obstacles such as smoke at night or in low-light environments. By combining visible light images and infrared images, more comprehensive target information can be obtained, improving the accuracy and reliability of target detection and recognition.

[0055] The present invention utilizes the correlation and complementarity of visible light images and infrared images to extract common features and improve the detection effect. Through the features of visible light images and infrared images, the accuracy of the center point can be improved. Key features can be extracted from both the visible light modality of visible light images and the infrared modality of infrared images, enabling the target detection model to effectively extract the common features of the target. Even when a specific modality is missing, the target detection model still has good detection capabilities.

[0056] Furthermore, step S4 includes the following steps:

[0057] Step S41: Predict the target feature center point through the features of visible light images and infrared images.

[0058] Step S411: Construct corresponding feature sets according to the features of visible light images and infrared images.

[0059] The visible light image feature set is where n vis represents the total number of feature vectors of visible light image features; R represents the set of real numbers, d represents the feature vector dimension, represents the i-th visible light image feature. For dimension m ∈ [1, d], the i-th visible light image feature in the m-th dimension is represented as

[0060] The infrared image feature set is where n ir represents the total number of feature vectors of infrared image features; represents the j-th infrared image feature. For dimension m ∈ [1, d], the j-th infrared image feature in the m-th dimension is represented as

[0061] Step S412: Calculate the degree of difference between the visible light image features and the infrared image features.

[0062] To measure the degree of difference between the visible light image features and the infrared image features, for any (i.e., the k-th feature vector in the visible light image feature set) and (the l-th feature vector in the infrared image feature set), the Euclidean distance is used for measurement, and its Euclidean distance calculation formula is:

[0063]

[0064] where, represents the k-th visible light image feature, represents the l-th infrared image feature, represents the Euclidean distance between the k-th visible light image feature and the l-th infrared image feature, represents the k-th visible light image feature in the m-th dimension, represents the l-th infrared image feature in the m-th dimension.

[0065] The Euclidean distance can intuitively reflect the spatial distance relationship between different modality image features.

[0066] Step S413: Predict the target feature center point according to the degree of difference between the visible light image features and the infrared image features.

[0067] Using the Euclidean distance calculated in step S412, form a distance matrix D with the Euclidean distances of all feature pairs, and then use the K-Means clustering algorithm to predict the feature center point of the target. The specific steps are as follows:

[0068] Initialize the cluster centers: Set the desired number of clusters K, and randomly initialize K cluster centers μ 1 ,…,μ i ,…,μ K ∈R d , where i ∈ [1, K]. These cluster centers are also d-dimensional vectors, representing the initial center positions of different classes.

[0069] Assign feature points to clusters: For each feature point f (which can be or ), calculate its Euclidean distance d Euclidean (f, μ i ) to each cluster center, and assign f to the class where the nearest cluster center is located. For example, if the distance from f to the cluster center μ i is the smallest, then f is classified into C i cluster category.

[0070] Update the cluster center: For each cluster category C i , its cluster center μ i is updated by calculating the average value of all feature points in this cluster, that is where |C i | represents the number of feature points included in the cluster category C i .

[0071] Iterate until convergence: Repeat the above two steps of assigning feature points and updating the cluster center until the cluster center no longer changes significantly. The finally determined cluster center, which is the predicted target feature center point, generally summarizes the core position of the feature distribution in space

[0072] Step S42: Map the visible light image features and infrared image features to the feature distribution space and screen the core feature set.

[0073] Step S421: Map the visible light image features and infrared image features to the feature distribution space through linear features.

[0074] In order to further explore the relationship between the visible light image features and infrared image features in different dimensional spaces, the original features are mapped to a new feature distribution space by means of linear mapping.

[0075] The linear mapping matrix is W ∈ R d′×d , where d′ represents the dimension of the mapped feature. Each element of this matrix determines the weight distribution of each dimension of the original feature vector in the mapping process, thus realizing the transformation from the d-dimensional space to the d′-dimensional space.

[0076] The feature of the i-th visible light image after linear mapping is represented as The feature of the j-th infrared image after linear mapping is represented as and are in the new d′-dimensional feature distribution space.

[0077] Step S422: Calculate the Mahalanobis distance between the visible light image features and infrared image features and screen the core feature set.

[0078] In the new feature distribution space, it is necessary to measure the distance relationship between features. The present invention introduces the Mahalanobis distance to comprehensively consider the feature distribution and the correlation between dimensions.

[0079] First, calculate the covariance matrix Σ of the mapped features, and its calculation formula is based on the mapped feature set

[0080] The calculation method of the covariance matrix is as follows:

[0081]

[0082] Among them, f′ represents any feature after mapping; is the mean vector of the features after mapping, with a dimension of d′, obtained by averaging all the feature vectors after mapping, reflecting the overall concentration trend of the features in the new feature distribution space. The covariance matrix Σ is a d′×d′ matrix, used to characterize the changes of features in different dimensions and the degree of mutual correlation between dimensions.

[0083] For any pair of features of visible light image and infrared image after mapping and The Mahalanobis distance calculation formula between them is:

[0084]

[0085] Among them, represents the k-th visible light image feature after mapping, represents the l-th infrared image feature after mapping; represents the Mahalanobis distance between the k-th visible light image feature and the l-th infrared image feature after mapping; T represents transpose, and Σ -1 represents the inverse matrix of the covariance matrix.

[0086] The Mahalanobis distance synthesizes the feature distribution information contained in the covariance matrix, and can more accurately reflect the true difference of features considering the dimension correlation compared with the Euclidean distance.

[0087] Set a distance threshold θ. By comparing the Mahalanobis distance between all pairs of features of visible light image and infrared image with this threshold, select the pairs of visible light image features and infrared image features whose Mahalanobis distance is less than θ, and form the core feature set F core . There are n core core visible light image features and core infrared image features in the core feature set, which are expressed as:

[0088]

[0089] The core feature set selected in this way contains core visible light image features and core infrared image features with relatively close distances and relatively tight feature distribution relationships in the new feature distribution space.

[0090] Step S43: Screen representative features through Laplacian feature constraints.

[0091] Step S431: Calculate the feature similarity between the core visible light image features and the core infrared image features in the core feature set.

[0092] To further analyze the internal relationship between the core visible light image features and the core infrared image features, it is necessary to calculate the similarity between the core visible light image features and the core infrared image features. The present invention uses a Gaussian kernel function to measure any two core visible light image features and the core infrared image features The feature similarity W pq between them is calculated as follows:

[0093]

[0094] where represents the Euclidean distance between the core visible light image feature and the core infrared image feature which is obtained through the Euclidean distance calculation method introduced before and reflects the degree of difference in the spatial position of the features. σ represents the bandwidth parameter of the Gaussian kernel, and σ determines the shape of the Gaussian kernel function and the sensitivity to distance. A smaller σ value makes the similarity change more violently when the feature distance is small, meaning it is more sensitive to small changes in distance; while a larger σ value makes the similarity change relatively smoothly with distance, affecting the quantization result of the similarity between features and the subsequent relationship structure constructed based on the similarity.

[0095] Step S432: Construct the Laplacian matrix.

[0096] Based on the calculated feature similarity between all core visible light image features and core infrared image features, construct an adjacency matrix The dimension of the adjacency matrix is n core ×n core which depicts the strength of the connection relationship between features.

[0097] Construct the degree matrix D. The degree matrix is a diagonal matrix, and its diagonal element represents the degree of the corresponding node (i.e., each feature point), that is, the sum of the weights of the edges connected to this feature point, and its dimension is also n core ×n core .

[0098] The Laplacian matrix L is obtained by subtracting the degree matrix D from the adjacency matrix W. The Laplacian matrix calculation formula is:

[0099] L = D - W

[0100] The dimension of the Laplacian matrix L is n core ×n core, which plays a core role in the feature screening process based on graph theory, can measure the local changes between features and reflect certain properties of the graph structure (a graph constructed with features as nodes), such as the local smoothness of features, etc., providing a basis for screening out the most representative core features.

[0101] Step S433: Screen representative features according to the Laplacian matrix.

[0102] Solve the eigenvalues and corresponding eigenvectors of the Laplacian matrix L. These eigenvalues are arranged in ascending order, and each eigenvalue corresponds to a specific eigenvector, which reflects different properties of the graph structure and the distribution characteristics of features in space. By selecting the eigenvectors corresponding to the first M smallest eigenvalues, the features corresponding to these eigenvectors are used as the most core M representative features.

[0103] Step S5: Input the obtained representative features into the YOLO object detector.

[0104] Step S6: Optimize the model training.

[0105] Step S61: Optimize the feature differences and commonalities using the comprehensive loss function.

[0106] The comprehensive loss function plays a key role in optimizing the relationship between features of two modalities (visible light and infrared). The formula of the comprehensive loss function is:

[0107]

[0108] where y ∈ {0, 1} is an indicator variable used to distinguish different sample categories or situations. y determines the emphasis of the two terms in the comprehensive loss function, guiding the model to optimize the feature distance in different situations. For example, when y = 1, it means that in the current sample situation, more attention is paid to making the features of the two modalities approach each other and reducing the distance between them; while when y = 0, in another situation, the distance between features is controlled within a reasonable range to avoid excessive differences.

[0109] is the distance metric between the features of the two modalities, which is the Mahalanobis distance of the mapped features, reflects the actual difference degree of the features of the two modalities in the feature space, is the key quantity in the comprehensive loss function for measuring the feature relationship, and its value directly affects the calculation result of the loss function and the subsequent optimization direction of feature commonalities and differences.

[0110] α is a preset threshold parameter, which is used in conjunction with and max means taking the maximum value. In the case of y = 0, when When it is greater than α, this term will punish the excess part to control the difference between the two modal features within a reasonable range, so as to optimize the feature difference, enabling the model to utilize the respective advantages of the two modal features while avoiding affecting the overall performance due to excessive differences.

[0111] During the model training stage, define the total loss function L of the model total , and the total loss function combines the comprehensive loss function L c and the loss function L of the YOLO detector itself yilo , which is used to measure the difference between the model prediction result and the true label. The goal of training is to minimize this loss function by continuously adjusting the model parameters.

[0112] To balance the relative importance of the comprehensive loss function L c and the loss function L of the YOLO detector yolo in the total loss function L total , introduce the weight coefficient β ∈ [0, 1]. By minimizing the total loss function, the formula for the total loss function is:

[0113] L total = βL c + (1 - β)L yolo

[0114] Continuously update the model parameters to enhance the commonality and difference of the two modal features, thereby improving the performance of the model in the object detection task, enabling it to more accurately detect the target object and locate its position and other information.

[0115] Step S62: Update the weights through backpropagation.

[0116] Let the parameters of the model be θ. The parameters of the model cover all variables involved in the model calculation, such as the convolution kernel parameters of the convolutional layer, the weights and biases of the fully connected layer, and their dimensions and specific structures depend on the network architecture design of the model.

[0117] After calculating the loss value L total (θ) through forward propagation, calculate the gradient of the loss function with respect to the parameters according to the backpropagation algorithm This gradient represents the direction in which the loss function rises fastest at the current parameter values. Therefore, when updating the parameters, adjust them along its opposite direction (i.e., the negative gradient direction) to gradually reduce the loss function value. The specific parameter update formula is:

[0118]

[0119] Among them, θ t+1 represents the parameter value of the model at the (t + 1)-th iteration, and θ t$\theta^{(t)}$ is the parameter value of the model at the $t$-th iteration, and $\eta$ is the learning rate. It is a positive number used to control the step size of parameter update. If the learning rate is too large, it may cause the parameter update to be too large, resulting in the model not converging or even increasing the loss function value. If the learning rate is too small, the parameter update is slow, and the training process will take a long time. Therefore, it is necessary to reasonably select the value of the learning rate according to the specific model and dataset. As the number of iterations $t$ increases, the model parameters are continuously updated and optimized, and the model performance will gradually be improved.

[0120] A visible light-infrared modal object detection method based on deep learning proposed by the present invention, based on the above-trained model, the method includes the following steps:

[0121] Step S11: Obtain a visible light image or an infrared image to be detected.

[0122] Step S12: Input the visible light image or the infrared image into a convolutional neural network, and obtain the original visible light image features or the original infrared image features through feature extraction by the convolutional neural network.

[0123] Step S13: Input the original visible light image features into the corner attention module to extract visible light corner features, and input the original visible light image features into the neighborhood attention module to extract visible light edge features; or input the original infrared image features into the corner attention module to extract infrared corner features, and input the original infrared image into the neighborhood attention module to extract infrared edge features.

[0124] Step S14: Input the extracted visible light corner features, visible light edge features or infrared corner features, infrared edge features into the trained YOLO object detector. The YOLO object detector outputs the object type, confidence level, and bounding box position, and outputs the final object detection result.

[0125] In this embodiment, 3 groups of control experiments are set up to verify the effectiveness of the present invention. The experimental setup table is as follows in Table 1:

[0126] Table 1 Experimental Setup Table

[0127]

[0128] Evaluation criteria:

[0129] The precision rate (P) represents the proportion of true target samples among the detected target samples, and is used to evaluate whether the infrared image target prediction is accurate. The precision rate P is calculated in the following way:

[0130]

[0131] where, TP represents true positive, and FP represents false positive.

[0132] The recall rate (R) represents the proportion of the detected true target samples in the total true target samples, reflecting whether all targets are discovered. The recall rate R is calculated as follows:

[0133]

[0134] where FN represents false negatives.

[0135] mAP50 refers to the average value of mAP (mean Average Precision) under 10 different IOU (Intersection over Union) thresholds with an IOU of 0.5 as the judgment criterion. IOU is an index to measure the degree of overlap between the target detection result and the true annotation, and the larger its value, the more accurate the detection result. An IOU of 0.5 is a commonly used judgment criterion, indicating that the area of the detection result region and the true annotation region has at least half of the overlapping area.

[0136] The experimental results are shown in Table 2:

[0137] Table 2 Experimental Results Table

[0138]

[0139] The experimental results show that through three groups of control experiments, compared with the original model, the present invention has significantly improved in terms of accuracy, recall rate and mAP50 indicators, fully verifying its advantages in feature extraction and target detection effect. It can be seen from the experimental data that:

[0140] In terms of accuracy: In the three groups of experiments, the accuracy of the present invention has been significantly improved compared with the original model. In the first group of experiments, the accuracy of the original model was 60.3%, and the present invention was improved to 73.5%; in the second group, it was improved from 65.4% to 77.1%; in the third group, it was improved from 79.4% to 89.4%. This shows that the present invention can more accurately identify the true target samples among the detected target samples, effectively reducing the misjudgment situation.

[0141] In terms of recall rate: In the first group of experiments, the recall rate of the original model was 37.3%, and the present invention reached 52.7%; in the second group, it was improved from 40.2% to 55.6%; in the third group, it was improved from 59.9% to 78.8%. This shows that the present invention can detect more true target samples, has more advantages in target discovery ability, and reduces the missed detection situation.

[0142] In terms of mAP50: mAP50 measures the mean average precision under the criterion of IOU being 0.5. The higher the value, the better the detection effect. In the first group of experiments, the mAP50 of the original model was 46.8%, and this invention improved it to 67.3%; in the second group, it was improved from 50.7% to 69.4%; in the third group, it was improved from 74.9% to 85.2%. This fully proves that this invention has a significant improvement in detection accuracy, and the coincidence degree between the detected area and the true annotation area is higher.

[0143] Compared with using the original model for training and detection, this invention has a great improvement in feature extraction and target detection effect.

[0144] Figure 2 It is a schematic diagram of the visualization effect of the key features of the model of this invention. From Figure 2 it can be seen that after using the feature extraction method optimized by the features in this article, the feature heat map of the model can clearly show that the optimized detection model can more accurately extract the effective features of the target and improve the detection ability of the model.

[0145] It can be understood that this invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of this invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of this invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of this invention. Therefore, this invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by this invention.

Claims

1. A visible light-infrared modality target detection method based on deep learning, characterized in that: The method comprises the following steps: Step S11: obtaining a visible light image or infrared image to be detected; Step S12: inputting the visible light image or the infrared image into a convolutional neural network, and extracting features of the original visible light image or the original infrared image through the convolutional neural network; Step S13: inputting the original visible light image features into the corner attention module to extract visible light corner features, and inputting the original visible light image features into the neighborhood attention module to extract visible light edge features; or inputting the original infrared image features into the corner attention module to extract infrared corner features, and inputting the original infrared image into the neighborhood attention module to extract infrared edge features; Step S14: The extracted visible light corner point features, visible light edge features or infrared corner point features, infrared edge features are input into the trained YOLO target detector, and the YOLO target detector outputs the target type, confidence level, and bounding box position, and outputs the final target detection result.

2. The visible light-infrared modality target detection method based on deep learning according to claim 1, characterized in that: The YOLO target detector in step S14 is trained as follows: Step S1: obtaining a paired visible light image and infrared image; Step S2: inputting the visible light image and the infrared image into the convolutional neural network respectively, and obtaining the original visible light image features and the original infrared image features respectively through the convolutional neural network feature extraction; Step S3: inputting the original visible light image features into the corner attention module to extract visible light corner features, and inputting the original visible light image features into the neighborhood attention module to extract visible light edge features; inputting the original infrared image features into the corner attention module to extract infrared corner features, and inputting the original infrared image into the neighborhood attention module to extract infrared edge features; the extracted visible light corner features and visible light edge features constitute visible light image features; the extracted infrared corner features and infrared edge features constitute infrared image features; Step S4: Processing the visible light image features and the infrared image features to obtain representative features; Step S5: input the obtained representative features into the YOLO target detector; Step S6: Optimize model training.

3. The visible light-infrared modality target detection method based on deep learning according to claim 2 is characterized in that: Step S4 includes the following steps: Step S41: predicting the target feature center point through visible light image features and infrared image features; Step S42: Mapping visible light image features and infrared image features to feature distribution space and screening a core feature set; Step S43: Filter representative features through Laplace feature constraints.

4. The visible light-infrared modality target detection method based on deep learning according to claim 3 is characterized in that: Step S41 includes the following steps: Step S411: constructing a corresponding feature set according to the visible light image features and the infrared image features; Step S412: Calculating the difference between the visible light image features and the infrared image features; Step S413: predicting the target feature center point according to the difference between the visible light image feature and the infrared image feature.

5. The visible light-infrared modality target detection method based on deep learning according to claim 4 is characterized in that: Step S42 includes the following steps: Step S421: mapping the visible light image features and the infrared image features to the feature distribution space through linear feature mapping; Step S422: Calculate the Mahalanobis distance between the visible light image features and the infrared image features and select the core feature set.

6. The visible light-infrared modality target detection method based on deep learning according to claim 5, characterized in that: Step S43 includes the following steps: Step S431: calculating the feature similarity between the core visible light image feature and the core infrared image feature in the core feature set; Step S432: constructing a Laplace matrix; Step S433: Screening representative features according to the Laplacian matrix.

7. The visible light-infrared modality target detection method based on deep learning according to claim 2, characterized in that: Step S6 includes the following steps: Step S61: optimizing feature differences and commonalities using a comprehensive loss function; Step S62: Back propagation to update weights.

Citation Information

Patent Citations

  • Infrared and visible light image variation fusion method capable of keeping saliency information

    CN109493309A

  • Infrared small target detection method and device based on image information entropy and multi-scale local contrast amount

    CN115731174A

  • Multispectral target detection model training method, target detection method and system

    CN117911710A

  • UAV (unmanned aerial vehicle) cross-modal fusion detection method based on CFT-OfficientDet

    CN118485932A