Infrared target component identification method and system based on knowledge graph, and storage medium
By introducing a two-stage recognition strategy of the knowledge graph and component correlation attention module in infrared target component recognition technology, the problem of poor recognition performance in complex scenarios is solved, and high-precision target component recognition is achieved.
Patent Information
- Application Number
- CN202510178928.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-03
AI Technical Summary
The existing infrared target component recognition technology has poor recognition performance in complex scenarios, especially under conditions such as self-occlusion of targets, lack of obvious visual features, and large changes in features.
A two-stage recognition strategy based on knowledge graph is adopted, first identifying the target as a whole and expanding to high resolution, and then using the component recognition model to combine the target knowledge graph and component relevance attention module for component recognition.
It significantly improves the accuracy and recall rate of target components, solves the recognition problems caused by insufficient visual features, and the accuracy reaches 92.2%.
Smart Images

Figure CN120088458A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target recognition, and particularly relates to an infrared target component recognition method, system and storage medium based on a knowledge graph. Background Art
[0002] The technology of recognizing target components in infrared images plays an important role in the industrial field and military applications, and can help to more accurately understand and analyze the target structure, so as to achieve precise pose and position determination. Existing target component recognition methods are mainly divided into two categories: traditional algorithms and deep learning algorithms.
[0003] Traditional algorithms mainly rely on manually designed feature extraction means such as sliding windows and histograms of oriented gradients to recognize target components. However, such methods have significant limitations, such as the need for manual intervention to define features and weak adaptability to the features of different components, which makes it difficult for them to meet the complex and variable recognition requirements in practical applications.
[0004] Deep learning algorithms automatically learn more abstract and higher-dimensional features through models such as convolutional neural networks, showing strong feature expression capabilities and good generalization performance. Although deep learning has made remarkable progress in the field of target component recognition, existing algorithms still face many challenges. Specifically, existing algorithms often ignore the structural and semantic relationships between the components of the target, resulting in a decline in recognition ability when the target is self-occluded, the imaging is blurred, etc.
[0005] As a knowledge base represented by a graph structure, a knowledge graph can describe entities, attributes and the relationships between them, and has powerful knowledge reasoning capabilities. The knowledge graph has significant advantages in information integration, semantic understanding, relationship distribution, etc., and can deepen the connection from the whole to the part. Although the knowledge graph shows great potential in target recognition, due to reasons such as missing dataset information and recognition methods, the problem of target component recognition has not been effectively solved. Especially under conditions such as low-resolution small targets, target self-occlusion, unclear visual features, and large feature changes, the recognition performance of target components is poor.
[0006] In summary, the existing technologies have the following deficiencies:
[0007] 1. Traditional algorithms rely too much on manual feature extraction, resulting in limited adaptability;
[0008] 2. Deep learning algorithms ignore the structural and semantic connections between components and the decline in recognition performance in complex scenarios;
[0009] 3. The application potential of the knowledge graph in the field of target recognition has not been fully exploited.
[0010] Therefore, it is necessary to develop a new method, system and storage medium for infrared target component recognition based on knowledge graph. Summary of the Invention
[0011] The purpose of the present invention is to provide a method, system and storage medium for infrared target component recognition based on knowledge graph to improve the recognition performance in complex scenarios.
[0012] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0013] In the first aspect, a method for infrared target component recognition based on knowledge graph according to the present invention includes the following steps:
[0014] Obtain an infrared image;
[0015] Identify the overall target in the infrared image, and expand the target area to high resolution to obtain an expanded high-resolution target image;
[0016] Construct a component recognition model, and perform component recognition on the expanded high-resolution target image based on the component recognition model;
[0017] Among them, the component recognition model incorporates a target knowledge graph, and uses the structural relationship of target components to infer the co-occurrence relationship of components; at the same time, a component correlation attention module is added to the component recognition model. The component correlation attention module uses the target knowledge graph to infer the existence relationship of components, and incorporates the co-occurrence relationship of components into the visual features to enhance the component recognition ability by inferring the relationship between target components.
[0018] Optionally, identifying the overall target in the infrared image and expanding the target area to high resolution includes:
[0019] Identify the overall target in the infrared image and its position;
[0020] Separate the target from the infrared image;
[0021] Perform magnification and image enhancement processing on the separated target area;
[0022] Expand the enhanced target area to high resolution and perform enhancement processing to ensure sufficient resolution of target components at different distances, which is beneficial to improving component recognition performance.
[0023] Optionally, the enhancement processing includes sharpening filtering to enhance detail and edge information.
[0024] Optionally, the component recognition model incorporates a target knowledge graph, and uses the structural relationship of target components to infer the co-occurrence relationship of components, including:
[0025] Construct a target knowledge graph, where the target knowledge graph is used to describe the entities, attributes, and the relationships between them of the target components;
[0026] Use a global vector model based on word representation to map component information into word vectors;
[0027] Integrate the word vectors into a spatial attention module, associate component category information, and calculate the empirical vectors associated with the components and the co-occurrence matrix of the components;
[0028] Use the empirical vectors and the knowledge graph matrix to construct an empirical relationship graph and a fusion feature graph to improve the component recognition performance.
[0029] Optionally, the component recognition model includes a Backbone part, a Head part, and a Detect part;
[0030] The Backbone part includes a Conv module, a C2f module, and an SPPF module; the Conv module is composed of convolution, batch normalization, and an activation function spliced together; the C2f module is composed of two Conv modules and a Bottleneck structure spliced together; the SPPF module is a fast spatial pyramid pooling process;
[0031] The Head part includes a Conv module, a C2f module, and a component relevance attention module.
[0032] Optionally, when training and testing the component recognition model, it further includes a learning rate adjustment strategy based on self-occlusion, specifically including:
[0033] When making the training and testing data sets, set self-occlusion labels for the components to indicate whether the components are self-occluded;
[0034] Combine the cosine learning rate scheduling strategy with the self-occlusion labels to dynamically adjust the learning rate of each batch; when there are more occluded components, increase the learning rate to promote the component recognition model to learn the occluded features; when there are more non-occluded components, decrease the learning rate to reduce the impact of occlusion, thereby enhancing the learning performance and convergence of the component recognition model for occlusion.
[0035] Optionally, it further includes an evaluation step, using precision and recall to evaluate the component recognition accuracy index, and drawing a precision-recall curve, and calculating the mean average precision for all target components to evaluate the recognition ability of the component recognition model for all components.
[0036] Second aspect, an infrared target component recognition system based on a knowledge graph according to the present invention includes a memory and a controller. A computer-readable program is stored in the memory. When the computer-readable program is called by the controller, it can execute the steps of the infrared target component recognition method based on the knowledge graph according to the present invention.
[0037] Third aspect, a storage medium according to this aspect stores a computer-readable program. When the computer-readable program is called by the controller, it can execute the steps of the infrared target component recognition method based on the knowledge graph according to the present invention.
[0038] Advantages of the present invention:
[0039] 1. The present invention adopts a two-stage recognition strategy of first recognizing the target as a whole and then recognizing the target components. After recognizing the target as a whole, it expands to a high-resolution enhanced signal to detail the information, improves the target recognition ability, and associates and integrates the knowledge graph corresponding to the target category. It uses a global vector model to infer the co-occurrence relationship of the target component structure relationship, and fuses the component relevance attention to improve the component recognition performance, solving the component recognition problem caused by insufficient visual features; adding occlusion information of the target components to the label, regulating the learning rate during model training through the occlusion information, and enhancing the learning performance and convergence of the model to occlusion. The present invention effectively solves the problems such as target self-occlusion, unclear visual features, and large changes in features with distance, which lead to a decrease in the accuracy of component recognition. Through the indoor target equivalent scaled model system test verification, this method significantly improves the accuracy and recall rate of target component recognition, and the accuracy reaches 92.2%.
[0040] 2. The deep learning model in the present invention can automatically learn hierarchical feature representations from the original data without manual feature design, thus solving the problem that traditional algorithms are overly dependent on manual feature extraction, resulting in limited adaptability.
[0041] 3. Based on the Glove model, the present invention uses a dataset to create a component knowledge graph of the target, and uses the CAM model to apply it to the algorithm, thus giving full play to the role of the knowledge graph in the field of target recognition. Description of the Drawings
[0042] Figure 1 is a flowchart of the method in the embodiment of the present application;
[0043] Figure 2 is a flowchart of the overall and local recognition method in the embodiment of the present application;
[0044] Figure 3 is a structural diagram of the component recognition model in the embodiment of the present application;
[0045] Figure 4 It is the structure diagram of the CAM in the embodiment of the present application;
[0046] Figure 5 It is an example diagram of the aircraft attitude;
[0047] Figure 6 It is the inference flow chart of the target knowledge graph in the embodiment of the present application;
[0048] Figure 7 It is the label format diagram in the embodiment of the present application;
[0049] Figure 8 It is the example diagram of each attitude and distance of the target in the embodiment of the present application;
[0050] Figure 9 It is the performance result diagram of the comparison between the component recognition model algorithm and the original algorithm in the embodiment of the present application;
[0051] Figure 10 It is the mAP curve diagram of each component in the embodiment of the present application;
[0052] Figure 11 It is the comparison result diagram of aircraft component recognition at each attitude distance in the embodiment of the present application;
[0053] Figure 12 It is the schematic diagram of the system in the embodiment of the present application.
[0054] In the figure, 1 is the controller, and 2 is the memory. Specific implementation manner
[0055] The following will describe the implementation manners of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for explaining the present invention, rather than for limiting the protection scope of the present invention.
[0056] As Figure 1 shown, in the embodiment of the present application, a method for infrared target component recognition based on a knowledge graph includes the following steps:
[0057] Step 1. Obtain an infrared image.
[0058] Step 2. Identify the overall target in the infrared image, and expand the target area to high resolution to obtain an expanded high-resolution target image.
[0059] Step 3. Build a component recognition model, and perform component recognition on the expanded high-resolution target image based on the component recognition model. Among them, the component recognition model incorporates the target knowledge graph, and uses the target component structure relationship to infer the co-occurrence relationship of components; at the same time, a component relevance attention module is added to the component recognition model, and the component relevance attention module uses the target knowledge graph to infer the existence relationship of components, and integrates the co-occurrence relationship of components into the visual features, and enhances the component recognition ability by inferring the relationship between target components.
[0060] The following is a detailed description of the infrared target component recognition method based on the knowledge graph:
[0061] An indoor target equivalent scaled-down model verification system is built, and a large number of tests are carried out on the aircraft under different postures and distances. The results show that the proposed method has better ability to recognize target components, and significantly improves the accuracy and recall rate.
[0062] 1. Target component recognition based on the knowledge graph
[0063] Existing deep learning network models ignore the relationship information between target components, resulting in a decline in recognition ability when the target is self-occluded, the imaging is blurred, etc., a significant drop in the recognition recall rate, and poor environmental state adaptability. To address this problem, the embodiment of this application proposes an infrared target component recognition method based on the knowledge graph, as Figure 2 shown. First, identify the overall target in the infrared image and its position; separate the target from the infrared image; perform magnification and image enhancement processing on the separated target area; expand the enhanced target area to high resolution and perform enhancement processing. Among them, the enhancement processing includes sharpening filtering to enhance detail and edge information. The problem of blurring or unclear details in the magnified image is solved, thereby improving the target recognition ability. After that, perform component recognition on the high-resolution target image after target recognition. The component recognition model incorporates the target knowledge graph and uses the target component structure relationship to infer the co-occurrence relationship of components. Specifically, it includes: constructing a target knowledge graph, where the target knowledge graph is used to describe the entities, attributes, and their relationships of target components; using a global vector model based on word representation to map component information into word vectors; integrating the word vectors into a spatial attention module, associating component category information, and calculating the empirical vector associated with the component and the co-occurrence number graph matrix of components; constructing an empirical relationship graph and a fusion feature graph using the empirical vector and the knowledge graph matrix. By fusing component relevance attention, the component recognition performance is improved, and the problem of component recognition difficulties caused by insufficient visual features is solved. It should be noted that after detecting and recognizing the overall target, this method improves the target resolution to ensure sufficient resolution of target components at different distances, which is beneficial to improving the component recognition performance.
[0064] 1.1 Component Recognition Strategy Based on Knowledge Graph
[0065] As shown in Figure 3 , the infrared target component recognition algorithm structure based on the knowledge graph is shown. Its main structure includes the Backbone part, the Head part, and the Detect part. The input is an infrared image of 640*640 (i.e., the image after magnification and sharpening filtering). In the Backbone part, the Conv module, the C2f module, and the SPPF module are used. The Conv module is composed of convolution, batch normalization (Batch Normalization, BN), and the activation function (Sigmoid Linear Unit, SiLU) spliced together. The C2f module is composed of two Conv modules and the Bottleneck structure spliced together; SPPF is the fast spatial pyramid pooling process. Based on the Conv module and the C2f module, the Head part combines the co-occurrence relationship (not occluded at the same time) in the knowledge graph and adds the Component-related Attention Module (CAM) to enhance the component recognition ability by reasoning about the relationship between target components.
[0066] Construct co-occurrence knowledge for the component box distribution of the target, perform graph convolution operations on the target features using the knowledge graph, and let the target features propagate along the edges of the knowledge graph nodes to finally obtain the target enhanced features that integrate the co-occurrence relationship. Combining the knowledge graph and visual features, in the embodiments of this application, CAM is constructed to statistically analyze the component association degree on one side of the target, use the target knowledge graph to infer the existence relationship of components, and improve the component recognition accuracy, as shown in Figure 4 . To make up for the deficiencies of visual features such as component self-occlusion, unclear visual features, and large feature changes, CAM constructs a relationship graph using the word embedding vector of the category label and the relationship matrix of the target category to enhance the recognition of potential co-occurring components on the same side. CAM integrates the co-occurrence relationship into the visual features, improving the component recognition ability in cases of self-occlusion, unclear visual features, and feature changes. Using the co-occurrence relationship between components, this method uses the word embedding vector obtained by training the Glove model on the global word-word co-occurrence count where n represents the number of categories and r represents the dimension. The co-occurrence matrix represents the co-occurrence probability between component categories in different poses of the target.
[0067] As shown in Figure 5 , there is a co-occurrence relationship between the left wing, the left tail wing, and the left vertical wing of the aircraft, while the right tail wing does not have a co-occurrence due to being occluded.
[0068] As shown in Figure 6As shown in the figure, it is the inference flow chart of the target knowledge graph, and the specific implementation method is as follows.
[0069] First, construct a co-occurrence adjacency matrix. The directed co-occurrence relationship graph (this directed co-occurrence relationship graph is obtained after performing convolution operations on this co-occurrence adjacency matrix by the following graph convolutional network) uses the word embedding vector as the node, and uses the asymmetric co-occurrence matrix to initialize it. Given a labeled training dataset, first calculate the number of times each component appears in all images:
[0070] A P ={A 1 , A 2 , A 3 , ……, A n} (1)
[0071] In the formula, A i (i ∈ {1, 2, 3 ……, n}) represents the number of times component i appears in the entire dataset.
[0072] Then construct an initial matrix C to calculate the number of times any two components appear:
[0073]
[0074] In the formula, C ij represents the number of times component i and component j co-occur. Specifically, when component i and component j are both unoccluded in a certain image, C ij and C ji are both incremented by 1. To establish the directed co-occurrence relationship between different components, by calculating their conditional probabilities, construct an asymmetric co-occurrence matrix A:
[0075]
[0076] A ij = P(C i | C j ) = C ij / A j (4)
[0077] In the formula, A ij represents the probability that component i is also unoccluded when component j is unoccluded.
[0078] In the embodiments of the present application, a graph convolutional network (Graph Convolutional Network, GCN) is used to infer the co-occurrence relationship graph. The graph convolutional network GCN operates on the node features along the corresponding edges and applies the non-linear function LeakyReLU to update the node features. The inference formula is expressed as:
[0079] F (l+1) = L(A F (l) W (l) )(5)
[0080] In the formula, F (l+1) represents the output feature, L represents the non-linear function LeakyReLU, A represents the normalized matrix of the co-occurrence matrix A, F (l) represents the input feature, and W (l) represents the weight matrix.
[0081] To guide the recognizer to pay more attention to potential co-occurring targets in the same image, the feature results of the relationship reasoning are further combined with the visual feature F vision . First, F vision is passed through the max pooling layer, and then re-weighted by multiplying by F (l +1) . After the weighted features are successively convolved through a linear transformation and a sigmoid activation function layer, finally, the convolved feature map and the visual feature F vision are linked through a concatenation operation to output a new feature map F CAM :
[0082]
[0083] where F CAM is the output feature of the CAM layer, F vision is the output feature of the C2f module, represents the concatenation operation, S represents the sigmoid activation function, F l represents the linear transformation, P represents the max pooling operation, represents the weighted multiplication, and F (l+1) is the output feature of the GCN network. In short, CAM integrates the co-occurrence relationship and visual features, improves the attention of visual features, and enhances the recognition ability of recognition targets.
[0084] 1.2 Learning Rate Adjustment Strategy Based on Self-Occlusion
[0085] Existing deep learning network models adopt a learning rate decay strategy to prevent the model loss from oscillating back and forth during the training process. Common learning rate decays include the cosine learning rate scheduling strategy, the adaptive scheduling strategy, etc. To address the problems of inaccurate part features caused by target self-occlusion and learning rate oscillation, an occlusion-oriented learning rate adjustment strategy (Self-occlusion based Learning Rate Decay, SLD) is proposed in the embodiments of this application. When making a dataset (for training and testing a part recognition model), in addition to the part category and the coordinates of the ground truth box in the label, a 0 / 1 label is also set to indicate whether the part is self-occluded (0 means not occluded). The label format is as Figure 7 shown:
[0086] Combining the cosine learning rate scheduling strategy with the self-occlusion label, the learning rate adjustment strategy based on self-occlusion is generated and expressed as:
[0087]
[0088] δ b = F 0 / (F 1 + F 0 ) (8)
[0089] In the formula, lr(b) is the learning rate of the current batch, and δ b is the self-occlusion adjustment parameter. In each batch, count the number of all 0 / 1 labels to dynamically adjust the learning rate parameter of each batch. min_lr is the minimum learning rate to prevent the model from stopping updating, and max_lr is the initial learning rate; t and T represent the current layer number and the total layer number respectively; F 0 and F 1 represent the number of 0s and 1s in Flag respectively. When F 0 is larger, it means that there are more occluded parts in this batch. At this time, δ b will also increase, increasing the learning rate. On the contrary, it means that there are more occluded parts, and δ b will decrease, reducing the learning rate to reduce the impact of occlusion.
[0090] 2 Experimental data and discussion and analysis
[0091] 2.1 Experimental dataset
[0092] An indoor target equivalent scaled-down model verification system was built. The 1:100 scaled-down model of the F-35 aircraft was fixed on a turntable, and the turntable was used to control the movement of the aircraft in azimuth and pitch; the turntable moved back and forth on the guide rail to simulate the movement of the target at different distances. The near-infrared camera Goldeye CL-030TEC1 was used to acquire aircraft images with a resolution of 640×512 pixels, which were then transmitted to a computer for display and processing.
[0093] The test shooting covered various flight postures, and 9 components such as the aircraft nose, cockpit, left wing, right wing, left tail wing, right tail wing, left vertical fin, right vertical fin, and engine were labeled. Figure 8 It was to obtain typical target images. Compared with the infrared camera, in Figure (a), it was a long-distance and right yaw situation, with self-occlusion on the left side of the aircraft; in Figure (b), it was a long-distance and left yaw, with less self-occlusion on the right side; in Figure (c), it was a close-distance and forward flight, without self-occlusion; in Figure (d), it was a close-distance and left yaw, with complete self-occlusion on the right side; in Figure (e), it was a close-distance and right yaw situation, with less self-occlusion on the left side.
[0094] In the embodiment of this application, 4860 images were selected for the experiment, and they were randomly divided into a training set, a validation set, and a test set according to a ratio of 8:1:1. A deep learning network was constructed, trained, and tested using PyTorch, and the GPU model was Titan X; the training method used the SGD stochastic gradient descent method, with an initial learning rate of 0.01, a batch size set to batch_size = 4, and the number of iterations epoch = 200.
[0095] 2.2 Evaluation Metrics
[0096] The test used Precision and Recall to evaluate the component recognition accuracy metrics. Precision represents the proportion of actual positive samples among the predicted positive samples, while Recall represents the proportion of actual positive samples that are predicted as positive samples:
[0097]
[0098] In the formula, TP is the number of correctly predicted positive samples, FP is the number of samples wrongly predicted as positive samples, and FN is the number of samples wrongly predicted as negative samples.
[0099] By using Precision and Recall to plot the Precision-Recall curve (Average Precision, AP), calculate the AP value for all target components, and then take the average to obtain mAP to evaluate the recognition ability of the component detection method for all components.
[0100]
[0101] Where \(P(R)\) represents the precision-recall curve of this category, \(R\) represents the recall rate, and \(P\) represents the precision rate.
[0102]
[0103] Where \(N\) represents the total number of categories, \(AP\) i represents the \(AP\) value of category \(i\)
[0104] 3.3 Ablation Experiment
[0105] Using yolov8 as the basic algorithm model, a series of ablation experiments were carried out to verify the effectiveness of each module of the proposed algorithm for the component recognition ability under the condition of target self-occlusion. The test results are shown in Table 1. After adding the WCP and CAM modules, the precision rates of component recognition increased by 4% and 5.7% respectively, and the recall rates increased by 6.1% and 6.6% respectively, indicating that the method of fusing the WCP and CAM modules improved the generalization ability of target component recognition at different distances and poses. At the same time, after adding the three modules of WCP, CAM and SLD, the precision rate and mAP of target component recognition of the proposed method reached 93.5% and 92.2% respectively, and the recall rate reached 94% when adding the CAM module.
[0106] Table 1 Results of Ablation Experiment
[0107]
[0108]
[0109] During the training process, the change curves of the bounding box loss \(box\_loss\) and the mean average precision \(mAP\) of this method and the Yolov8 model with the number of iterations are as Figure 9 shown. The loss value of the Yolov8 model finally stabilized at 1.51, while the loss value of this method finally stabilized at 1.09, as Figure 9 (a) shows. The mean average precision curves finally stabilized at 89.8% and 91.1% respectively, as Figure 9 (b) shows. The SLD module adjusted the learning rate of the model during the training process and improved the convergence speed of the model.
[0110] As Figure 10 shown, the PR curves of the algorithm in the embodiment of this application and the Yolov8 model for the recognition of each component. It can be seen that the recognition accuracy \(AP\) of each component by this algorithm has reached 90%, and the average recognition rate has reached 92.2%. Not only is the recognition probability higher, but it is also more stable. It should be noted that the recognition accuracy of each component by the Yolov8 model varies greatly. The recognition rate of the aircraft nose is 0.944, while the recognition rate of the engine is only 59.8%, and the average recognition rate is 87.3%
[0111] 3.4 Comparative Experiment
[0112] To verify the recognition performance of this algorithm for components with self-occluding infrared targets, the Yolov8 model, the WCL-Yolov4-tiny algorithm
[20] and the Hybrid Task Cascade (HTC) algorithm
[21] were selected for comparative tests. The test results are shown in Table 2, and the AP values and mAP@0.5 of each component were used as the performance indicators for comparison. As can be seen from Table 2, the average recognition accuracy of this algorithm for components with target self-occlusion reaches 92.2%, which is 4.9%, 2.6% and 3.2% higher than that of the Yolov8, WCL-Yolov4-tiny (abbreviated as tiny in the table) and HTC methods respectively. From the relationship of the AP values of each component, it can be seen that although the Yolov8 and WCL-Yolov4-tiny algorithms perform well on some specific components, such as the left vertical wing, there are large differences in the recognition of components such as the engine, which is a problem caused by relying solely on visual features. The recognition accuracy of this algorithm for each component reaches more than 90%, and the mAP value is also the best among all algorithms.
[0113] Table 2 Comparison results of different algorithms Tab.2 Comparison results of different algorithms
[0114]
[0115]
[0116] Figure 11 are the recognition result diagrams of aircraft target components at different postures and distances. Among them, rows (1)-(5) represent various distances and postures of the target. (a) is the original image, (b) is the ground truth label, and (c), (d), (e) and (f) are the recognition results of the yolov8 algorithm, the WCL-Yolov4-tiny algorithm, the HTC algorithm and this algorithm respectively. For easy observation, the components were visually marked with colors such as dark red, green, yellow, blue, purple, cyan, gray, brown, red, etc. for the target components, and the recognition areas of the three comparison algorithms were enlarged proportionally.
[0117] For those with less occlusion and closer postures, such as the first and second rows, all four algorithms show good recognition performance, but there are a small number of cases where the Yolov8 and HTC repeatedly detect the aircraft head and cockpit.
[0118] In the third row, the target is relatively far away and the degree of self-occlusion of the component is small. The recognition accuracy of the three comparison algorithms for the aircraft tail drops significantly, and there are many cases of missed detection of the tail wing. Y olov8The WCL-Yolov4-tiny algorithm missed detecting the right tail wing, while the HTC algorithm missed detecting both tail wings. In comparison, the algorithm in this paper did not have any missed detection and was less affected by the distance.
[0119] In the fourth and fifth rows, the target was relatively close, but the degree of self-occlusion was relatively high. For the unoccluded aircraft head and cockpit, all four algorithms showed good performance. For the occlusion situation on the left side of the aircraft in the fourth row, the three comparison algorithms did not identify the components on the left side of the aircraft, such as the left wing and the left tail wing. For the occlusion situation on the right side of the aircraft in the fifth row, Yolov8 only identified the left wing and the vertical fin, WCL-Yolov4-tiny and HTC misidentified the left wing and the vertical fin, while the algorithm in this paper could better identify the components on both the left and right sides.
[0120] From the above data, it can be seen that the algorithm in this paper has better recognition performance for components with more self-occlusion, such as engines, left and right wings, etc. It has the highest recognition ability for components in different postures and distances and maintains a low false detection rate, making it suitable for identifying target components with self-occlusion.
[0121] 3. Conclusion
[0122] Aiming at problems such as self-occlusion of the target, unclear visual features, and large changes in features with distance in component recognition, an infrared target component recognition method based on a knowledge graph is proposed in the embodiments of this application. This method uses a holistic and local recognition method, associates and integrates the knowledge graph corresponding to the target, and combines a component correlation attention module with a learning rate adjustment strategy for self-occlusion to solve the problem of component recognition difficulties caused by insufficient visual features in the case of self-occlusion. Through the test and verification of the indoor target equivalent scaled model system, this method significantly improves the accuracy and recall rate of target component recognition, and the accuracy reaches 92.2%.
[0123] As Figure 12 shown, in the embodiments of this application, an infrared target component recognition system based on a knowledge graph includes a memory 1 and a controller 2. The memory 1 stores a computer-readable program, and when the computer-readable program is called by the controller 2, it can execute the steps of the infrared target component recognition method described in the embodiments of this application.
[0124] In the embodiments of this application, a storage medium stores a computer-readable program, and when the computer-readable program is called by the controller, it can execute the steps of the infrared target component recognition method described in the embodiments of this application.
[0125] This system combines the relationship network of the target knowledge graph and helps to solve the problem of component recognition under conditions such as target self-occlusion and blurred imaging through the relationship reasoning between components. This method adopts a two-stage recognition strategy of first recognizing the target as a whole and then recognizing the target components. That is, first detect the target as a whole, identify the target and its position, and expand the target area to a high-resolution enhanced signal to detail the information, which is used to associate the target knowledge graph and solve the problem of excessive size differences of target components at different distances. Then, use the global vector model based on word representation (GloVe) to map the component information into word vectors, integrate them into the spatial attention module, associate the component category information, and calculate the empirical vector associated with the component and the co-occurrence number graph matrix of the components; use the empirical vector and the knowledge graph matrix to construct an empirical relationship graph and a fusion feature graph to improve the component recognition performance. At the same time, fully consider the problem of component self-occlusion caused by the target pose, add self-occlusion / non-occlusion labels to represent the appearance state of the components in the dataset, and adaptively adjust the learning rate to reduce the impact of misleading features caused by occlusion on the model.
[0126] In the embodiments of the present application, the storage medium may be a tangible storage medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The storage medium may be a machine-readable signal storage medium or a machine-readable storage medium. The storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. More specific examples of the storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0127] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for identifying an infrared target component, characterized in that: The following steps are involved: Acquire infrared images; Recognize the entire target in the infrared image, and expand the target area to a high resolution to obtain an expanded high-resolution target image; A component recognition model is constructed, and components are recognized on the expanded high-resolution target image based on the component recognition model; Among them, the component recognition model is integrated into the target knowledge graph, and the target component structural relationship is used to infer the component co-occurrence relationship; at the same time, a component relevance attention module is added to the component recognition model, and the component relevance attention module uses the target knowledge graph to infer the component existence relationship and integrates the component co-occurrence relationship into the visual features, thereby enhancing the component recognition ability by inferring the relationship between target components.
2. The method according to claim 1, characterized in that The target in the infrared image is identified as a whole, and the target area is expanded to a high resolution, including: Identify the entire target and its location in the infrared image; separating the target from the infrared image; Enlarge and enhance the image of the separated target area; The enhanced target area is expanded to a high resolution and then enhanced.
3. The method according to claim 2, characterized in that The enhancement processing includes sharpening filtering to enhance details and edge information.
4. The method according to claim 1, characterized in that: The component recognition model is integrated into the target knowledge graph, and the component co-occurrence relationship is inferred using the target component structure relationship, including: Construct a target knowledge graph, where the target knowledge graph is used to describe entities, attributes and relationships between target components; Use a global vector model based on word representation to map component information into word vectors; Integrate word vectors into the spatial attention module, associate component category information, and calculate component-related experience vectors and component co-occurrence map matrices; The experience vector and knowledge graph matrix are used to construct the experience relationship graph and the fusion feature graph.
5. The method according to claim 1, characterized in that The component recognition model includes a Backbone part, a Head part and a Detect part; The Backbone part includes a Conv module, a C2f module and an SPPF module; the Conv module is composed of convolution, batch normalization and activation function; the C2f module is composed of two Conv modules and a Bottleneck structure; the SPPF module is a fast spatial pyramid pooling process; The Head part includes a Conv module, a C2f module and a component relevance attention module.
6. The method according to claim 1, characterized in that When training and testing the component recognition model, a learning rate adjustment strategy based on self-occlusion is also included, specifically including: When preparing training and testing data sets, set self-occlusion labels for components to indicate whether the components are self-occluded; The cosine learning rate scheduling strategy is combined with the self-occlusion label to dynamically adjust the learning rate of each batch. When there are many occluded parts, the learning rate is increased; when there are many unoccluded parts, the learning rate is reduced.
7. The method according to claim 1, characterized in that The method also includes an evaluation step, using precision and recall to evaluate the component recognition accuracy index, and drawing a precision-recall curve, and calculating the average precision mean for all target components to evaluate the recognition ability of the component recognition model for all components.
8. An infrared target component recognition system based on knowledge graph, characterized in that: It includes a memory and a controller, wherein the memory stores a computer-readable program, and when the computer-readable program is called by the controller, it can execute the steps of the infrared target component recognition method based on the knowledge graph as described in any one of claims 1 to 7.
9. A storage medium, characterized in that: A computer-readable program is stored therein, and when the computer-readable program is called by the controller, the steps of the infrared target component recognition method based on the knowledge graph as described in any one of claims 1 to 7 can be executed.