Contour-based fine-grained target classification and recognition methods, systems, and storage media

By employing a contour-based fine-grained classification and recognition method, which combines contour extraction and recognition sub-networks with 3D model projection image data, the problem of perspective differences in aircraft model recognition is solved, achieving high-precision aircraft model recognition.

CN116824201BActive Publication Date: 2025-12-02SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310173593.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-12-02
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately identify aircraft models from different perspectives, especially military aircraft, due to the large differences in angles and between different types of aircraft, as well as the lack of multi-view optical image data, which makes identification difficult.

Method used

By employing a contour-based fine-grained target classification and recognition method, the target object contour extraction sub-network is used to locate and segment the target object contour, and the target object contour recognition sub-network is combined for fine classification. Finally, by combining the 3D model projection image data, a contour-based fine-grained target classification and recognition network model is established.

Benefits of technology

It enables fine-grained classification of aircraft models in the air, improves recognition accuracy, and can accurately identify aircraft models from different perspectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824201B_ABST
    Figure CN116824201B_ABST
Patent Text Reader

Abstract

This invention discloses a contour-based fine-grained target classification and recognition method, system, and storage medium. The method includes: acquiring target object image data; inputting the target object image data into a preset contour-based fine-grained target classification and recognition network model for analysis, locating the target object and segmenting its contour through a target object contour extraction sub-network to obtain target object contour image data; performing fine-grained classification on the target object contour image data through a target object contour recognition sub-network to obtain classification data of the target object in the target object image data; and sending the classification data of the target object to a preset terminal for display. This invention achieves the purpose of fine-grained classification of target objects by locating them and identifying their types based on their contours.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of data processing and data transmission, and more specifically, to contour-based fine-grained target classification and recognition methods, systems, and storage media. Background Technology

[0002] Target object identification (aircraft type) plays a crucial role in air combat command systems, national defense and military early warning and defense systems, and air traffic control. With the development of optical imaging systems, optical instruments can capture images of target objects, and then identify aircraft in these images to detect flight paths and distinguish between friendly and enemy aircraft. This provides more detailed combat information, enabling rapid response and the implementation of measures to engage enemy targets, thus gaining a battlefield advantage. Aircraft type identification based on optical images has attracted attention from both military and academic communities.

[0003] Current target object image recognition technologies mainly include recognition algorithms based on traditional image processing and recognition algorithms based on deep learning, both of which have achieved good results. However, image-based aircraft model recognition technology still faces difficulties and challenges. Unlike the single top-down view of aircraft remote sensing images from the ground, aircraft in the air exhibit diverse attitude variations from different perspectives, with varying coating colors and textures, as well as significant intra-class differences and smaller inter-class differences, making accurate aircraft model identification difficult. Even for the same aircraft model, images show significant visual differences due to different shooting angles. Conversely, different aircraft models may have similar appearances from certain shooting angles. Furthermore, images of aircraft in flight are relatively scarce, especially for military aircraft, where the secrecy makes it difficult to obtain a large number of multi-view optical aircraft images from publicly available sources. In particular, with the rapid development of science and technology, many countries have developed new fighter jets. Obtaining real-scene images of these aircraft is extremely difficult; at most, 3D model information can be acquired. How to utilize 3D model information for aircraft model identification and prepare reconnaissance mechanisms for defense systems is a worthy area of ​​research.

[0004] Therefore, the existing technology has defects and urgently needs improvement. Summary of the Invention

[0005] In view of the above problems, the purpose of this invention is to provide a contour-based fine-grained classification and recognition method, system and storage medium for target objects, which locates target objects and identifies target object types by using the target object contour, thereby achieving the purpose of fine-grained classification of target objects.

[0006] The first aspect of this invention provides a contour-based fine-grained target classification and recognition method, comprising:

[0007] Acquire image data of the target object;

[0008] The target object image data is input into a preset contour-based fine-grained target classification and recognition network model for analysis. The target object is located and its contour is segmented through the target object contour extraction sub-network to obtain the target object contour image data.

[0009] The target object contour image data is refined by using a target object contour recognition sub-network to obtain the classification data of the target object in the target object image data;

[0010] The classification data of the target object is sent to a preset terminal for display.

[0011] This plan also includes:

[0012] Acquire sample image data;

[0013] The sample image data is preprocessed to obtain the target object image data;

[0014] The target object image data is input into a preset E2EC instance segmentation model to obtain the target object contour image data;

[0015] A target object contour image dataset is established based on the target object contour image data.

[0016] This plan also includes:

[0017] Obtain the 3D model data of the target object;

[0018] The three-dimensional model data of the target object is projected using a first preset method to obtain multiple projected image data of the target object model.

[0019] The target object model's multiple projected image data are preprocessed using a second preset method to obtain a three-dimensional model contour image dataset of the target object.

[0020] A preset contour-based fine-grained target classification and recognition network model is established based on the target object contour image dataset and the target object 3D model contour image dataset.

[0021] In this solution, the step of locating the target object and segmenting its contour through a target object contour extraction sub-network to obtain target object contour image data includes:

[0022] Feature extraction is performed on the target object image data to obtain image feature information of the image data, and a heat map is generated based on the image feature information;

[0023] Based on the analysis of the heat map, the center point of the target object within the target object image is obtained;

[0024] Based on the analysis of the heat map and the center point, the initial target object contour image data is obtained;

[0025] The initial target object contour image data is refined to obtain the final target object contour image data.

[0026] In this solution, the step of performing refined classification of the target object contour image data through a target object contour recognition sub-network to obtain classification data of the target object in the target object image data includes:

[0027] Feature extraction is performed on the target object contour image data to obtain the global features of the target object contour image and the feature map of the target object contour image;

[0028] The feature map of the target object contour image is cropped according to the third preset method to obtain multiple target object contour local region images with discriminative information.

[0029] Feature extraction is performed on the local region images of the target object contours with discriminative information to obtain the local features of the target object contour images and the local feature maps of the target object contour images.

[0030] The feature map and local feature map of the target object contour image are fused to obtain fine-grained features of different parts and scales of the target object contour image.

[0031] Aircraft model data is obtained by analyzing the fine-grained features of different parts and scales of the target object's contour image.

[0032] In this scheme, the step of cropping the feature map of the target object contour image according to the third preset method to obtain multiple target object contour local region images with discriminative information includes:

[0033] Based on the feature map of the target object contour image, extract the features with the maximum recognizability and minimum redundancy of the local region of the target object contour to obtain multiple feature maps;

[0034] The activation map K of the local region of the target object contour is obtained by calculating the multiple feature maps;

[0035] The average activation value of the local region activation map K of the target object contour is obtained by aggregation calculation.

[0036] According to the activation average The feature map of the target object contour image is cropped to obtain multiple target object contour local region images with discriminative information.

[0037] A second aspect of the present invention provides a contour-based fine-grained target classification and recognition system, comprising a memory and a processor, wherein the memory includes a contour-based fine-grained target classification and recognition method program, which, when executed by the processor, performs the following steps:

[0038] Acquire image data of the target object;

[0039] The target object image data is input into a preset contour-based fine-grained target classification and recognition network model for analysis. The target object is located and its contour is segmented through the target object contour extraction sub-network to obtain the target object contour image data.

[0040] The target object contour image data is refined by using a target object contour recognition sub-network to obtain the classification data of the target object in the target object image data;

[0041] The classification data of the target object is sent to a preset terminal for display.

[0042] This plan also includes:

[0043] Acquire sample image data;

[0044] The sample image data is preprocessed to obtain the target object image data;

[0045] The target object image data is input into a preset E2EC instance segmentation model to obtain the target object contour image data;

[0046] A target object contour image dataset is established based on the target object contour image data.

[0047] This plan also includes:

[0048] Obtain the 3D model data of the target object;

[0049] The three-dimensional model data of the target object is projected using a first preset method to obtain multiple projected image data of the target object model;

[0050] The target object model's multiple projected image data are preprocessed using a second preset method to obtain a three-dimensional model contour image dataset of the target object.

[0051] A preset contour-based fine-grained target classification and recognition network model is established based on the target object contour image dataset and the target object 3D model contour image dataset.

[0052] In this solution, the step of locating the target object and segmenting its contour through a target object contour extraction sub-network to obtain target object contour image data includes:

[0053] Feature extraction is performed on the target object image data to obtain image feature information of the image data, and a heat map is generated based on the image feature information;

[0054] Based on the analysis of the heat map, the center point of the target object within the target object image is obtained;

[0055] Based on the analysis of the heat map and the center point, the initial target object contour image data is obtained;

[0056] The initial target object contour image data is refined to obtain the final target object contour image data.

[0057] In this solution, the step of performing refined classification of the target object contour image data through a target object contour recognition sub-network to obtain classification data of the target object in the target object image data includes:

[0058] Feature extraction is performed on the target object contour image data to obtain the global features of the target object contour image and the feature map of the target object contour image;

[0059] The feature map of the target object contour image is cropped according to the third preset method to obtain multiple target object contour local region images with discriminative information.

[0060] Feature extraction is performed on the local region images of the target object contours with discriminative information to obtain the local features of the target object contour images and the local feature maps of the target object contour images.

[0061] The feature map and local feature map of the target object contour image are fused to obtain fine-grained features of different parts and scales of the target object contour image.

[0062] Aircraft model data is obtained by analyzing the fine-grained features of different parts and scales of the target object's contour image.

[0063] In this scheme, the step of cropping the feature map of the target object contour image according to the third preset method to obtain multiple target object contour local region images with discriminative information includes:

[0064] Based on the feature map of the target object contour image, extract the features with the maximum recognizability and minimum redundancy of the local region of the target object contour to obtain multiple feature maps;

[0065] The activation map K of the local region of the target object contour is obtained by calculating the multiple feature maps;

[0066] The average activation value of the local region activation map K of the target object contour is obtained by aggregation calculation.

[0067] According to the activation average The feature map of the target object contour image is cropped to obtain multiple target object contour local region images with discriminative information.

[0068] A third aspect of the present invention provides a computer-readable storage medium comprising a contour-based fine-grained target classification and recognition method program, wherein when the contour-based fine-grained target classification and recognition method program is executed by a processor, it implements the steps of the contour-based fine-grained target classification and recognition method as described in any of the preceding claims.

[0069] This invention discloses a contour-based fine-grained target classification and recognition method, system, and storage medium. The method includes: acquiring target object image data; inputting the target object image data into a preset contour-based fine-grained target classification and recognition network model for analysis, locating the target object and segmenting its contour through a target object contour extraction sub-network to obtain target object contour image data; performing fine-grained classification on the target object contour image data through a target object contour recognition sub-network to obtain classification data of the target object in the target object image data; and sending the classification data of the target object to a preset terminal for display. This invention achieves the purpose of fine-grained classification of target objects by locating them and identifying their types based on their contours. Attached Figure Description

[0070] Figure 1 A flowchart of the contour-based fine-grained target classification and recognition method of the present invention is shown;

[0071] Figure 2 A flowchart of a method for acquiring contour image data of a target object according to the present invention is shown;

[0072] Figure 3 A flowchart of a method for acquiring a local region image of the contour of a target object according to the present invention is shown;

[0073] Figure 4 A block diagram of the contour-based fine-grained target classification and recognition system of the present invention is shown;

[0074] Figure 5A block diagram of the contour-based target fine-grained classification and recognition deep neural network structure in this invention is shown.

[0075] Figure 6 The diagram shows the effect of the output results of three sub-modules of an E2EC according to the present invention;

[0076] Figure 7 A block diagram of a target object contour recognition sub-network structure according to the present invention is shown;

[0077] Figure 8 The diagram illustrates the process of creating a dataset of actual contour images of a target object according to the present invention.

[0078] Figure 9 This image shows a rendering of a simplified model of an aircraft configuration according to the present invention.

[0079] Figure 10 This diagram shows a partial projection image of a target object model according to the present invention.

[0080] Figure 11 The image shown is a rendering of a partial outline image of a target object model according to the present invention. Detailed Implementation

[0081] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0082] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0083] Figure 1 A flowchart of the contour-based fine-grained target classification and recognition method of the present invention is shown.

[0084] like Figure 1 As shown, this invention discloses a contour-based fine-grained target classification and recognition method, including:

[0085] S102, acquire image data of the target object;

[0086] S104, the target object image data is input into a preset contour-based fine-grained classification and recognition network model for analysis. The target object is located and the target object contour is segmented through the target object contour extraction sub-network to obtain the target object contour image data.

[0087] S106, The target object contour image data is refined by the target object contour recognition sub-network to obtain the classification data of the target object in the target object image data;

[0088] S108, the classification data of the target object is sent to a preset terminal for display.

[0089] According to an embodiment of the present invention, a preset contour-based fine-grained target classification and recognition network model is used to process the acquired target object image data, thereby obtaining the classification data of the target object in the target object image data. The preset contour-based fine-grained target classification and recognition network model comprises two sub-networks. The first sub-network is a target object contour extraction sub-network, which locates the target object from the optical image and segments the target object contour based on a contour-level instance segmentation network. The second sub-network is a target object contour recognition sub-network, which, based on the proposed Dual-Stream Target Object Contour Recognition Network (DACR-Net), performs fine-grained classification of the target object contour image, further recognizes the actual contour of the target object, and finally outputs the target object type corresponding to the target object contour image. The combination of the two can effectively complete the target object type recognition task based on optical images. This solution explains the contour-based fine-grained target classification and recognition method through a contour-based aircraft model recognition method, wherein the preset terminal is a display terminal.

[0090] According to an embodiment of the present invention, it further includes:

[0091] Acquire sample image data;

[0092] The sample image data is preprocessed to obtain the target object image data;

[0093] The target object image data is input into a preset E2EC instance segmentation model to obtain the target object contour image data;

[0094] A target object contour image dataset is established based on the target object contour image data.

[0095] It should be noted that, taking contour-based aircraft model recognition as an example, the sample image data is publicly available image data from the internet, consisting of optical aircraft images of various models, scenes, and attitudes. Ten different aircraft models were selected: F-15 fighter jet, F-16 fighter jet, F-22 fighter jet, JAS 39 fighter jet, IL-76 transport aircraft, 747 passenger aircraft, A-10 attack aircraft, B-2 bomber, E-2C early warning aircraft, and B-52 bomber. For each aircraft type, 240 optical images were selected, and each optical image contained one aircraft target. The images underwent correction and adjustment preprocessing to ensure the target was roughly centered in the image, and the images were resized to 500×500 pixels to obtain optical aircraft image data. Ten types of optical aircraft image datasets were then established based on this data.

[0096] Then, the ten types of optical aircraft image datasets are input into the trained E2EC instance segmentation model to obtain the actual contours of the aircraft targets. For example... Figure 8 As shown, E2EC outputs in two ways, as shown in Figure (b) and Figure (c). The method shown in Figure (c) yields the aircraft target contour image, which meets the experimental requirements of this paper. Next, a contour detection algorithm is used to adjust the size of the aircraft contour, as shown in Figure (d). Finally, an aircraft target contour image dataset containing 2400 actual aircraft target contour images is obtained.

[0097] According to an embodiment of the present invention, it further includes:

[0098] Obtain the 3D model data of the target object;

[0099] The three-dimensional model data of the target object is projected using a first preset method to obtain multiple projected image data of the target object model;

[0100] The target object model's multiple projected image data are preprocessed using a second preset method to obtain a three-dimensional model contour image dataset of the target object.

[0101] A preset contour-based fine-grained target classification and recognition network model is established based on the target object contour image dataset and the target object 3D model contour image dataset.

[0102] It should be noted that, taking contour-based aircraft model recognition as an example, AC3D software allows for scaling, rotation, translation, and material processing for each aircraft type. Here, the same aircraft models as those in the ten-type optical aircraft image dataset are selected, and 3D models of the ten aircraft types are constructed using AC3D software.

[0103] When observing an aircraft in three-dimensional space from different viewpoints, the aircraft will exhibit different attitudes. Therefore, for multi-view aircraft recognition, it is necessary to acquire aircraft attitude images from different viewpoints. Morphology mapping is an effective method for obtaining two-dimensional projection images of three-dimensional targets, but morphology mapping modeling is relatively complex. Therefore, simplified models are often used in practical engineering. Thus, this solution uses... Figure 9 The simplified model shown in the morphological diagram places the target object at the center of a unit observation sphere and then projects it along the direction from the sphere surface to the center of the sphere to obtain two-dimensional projected images from different viewpoints.

[0104] Then, the three-dimensional model data of the aircraft (target object) is projected using the first preset method, that is, using... Figure 9 The straight line from the center of the observation sphere to the vertex of the Y-axis is used as the baseline. First, the 3D model of the aircraft is displayed using the Wi-Fi reframe function, such as... Figure 10 As shown, to avoid the influence of the 3D model's built-in lighting and textures on subsequent contour extraction, the model is rotated and projected 20° around the X-axis 18 times. This process is called Step 1. Next, it is rotated 20° around the Z-axis once, and Step 1 is repeated until the Z-axis has been rotated 18 times. This process is called Step 2. Then, the X and Z-axis positions are swapped, and Steps 1 and 2 are repeated. Afterward, all projected images are rotated clockwise 18 times. Finally, 11,664 projected images are obtained for each type of aircraft model, containing the attitude of each aircraft model from different viewpoints.

[0105] After obtaining the projected images of various aircraft models, the second preset method is used to preprocess the multiple projected image data of each aircraft model. This involves using edge detection and contour detection algorithms to extract the aircraft contours from the projected images and adjusting their dimensions. Partial contour images of various aircraft models are shown below. Figure 11 As shown, its outline image size is 500×500.

[0106] Figure 2 A flowchart of a method for acquiring contour image data of a target object according to the present invention is shown.

[0107] like Figure 2 As shown in the embodiment of the present invention, the step of locating the target object and segmenting the target object contour through the target object contour extraction sub-network to obtain target object contour image data includes:

[0108] S202, perform feature extraction on the target object image data to obtain image feature information of the image data, and generate a heat map based on the image feature information;

[0109] S204, Analyze the heat map to obtain the center point of the target object within the target object image;

[0110] S206, Based on the heat map and the center point, analyze the initial target object contour image data to obtain the target object contour image data;

[0111] S208, refine the initial target object contour image data to finally obtain the target object contour image data.

[0112] It should be noted that the target object contour extraction subnetwork's role is to detect whether the optical image contains a target object and segment and extract the target object's contour. E2EC is a multi-stage, efficient end-to-end contour-based instance segmentation model that treats instance segmentation as a regression task, i.e., regressing the vertex coordinates of the contour represented by a series of discrete vertices to obtain the instance's contour. E2EC includes a contour initialization module, a global contour deformability module, and a contour thinning module, outputting a contour represented by 128 contour coordinate points.

[0113] Therefore, taking contour-based aircraft model recognition as an example, the target object contour extraction sub-network obtains the target object contour from the target object image based on E2EC. First, the target object image is input, and the DLA-34 feature extraction network is selected to extract image feature information. The generated heatmap is then input into the CenterNet object detection network to obtain the center point of the target object. The contour initialization module directly regresses the complete initial target object contour based on the center point features, such as... Figure 6 As shown in Figure (a), the global contour deformable module fine-tunes the contour of the target object based on all the initial target object contour points and center points to obtain a rough target object contour, as shown in Figure (a). Figure 6 As shown in Figure (b), the contour refinement module performs two refinements based on all the rough target object contour points and combines multi-directional alignment and dynamic matching loss functions to obtain the final target object contour, as shown in Figure (b). Figure 6 As shown in Figure (c), E2EC can perform excellent instance segmentation of target objects in optical images, obtaining high-quality actual contours of the target objects.

[0114] According to an embodiment of the present invention, the step of performing refined classification of the target object contour image data through a target object contour recognition sub-network to obtain classification data of the target object in the target object image data includes:

[0115] Feature extraction is performed on the target object contour image data to obtain the global features of the target object contour image and the feature map of the target object contour image;

[0116] The feature map of the target object contour image is cropped according to the third preset method to obtain multiple target object contour local region images with discriminative information.

[0117] Feature extraction is performed on the local region images of the target object contours with discriminative information to obtain the local features of the target object contour images and the local feature maps of the target object contour images.

[0118] The feature map and local feature map of the target object contour image are fused to obtain fine-grained features of different parts and scales of the target object contour image.

[0119] Aircraft model data is obtained by analyzing the fine-grained features of different parts and scales of the target object's contour image.

[0120] It should be noted that, taking contour-based aircraft model recognition as an example, feature description of the target object image is a crucial step in the aircraft model recognition task. This scheme uses contours as the feature description to represent the differences between different aircraft models and different viewpoints. Therefore, the role of the target object contour recognition sub-network is to further classify the actual aircraft contours obtained by the target object contour extraction sub-network to determine the specific aircraft model to which the contour image belongs. It is based on the proposed Dual-Stream Target Object Contour Recognition Network (DACR-Net), and its network structure is as follows: Figure 7 As shown, DACR-Net contains one global target branch and one local target branch. Because ResNet uses residual structural units to effectively solve the gradient vanishing problem in deep networks, it can improve feature extraction capabilities while continuously increasing the number of network layers, and it also has a small number of model parameters and good portability. Therefore, this paper chooses ResNet-50 as the feature extraction network for both branches, that is, removing the final fully connected layer from ResNet-50 and using it as the backbone of the feature extraction network, which can comprehensively extract the key features of the target object contour image.

[0121] First, the global branch extracts global features from the target object contour image and outputs the feature map of the target object contour image. Second, using F∈R... C×H×W Let represent the feature map with C channels and spatial size H×W output from the last convolutional layer of the input image S after passing through the ResNet-50-based feature extraction network.

[0122] Then, the Non-Maximum Suppression (NMS) algorithm is used to select a fixed number of windows as part regions of different scales, thereby obtaining multiple local region images of the target object contour with discriminative information through cropping, which is the third preset method.

[0123] Subsequently, the target local branch mainly extracts local features of the target object contour image and outputs feature maps based on the local region image of the target object contour. Finally, the convolutional features obtained from the feature extraction networks of the two branches are fused to obtain fine-grained features of different parts and scales of the target object contour image, which improves the feature representation ability to a certain extent and fully considers the impact of the distinguishable parts of the target object contour on classification.

[0124] Figure 3 A flowchart of a method for acquiring a local region image of the contour of a target object according to the present invention is shown.

[0125] like Figure 3 As shown in the embodiment of the present invention, the step of cropping the feature map of the target object contour image according to the third preset method to obtain multiple target object contour local region images with discriminative information includes:

[0126] S302, extract the features with maximum recognizability and minimum redundancy of the local region of the target object contour based on the feature map of the target object contour image to obtain multiple feature maps;

[0127] S304, Calculate the multiple feature maps to obtain the local region activation map K of the target object contour;

[0128] S306, perform aggregation calculation on the activation map K of the local region of the target object contour to obtain the average activation value of the activation map of the local region of the target object contour.

[0129] S308, based on the activation average value The feature map of the target object contour image is cropped to obtain multiple target object contour local region images with discriminative information.

[0130] It should be noted that, based on the feature map, the proposed Hybrid Attention Block Module (HACM) is used to extract features with maximum recognizability and minimum redundancy in the local region of the target object contour. Specifically, feature map F1 is compressed into feature map F2 through an average pooling layer. Feature map F2 then undergoes CBAM, where the input feature map F2 first passes through a channel attention module, multiplying the channel weights with the input feature map before being fed into a spatial attention module. The normalized spatial weights are multiplied with the input feature map of the spatial attention mechanism for adaptive feature refinement, outputting the final weighted feature map F3, thus obtaining richer attention features with different dimensions. The activation map K is then obtained by aggregating the feature maps F. Since regions with high activation values ​​in the activation map are often key areas, the sliding window concept from object detection is used to find informational windows as partial images. A fully convolutional network is used to implement the traditional sliding window method, obtaining feature maps of different windows from the feature map output from the previous branch. All window values ​​are sorted from largest to smallest; larger values ​​indicate greater information content in the region. Then, the Non-Maximum Suppression (NMS) algorithm, also known as the third preset method, is used to select a fixed number of windows as part regions of different scales, thereby obtaining multiple target object contour local region images with discriminative information through cropping. Subsequently, the target local branch mainly extracts local features of the target object contour image and outputs feature maps based on the target object contour local region images.

[0131] According to an embodiment of the present invention, the activation map K of the local region of the target object contour and the average activation value of the local region of the target object contour are... The calculation method is as follows:

[0132] The specific method for calculating the activation map K of the local region of the target object contour is as follows:

[0133]

[0134] Among them, f x It is the x-th feature map of the channel corresponding to feature map F.

[0135] The average activation value of the local activation map of the target object's contour. The specific calculation method is as follows:

[0136]

[0137] Where W and H are the width and height of a window’s feature map, and (x,y) is a specific position in the window’s activation map.

[0138] It should be noted that the activation map K of the local region of the target object contour is obtained by aggregating the feature maps F of the local region of the target object contour, where fx It is the x-th feature map of the corresponding channel of feature map F. The average activation value of the local region activation map of the target object contour. It is obtained by aggregating the activation map of the current window along the channel dimension, where W and H are the width and height of the feature map of a window, and (x,y) is a specific position in the activation map of a window.

[0139] According to an embodiment of the present invention, before establishing a contour-based target fine-grained classification and recognition deep neural network model, the method further includes:

[0140] The contour smoothing process is performed on the three-dimensional model contour image dataset of the target object using the fourth preset method.

[0141] It should be noted that, in order to make the target object's 3D model contour image dataset closer to the actual target object contour image dataset—that is, to minimize the domain gap between the two and approximate the contour shape of the target object in the real scene, thereby improving the model's recognition accuracy—a fourth preset method is used to smooth the contour of the target object's 3D model contour image dataset. This fourth preset method is the Savitzky-Golay (SG) smoothing filtering algorithm, which is a polynomial smoothing algorithm based on the least squares principle. The SG smoothing filtering algorithm mainly adjusts the degree of contour smoothing by changing the values ​​of the `window_length` and `polyorder` parameters.

[0142] Figure 4 A block diagram of the contour-based fine-grained target classification and recognition system of the present invention is shown.

[0143] like Figure 4 As shown, a second aspect of the present invention provides a contour-based fine-grained target classification and recognition system 4, including a memory 41 and a processor 42. The memory includes a contour-based fine-grained target classification and recognition method program, which, when executed by the processor, performs the following steps:

[0144] Acquire image data of the target object;

[0145] The target object image data is input into a preset contour-based fine-grained target classification and recognition network model for analysis. The target object is located and its contour is segmented through the target object contour extraction sub-network to obtain the target object contour image data.

[0146] The target object contour image data is refined by using a target object contour recognition sub-network to obtain the classification data of the target object in the target object image data;

[0147] The classification data of the target object is sent to a preset terminal for display.

[0148] According to an embodiment of the present invention, a preset contour-based fine-grained target classification and recognition network model is used to process the acquired target object image data, thereby obtaining the classification data of the target object in the target object image data. The preset contour-based fine-grained target classification and recognition network model comprises two sub-networks. The first sub-network is a target object contour extraction sub-network, which locates the target object from the optical image and segments the target object contour based on a contour-level instance segmentation network. The second sub-network is a target object contour recognition sub-network, which, based on the proposed Dual-Stream Target Object Contour Recognition Network (DACR-Net), performs fine-grained classification of the target object contour image, further recognizes the actual contour of the target object, and finally outputs the target object type corresponding to the target object contour image. The combination of the two can effectively complete the target object type recognition task based on optical images. This solution explains the contour-based fine-grained target classification and recognition method through a contour-based aircraft model recognition method, wherein the preset terminal is a display terminal.

[0149] According to an embodiment of the present invention, it further includes:

[0150] Acquire sample image data;

[0151] The sample image data is preprocessed to obtain the target object image data;

[0152] The target object image data is input into a preset E2EC instance segmentation model to obtain the target object contour image data;

[0153] A target object contour image dataset is established based on the target object contour image data.

[0154] It should be noted that, taking contour-based aircraft model recognition as an example, the sample image data is publicly available image data from the internet, consisting of optical aircraft images of various models, scenes, and attitudes. Ten different aircraft models were selected: F-15 fighter jet, F-16 fighter jet, F-22 fighter jet, JAS 39 fighter jet, IL-76 transport aircraft, 747 passenger aircraft, A-10 attack aircraft, B-2 bomber, E-2C early warning aircraft, and B-52 bomber. For each aircraft type, 240 optical images were selected, and each optical image contained one aircraft target. The images underwent correction and adjustment preprocessing to ensure the target was roughly centered in the image, and the images were resized to 500×500 pixels to obtain optical aircraft image data. Ten types of optical aircraft image datasets were then established based on this data.

[0155] Then, the ten types of optical aircraft image datasets are input into the trained E2EC instance segmentation model to obtain the actual contours of the aircraft targets. For example... Figure 8 As shown, E2EC outputs in two ways, as shown in Figure (b) and Figure (c). The method shown in Figure (c) yields the aircraft target contour image, which meets the experimental requirements of this paper. Next, a contour detection algorithm is used to adjust the size of the aircraft contour, as shown in Figure (d). Finally, an aircraft target contour image dataset containing 2400 actual aircraft target contour images is obtained.

[0156] According to an embodiment of the present invention, it further includes:

[0157] Obtain the 3D model data of the target object;

[0158] The three-dimensional model data of the target object is projected using a first preset method to obtain multiple projected image data of the target object model;

[0159] The target object model's multiple projected image data are preprocessed using a second preset method to obtain a three-dimensional model contour image dataset of the target object.

[0160] A preset contour-based fine-grained target classification and recognition network model is established based on the target object contour image dataset and the target object 3D model contour image dataset.

[0161] It should be noted that, taking contour-based aircraft model recognition as an example, AC3D software allows for scaling, rotation, translation, and material processing for each aircraft type. Here, the same aircraft models as those in the ten-type optical aircraft image dataset are selected, and 3D models of the ten aircraft types are constructed using AC3D software.

[0162] When observing an aircraft in three-dimensional space from different viewpoints, the aircraft will exhibit different attitudes. Therefore, for multi-view aircraft recognition, it is necessary to acquire aircraft attitude images from different viewpoints. Morphology mapping is an effective method for obtaining two-dimensional projection images of three-dimensional targets, but morphology mapping modeling is relatively complex. Therefore, simplified models are often used in practical engineering. Thus, this solution uses... Figure 9 The simplified model shown in the morphological diagram places the target object at the center of a unit observation sphere and then projects it along the direction from the sphere surface to the center of the sphere to obtain two-dimensional projected images from different viewpoints.

[0163] Then, the three-dimensional model data of the aircraft (target object) is projected using the first preset method, that is, using... Figure 9 The straight line from the center of the observation sphere to the vertex of the Y-axis is used as the baseline. First, the 3D model of the aircraft is displayed using the Wi-Fi reframe function, such as... Figure 10As shown, to avoid the influence of the 3D model's built-in lighting and textures on subsequent contour extraction, the model is rotated and projected 20° around the X-axis 18 times. This process is called Step 1. Next, it is rotated 20° around the Z-axis once, and Step 1 is repeated until the Z-axis has been rotated 18 times. This process is called Step 2. Then, the X and Z-axis positions are swapped, and Steps 1 and 2 are repeated. Afterward, all projected images are rotated clockwise 18 times. Finally, 11,664 projected images are obtained for each type of aircraft model, containing the attitude of each aircraft model from different viewpoints.

[0164] After obtaining the projected images of various aircraft models, the second preset method is used to preprocess the multiple projected image data of each aircraft model. This involves using edge detection and contour detection algorithms to extract the aircraft contours from the projected images and adjusting their dimensions. Partial contour images of various aircraft models are shown below. Figure 11 As shown, its outline image size is 500×500.

[0165] According to an embodiment of the present invention, the step of locating the target object and segmenting the target object contour through the target object contour extraction sub-network to obtain target object contour image data includes:

[0166] Feature extraction is performed on the target object image data to obtain image feature information of the image data, and a heat map is generated based on the image feature information;

[0167] Based on the analysis of the heat map, the center point of the target object within the target object image is obtained;

[0168] Based on the analysis of the heat map and the center point, the initial target object contour image data is obtained;

[0169] The initial target object contour image data is refined to obtain the final target object contour image data.

[0170] It should be noted that the target object contour extraction subnetwork's role is to detect whether the optical image contains a target object and segment and extract the target object's contour. E2EC is a multi-stage, efficient end-to-end contour-based instance segmentation model that treats instance segmentation as a regression task, i.e., regressing the vertex coordinates of the contour represented by a series of discrete vertices to obtain the instance's contour. E2EC includes a contour initialization module, a global contour deformability module, and a contour thinning module, outputting a contour represented by 128 contour coordinate points.

[0171] Therefore, taking contour-based aircraft model recognition as an example, the target object contour extraction sub-network obtains the target object contour from the target object image based on E2EC. First, the target object image is input, and the DLA-34 feature extraction network is selected to extract image feature information. The generated heatmap is then input into the CenterNet object detection network to obtain the center point of the target object. The contour initialization module directly regresses the complete initial target object contour based on the center point features, such as... Figure 6 As shown in Figure (a), the global contour deformable module fine-tunes the contour of the target object based on all the initial target object contour points and center points to obtain a rough target object contour, as shown in Figure (a). Figure 6 As shown in Figure (b), the contour refinement module performs two refinements based on all the rough target object contour points and combines multi-directional alignment and dynamic matching loss functions to obtain the final target object contour, as shown in Figure (b). Figure 6 As shown in Figure (c), E2EC can perform excellent instance segmentation of target objects in optical images, obtaining high-quality actual contours of the target objects.

[0172] According to an embodiment of the present invention, the step of performing refined classification of the target object contour image data through a target object contour recognition sub-network to obtain classification data of the target object in the target object image data includes:

[0173] Feature extraction is performed on the target object contour image data to obtain the global features of the target object contour image and the feature map of the target object contour image;

[0174] The feature map of the target object contour image is cropped according to the third preset method to obtain multiple target object contour local region images with discriminative information.

[0175] Feature extraction is performed on the local region images of the target object contours with discriminative information to obtain the local features of the target object contour images and the local feature maps of the target object contour images.

[0176] The feature map and local feature map of the target object contour image are fused to obtain fine-grained features of different parts and scales of the target object contour image.

[0177] Aircraft model data is obtained by analyzing the fine-grained features of different parts and scales of the target object's contour image.

[0178] It should be noted that, taking contour-based aircraft model recognition as an example, feature description of the target object image is a crucial step in the aircraft model recognition task. This scheme uses contours as the feature description to represent the differences between different aircraft models and different viewpoints. Therefore, the role of the target object contour recognition sub-network is to further classify the actual aircraft contours obtained by the target object contour extraction sub-network to determine the specific aircraft model to which the contour image belongs. It is based on the proposed Dual-Stream Target Object Contour Recognition Network (DACR-Net), and its network structure is as follows: Figure 7 As shown, DACR-Net contains one global target branch and one local target branch. Because ResNet uses residual structural units to effectively solve the gradient vanishing problem in deep networks, it can improve feature extraction capabilities while continuously increasing the number of network layers, and it also has a small number of model parameters and good portability. Therefore, this paper chooses ResNet-50 as the feature extraction network for both branches, that is, removing the final fully connected layer from ResNet-50 and using it as the backbone of the feature extraction network, which can comprehensively extract the key features of the target object contour image.

[0179] First, the global branch extracts global features from the target object contour image and outputs the feature map of the target object contour image. Second, using F∈R... C×H×W Let represent the feature map with C channels and spatial size H×W output from the last convolutional layer of the input image S after passing through the ResNet-50-based feature extraction network.

[0180] Then, the Non-Maximum Suppression (NMS) algorithm is used to select a fixed number of windows as part regions of different scales, thereby obtaining multiple local region images of the target object contour with discriminative information through cropping, which is the third preset method.

[0181] Subsequently, the target local branch mainly extracts local features of the target object contour image and outputs feature maps based on the local region image of the target object contour. Finally, the convolutional features obtained from the feature extraction networks of the two branches are fused to obtain fine-grained features of different parts and scales of the target object contour image, which improves the feature representation ability to a certain extent and fully considers the impact of the distinguishable parts of the target object contour on classification.

[0182] According to an embodiment of the present invention, the step of cropping the feature map of the target object contour image according to the third preset method to obtain multiple target object contour local region images with discriminative information includes:

[0183] Based on the feature map of the target object contour image, extract the features with the maximum recognizability and minimum redundancy of the local region of the target object contour to obtain multiple feature maps;

[0184] The activation map K of the local region of the target object contour is obtained by calculating the multiple feature maps;

[0185] The average activation value of the local region activation map K of the target object contour is obtained by aggregation calculation.

[0186] According to the activation average The feature map of the target object contour image is cropped to obtain multiple target object contour local region images with discriminative information.

[0187] It should be noted that, based on the feature map, the proposed Hybrid Attention Block Module (HACM) is used to extract features with maximum recognizability and minimum redundancy in the local region of the target object contour. Specifically, feature map F1 is compressed into feature map F2 through an average pooling layer. Feature map F2 then undergoes CBAM, where the input feature map F2 first passes through a channel attention module, multiplying the channel weights with the input feature map before being fed into a spatial attention module. The normalized spatial weights are multiplied with the input feature map of the spatial attention mechanism for adaptive feature refinement, outputting the final weighted feature map F3, thus obtaining richer attention features with different dimensions. The activation map K is then obtained by aggregating the feature maps F. Since regions with high activation values ​​in the activation map are often key areas, the sliding window concept from object detection is used to find informative windows as partial images. A fully convolutional network is used to implement the traditional sliding window method, obtaining feature maps of different windows from the feature map output from the previous branch. All window values ​​are sorted from largest to smallest; larger values ​​indicate greater information content in the region. Then, the Non-Maximum Suppression (NMS) algorithm, also known as the third preset method, is used to select a fixed number of windows as part regions of different scales, thereby obtaining multiple target object contour local region images with discriminative information through cropping. Subsequently, the target local branch mainly extracts local features of the target object contour image and outputs feature maps based on the target object contour local region images.

[0188] According to an embodiment of the present invention, the activation map K of the local region of the target object contour and the average activation value of the local region of the target object contour are... The calculation method is as follows:

[0189] The specific method for calculating the activation map K of the local region of the target object contour is as follows:

[0190]

[0191] Among them, f x It is the x-th feature map of the channel corresponding to feature map F.

[0192] The average activation value of the local activation map of the target object's contour. The specific calculation method is as follows:

[0193]

[0194] Where W and H are the width and height of a window’s feature map, and (x,y) is a specific position in the window’s activation map.

[0195] It should be noted that the activation map K of the local region of the target object contour is obtained by aggregating the feature maps F of the local region of the target object contour, where f x It is the x-th feature map of the corresponding channel of feature map F. The average activation value of the local region activation map of the target object contour. It is obtained by aggregating the activation map of the current window along the channel dimension, where W and H are the width and height of the feature map of a window, and (x,y) is a specific position in the activation map of a window.

[0196] According to an embodiment of the present invention, before establishing a contour-based target fine-grained classification and recognition deep neural network model, the method further includes:

[0197] The contour smoothing process is performed on the three-dimensional model contour image dataset of the target object using the fourth preset method.

[0198] It should be noted that, in order to make the 3D model contour image dataset of the target object closer to the style of the actual contour image dataset of the target object, that is, to minimize the domain gap between the two and approximate the contour shape of the target object in the real scene, thereby improving the recognition accuracy of the model, the 3D model contour image dataset of the target object is smoothed using a fourth preset method. The fourth preset method is the Savitzky-Golay (SG) smoothing filtering algorithm, which is a polynomial smoothing algorithm based on the least squares principle. The SG smoothing filtering algorithm mainly adjusts the degree of contour smoothing by changing the values ​​of the two parameters: window_length and polyorder.

[0199] This invention discloses a contour-based fine-grained target classification and recognition method, system, and storage medium. The method includes: acquiring target object image data; inputting the target object image data into a preset contour-based fine-grained target classification and recognition network model for analysis, locating the target object and segmenting its contour through a target object contour extraction sub-network to obtain target object contour image data; performing fine-grained classification on the target object contour image data through a target object contour recognition sub-network to obtain classification data of the target object in the target object image data; and sending the classification data of the target object to a preset terminal for display. This invention achieves the purpose of fine-grained classification of target objects by locating them and identifying their types based on their contours.

[0200] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0201] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0202] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0203] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0204] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

Claims

1. A contour-based fine-grained target classification and recognition method, characterized in that, include: Acquire image data of the target object; The target object image data is input into a preset contour-based fine-grained target classification and recognition network model for analysis. The target object is located and its contour is segmented through the target object contour extraction sub-network to obtain the target object contour image data. The target object contour image data is refined by using a target object contour recognition sub-network to obtain the classification data of the target object in the target object image data; The classification data of the target object is sent to a preset terminal for display; The step of performing refined classification of the target object contour image data through a target object contour recognition sub-network to obtain classification data of the target object in the target object image data includes: Feature extraction is performed on the target object contour image data to obtain the global features of the target object contour image and the feature map of the target object contour image; The feature map of the target object contour image is cropped according to the third preset method to obtain multiple target object contour local region images with discriminative information. Feature extraction is performed on the local region images of the target object contours with discriminative information to obtain the local features of the target object contour images and the local feature maps of the target object contour images. The feature map and local feature map of the target object contour image are fused to obtain fine-grained features of different parts and scales of the target object contour image. Aircraft model data is obtained by analyzing the fine-grained features of different parts and scales of the target object's contour image. The step of cropping the feature map of the target object contour image according to the third preset method to obtain multiple target object contour local region images with discriminative information includes: Based on the feature map of the target object contour image, extract the features with the maximum recognizability and minimum redundancy of the local region of the target object contour to obtain multiple feature maps; The activation map K of the local region of the target object contour is obtained by calculating the multiple feature maps; The activation average value of the local region activation map K of the target object contour is obtained by performing aggregation calculation on the local region activation map of the target object contour. The feature map of the target object contour image is cropped based on the activation average value to obtain multiple target object contour local region images with discriminative information.

2. The contour-based fine-grained target classification and recognition method according to claim 1, characterized in that, Also includes: Acquire sample image data; The sample image data is preprocessed to obtain the target object image data; The target object image data is input into a preset E2EC instance segmentation model to obtain the target object contour image data; A target object contour image dataset is established based on the target object contour image data.

3. The contour-based fine-grained target classification and recognition method according to claim 2, characterized in that, Also includes: Obtain the 3D model data of the target object; The target object's three-dimensional model data is projected using a first preset method to obtain multiple projected image data of the target object's three-dimensional model. The target object's three-dimensional model is preprocessed using a second preset method to obtain a dataset of the target object's three-dimensional model contour images. A preset contour-based fine-grained target classification and recognition network model is established based on the target object contour image dataset and the target object 3D model contour image dataset.

4. The contour-based fine-grained target classification and recognition method according to claim 1, characterized in that, The step of locating the target object and segmenting its contour through a target object contour extraction sub-network to obtain target object contour image data includes: Feature extraction is performed on the target object image data to obtain image feature information of the image data, and a heat map is generated based on the image feature information; Based on the analysis of the heat map, the center point of the target object within the target object image is obtained; Based on the analysis of the heat map and the center point, the initial target object contour image data is obtained; The initial target object contour image data is refined to obtain the final target object contour image data.

5. A contour-based fine-grained target classification and recognition system, used to implement the contour-based fine-grained target classification and recognition method according to any one of claims 1-4, characterized in that, The system includes a memory and a processor. The memory contains a contour-based fine-grained target classification and recognition method program. When executed by the processor, the contour-based fine-grained target classification and recognition method program performs the following steps: Acquire image data of the target object; The target object image data is input into a preset contour-based fine-grained target classification and recognition network model for analysis. The target object is located and its contour is segmented through the target object contour extraction sub-network to obtain the target object contour image data. The target object contour image data is refined by using a target object contour recognition sub-network to obtain the classification data of the target object in the target object image data; The classification data of the target object is sent to a preset terminal for display.

6. The contour-based fine-grained target classification and recognition system according to claim 5, characterized in that, The step of locating the target object and segmenting its contour through a target object contour extraction sub-network to obtain target object contour image data includes: Feature extraction is performed on the target object image data to obtain image feature information of the image data, and a heat map is generated based on the image feature information; Based on the analysis of the heat map, the center point of the target object within the target object image is obtained; Based on the analysis of the heat map and the center point, the initial target object contour image data is obtained; The initial target object contour image data is refined to obtain the final target object contour image data.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a contour-based fine-grained target classification and recognition method program, which, when executed by a processor, implements the steps of the contour-based fine-grained target classification and recognition method as described in any one of claims 1 to 4.