Infrared sea and air target identification method based on MMdesection

Through a custom network and self-built data set based on MMdetection, the existing infrared sea-air target recognition method relies on expert experience and computation complex problems, and achieve high-precision and efficient infrared sea-air target recognition, especially in improving the detection rate of small-scale targets.

CN120047814APending Publication Date: 2025-05-27THE 53RD RES INST OF CHINA ELECTRONICS TECH CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411935276.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing infrared sea-air target recognition methods have problems such as relying on expert experience, complex calculations, and low detection rate of small-scale targets, which are difficult to meet the requirements of accurate, robust and rapid identification in practical applications.

Method used

The infrared sea-air target recognition method based on MMdetection is adopted, and by customizing the backbone network, establishing self-built data sets and performing training, a multi-scale cooperative feature network, a backbone feature extraction network, a feature fusion network, and a classification and identification network are built to improve the detection rate of small-scale targets.

Benefits of technology

It realizes high-precision infrared sea and air target recognition, improves the detection rate of small-scale targets, reduces the difficulty of development, and is suitable for scenarios with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047814A_ABST
    Figure CN120047814A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of target classification and recognition, and particularly relates to an infrared sea and air target recognition method based on MMdesection. The method comprises the following steps: making an infrared sea and air target data set, wherein classified targets comprise airplanes and ships; building a self-defined sea and air target convolutional neural classification network on the basis of MMdesign; testing after the OpenMMLab environment is configured, and determining that the environment is ready; a configuration file is compiled according to task requirements, the configuration file comprises a classification network, a data set, an optimization strategy and operation configuration, and network training is executed in the environment; and after the network training is completed, testing the pictures in the test set to realize the classification and identification of the sea and air targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target classification and recognition, and particularly relates to an infrared sea and air target recognition method based on MMdetection. Background Technique

[0002] In recent years, with the rise of artificial intelligence, image classification and target recognition based on deep learning have received extensive attention and research in the academic and industrial fields. Infrared images can work all-weather, can provide "infrared night vision" images for pilots, are less affected by weather, and can still observe the situation around the aircraft through infrared images under adverse weather conditions, and have high detail resolution ability in various complex environments. Sea and air targets, as a type of high-value targets, have wide applications in civil and military fields such as military early warning, infrared guidance, and maritime search and rescue. When detecting infrared sea and air targets at a long distance, due to factors such as unclear target features and low signal-to-noise ratio, they are easily submerged by the background. Therefore, an accurate, robust, and fast infrared sea and air target recognition method is still a challenging research topic.

[0003] Currently, the main infrared sea and air target recognition methods include traditional pattern recognition-based methods and deep learning-based methods. Among them, traditional methods generally include three steps: region screening, feature extraction, and recognition and classification. Region screening usually uses the sliding window method to initially determine the approximate location of the target; feature extraction is to manually extract target features through artificial experience, and common infrared target features mainly include features such as gray level and contrast; recognition and classification is to classify the extracted features according to a certain evaluation criterion, generally by setting thresholds or training feature classifiers. And deep learning-based methods automatically extract target features through convolutional neural networks. These current common methods generally have the following defects and deficiencies:

[0004] (1) Most traditional methods rely heavily on expert experience to manually extract features. However, infrared images usually have low resolution, and the size of the target of interest generally accounts for a small proportion in the image. Therefore, it is difficult to extract representative features to express and describe the target. In addition, the setting of classification discrimination rule thresholds and the adjustment of hyperparameters require a large amount of engineering practice accumulation and cannot be applied to all scenarios generally.

[0005] (2) Due to the strong confidentiality of infrared small target datasets, few public datasets, and the inherent characteristics of small target size and weak signal, directly applying conventional deep learning-based target detection and recognition methods to this field has unsatisfactory effects.

[0006] (3) In addition, some target recognition methods are computationally complex and have high requirements for computing resources. For scenarios with scarce computing resources, they do not have practical use value and are difficult to meet the real-time detection requirements.

[0007] Considering the complex scenarios and limited computing resources in actual applications, it is necessary to develop a sea-air target recognition model with accurate recognition, high robustness, flexible deployment, and fast calculation to provide reference for subsequent related research. Summary of the Invention

[0008] In view of this, the present invention aims to provide an infrared sea-air target recognition method based on MMdetection. By customizing the backbone network of MMdetection, establishing its own dataset and performing training, it has high detection accuracy. The self-built multi-level classification network has good detection effects on sea-air targets of various scales, especially improving the detection rate of small-scale targets.

[0009] The technical solution of the present invention is implemented as follows:

[0010] An infrared sea-air target recognition method based on MMdetection, the specific process is as follows:

[0011] Make an infrared sea-air target dataset, and the classification targets include airplanes and ships;

[0012] Build a custom sea-air target convolutional neural classification network based on MMdetection;

[0013] After configuring the OpenMMLab environment, test to determine that the environment is ready;

[0014] Write a configuration file according to the task requirements. The configuration file includes a classification network, a dataset, an optimization strategy, and a running configuration, and perform network training in the environment;

[0015] After the network training is completed, test the pictures in the test set to realize the classification and recognition of sea-air targets.

[0016] Furthermore, the custom sea-air target convolutional neural classification network built by the present invention based on MMdetection includes: a multi-scale co-feature network, a backbone feature extraction network, a feature fusion network, and a classification and recognition network; among them,

[0017] The multi-scale co-feature network is used to extract the shallow features of the image; specifically: the original image is downsampled to obtain the initial co-feature F 0 , the initial co-feature F 0 It is also used to generate the co-feature Fi in the process stage, where i represents the i-th process co-feature, and each co-feature is dimensionally aligned with the backbone feature extraction network through a convolution operation;

[0018] The backbone extraction network is used to extract the backbone features of the image and perform transformation and fusion with the shallow features extracted by the multi-scale co-feature network;

[0019] Feature fusion network, which uses a weighted fusion mechanism to add a weight to each scale feature that needs to be feature-fused to adjust the contribution degree of different scale feature maps to the output feature map;

[0020] The classification and recognition network includes a class prediction network and a location prediction network. The input of the class prediction network is the fused feature output by the feature fusion network, and the output is the class of the target. The input of the location prediction network is the fused feature output by the feature fusion network, and the output is the location of the target.

[0021] Further, the backbone extraction network of the present invention is EfficientNet, and the feature fusion network is BiFPN.

[0022] Further, the infrared sea and air target dataset produced by the present invention contains infrared sea and air target images under multiple different scenarios, different time periods, and different resolutions.

[0023] Further, after collecting the target images, the present invention also performs preprocessing, and the preprocessing includes data cleaning, data integration, data transformation, and data annotation;

[0024] Data cleaning: deleting duplicate data and unifying the data format;

[0025] Data integration: deleting and integrating redundant data and merging data from multiple different sources;

[0026] Data transformation: scaling the data to the same scale through normalization;

[0027] Data annotation: annotating the sea and air targets in the data according to the class.

[0028] Further, when the basic configuration file needs to be used in the present invention, it is carried out by inheritance.

[0029] Further, after configuring the OpenMMLab environment in the present invention, it is tested to determine that the environment is ready. Specifically: by calling an existing model and loading existing weights to perform object detection inference on the content of the picture. If the objects in the picture can be successfully detected, it means that the environment is set up.

[0030] Further, when establishing the dataset in the present invention, the collected data is divided into a training set, a validation set, and a test set according to a ratio of 6:2:2.

[0031] Further, after the configuration file of the present invention is written, the model is trained and verified through the entry script tools / train.py; in train.py, the dataset and the model are initialized, and an executor is built by calling the high-level API train_model, and this executor is responsible for scheduling the training and verification steps of the network model; in train_step, the forward_train method of the network model is called by setting return_loss=True to return the loss, and the data also enters the classification network structure during this process.

[0032] Further, the performance test of the model of the present invention is as follows: using the test dataset, the classification network model is called in the high-level API single_gpu_test or multi_gpu_test for testing.

[0033] Beneficial effects

[0034] First, compared with other existing infrared sea and air target recognition methods, the infrared sea and air target recognition method based on MMdetection proposed by the present invention does not need to rely on expert experience to manually extract features, avoiding the cumbersome process and low generalization ability of manually designed features; in addition, a multi-scale convolutional neural network is designed. Compared with neural network models with hundreds of layers, the model proposed by the present invention is more lightweight; finally, based on MMdetection for deployment, the development difficulty is greatly reduced, which can be used as a reference for relevant work and research.

[0035] Second, the detection accuracy of the present invention is high. The self-built multi-level classification network has good detection effects on sea and air targets of various scales, especially improving the detection rate of small-scale targets. Description of the drawings

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 is the network structure diagram customized by the present invention based on MMdetection;

[0038] Figure 2 is the network detection flow chart customized by the present invention based on MMdetection;

[0039] Figure 3 is the sea and air target detection result diagram based on MMdetection. Detailed implementation manners

[0040] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0041] It should be noted that, without conflict, the following embodiments and the features in the embodiments may be combined with each other; and, based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0042] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of the aspects described herein may be used to implement the device and / or practice the method. Additionally, this device may be implemented and this method may be practiced using other structures and / or functionality in addition to one or more of the aspects described herein.

[0043] The embodiments of the present application provide an infrared sea and air target recognition method based on MMdetection. The process of realizing infrared sea and air target recognition by the present invention will be briefly described below through examples.

[0044] Step S1: Set up the environment.

[0045] Use Anaconda to set up a virtual environment. First, install the mmcv library, and then install the mmdetection library and compile it. After the environment installation is completed, use the demo provided by OpenMMLab official to test whether the environment is ready. Perform object detection inference on the picture content by calling an existing model and loading existing weights. If the objects in the picture can be successfully detected, it means that the environment setup is completed.

[0046] Step S2: Create an infrared sea and air target dataset.

[0047] Generally, since the detection device is usually far from the target, the target occupies a small proportion in the whole image. Moreover, due to the few texture features in the infrared image and the low contrast with the background, the target is extremely likely to be submerged in the complex background environment. These characteristics make the infrared sea and air target detection more difficult. In addition, the currently publicly available infrared sea and air target data is relatively scarce. Therefore, the present invention collates multiple publicly available network datasets and the existing datasets in the laboratory, and classifies them according to the target types, mainly divided into two types of targets: ships and aircraft (ships are set as class 0 targets, and aircraft are set as class 1 targets), and establishes its own sea and air target dataset. This dataset contains infrared sea and air targets in multiple different scenarios, at different times, and with different resolutions for subsequent model training and generation.

[0048] The data collected for the first time is the original dataset, which cannot be directly used. According to the requirements of the target detection task, the original dataset is transformed into a standard dataset of infrared sea and air targets through steps such as data cleaning, data conversion, and data partitioning. Among them, ① Data cleaning: Since there are duplicate and inconsistently formatted data in the original data containing aircraft targets and ship targets, it is necessary to delete the duplicate data in the original data, check the data consistency for the inconsistently formatted data, and standardize and unify the data format and naming. ② Data integration: Since multiple data from different sources are merged, it is necessary to delete and integrate redundant data, resolve conflicts and inconsistencies between data, and also delete other irrelevant data in the original dataset. ③ Data transformation: Non-numerical data needs to be converted into numerical data. For example, the category data of aircraft and ships, ships are set as class 0 targets, and aircraft are set as class 1 targets; in addition, the data scales of infrared sea and air targets from different sources are different, and the data is scaled to the same scale through normalization. ④ Data annotation: The sea and air targets in the original dataset are annotated according to the categories. ⑤ Data partitioning: The nearly 12,000 photos collected are divided into a training set, a validation set, and a test set according to the ratio of 6:2:2.

[0049] Step S3: Customize the backbone network.

[0050] Create a custom backbone network using PyTorch under the models file. During the creation process, the backbone network and main modules need to inherit from mmcv.runner.BaseModule, which is a subclass of torch.nn.Module. It behaves exactly the same as the backbone network class except for additionally supporting the use of the init_cfg parameter to specify initialization methods including pre-trained models. Register the backbone network class to the mmcls.models.BACKBONES registerer. If additional modules or components need to be added, refer to models / det / layer_decay_optimizer_constructor.py and add them to the corresponding folder. Add the custom backbone network and custom components to the models / __init__.py file.

[0051] Due to the particularity of the application scenario, infrared sea and air targets are mostly small and weak targets. To improve the detection and perception capabilities of small and weak infrared sea and air targets, this invention performs feature fusion on shallow and deep layer features through cross-level cascading to enhance the expressive ability of the output feature map, enabling it to contain both shallow features such as pixel information and deep features such as semantic information, thereby enhancing the network model's attention to small-scale targets.

[0052] The specific implementation of the network is as follows: ① First, design a multi-scale co-feature network MSCLFNet (Multi-scale Co-light-feature Network) to extract shallow features such as spatial information, pixel information, and detail information. Downsample the original image and perform a series of combined operations including CONV (convolution), BN (batch normalization), and ReLU activation to extract the initial co-feature F0. F0 is also used to generate the co-feature Fi in the process stage, where i represents the i-th process co-feature. In addition, align the dimensions of the co-feature and backbone feature through 1*1 convolution operations. ② Subsequently, referring to the improvement of YOLOv7, select EfficientNet as the backbone feature extraction network to balance the model's accuracy and complexity. Extract the backbone feature through EfficientNet and fuse it with the shallow feature to fully retain the spatial information and small-scale target information. ③ Perform transformation fusion on the backbone feature extracted by EfficientNet and the shallow feature extracted by MSCLFNet through the feature fusion module, and input the feature map fused at each level into the BiFPN layer. ④ Input the output feature map of step ③ into the classification and regression network for direct classification and localization. The custom network structure of this invention is as Figure 1-2 shown.

[0053] Since the FPN (Feature Pyramid Network) was proposed, it has been applied to multi-scale feature fusion. However, in the past, when fusing different feature maps in FPN, it was simply adding them together. In fact, it is necessary to consider that the contributions of feature maps with different resolutions in the feature fusion process are usually uneven. To solve this problem, BiFPN is used instead of FPN in the present invention. In BiFPN, learnable weights are introduced for each input to learn the importance of feature maps at different scales.

[0054] In addition, the top-down and bottom-up multi-scale feature fusion is used repeatedly. In Figure 1 BiFPN layer is repeatedly used as a separate feature network layer to promote higher-level feature fusion, combining co-feature and backbone feature fusion to improve the model's information extraction ability for multi-scale targets. In addition, as a simplified bidirectional network, BiFPN has a lower network complexity. In the present invention, the overall accuracy and efficiency of the model are improved by adopting the BiFPN network structure.

[0055] In the specific implementation of the backbone network in this embodiment, first, a multi-scale co-feature network MSCLFNet is customized to extract shallow features; second, the output feature maps of the multi-scale co-feature network are not simply merged in channels, but are fused with the output feature maps of some layers of the backbone feature network (since the EfficientNet network can well balance the accuracy and complexity of the network, the EfficientNet network is used to extract backbone features) to fully retain spatial information and small-scale target information (many infrared targets in the dataset are small-scale); third, when fusing the feature maps of the co-feature network and the backbone feature network, it is not directly merged in channels, but after aligning the dimensions using 1*1 convolution, through the weighted feature fusion mechanism of the BiFPN network, a weight is added to each scale feature that needs to perform feature fusion, aiming to adjust the contribution degree of different scale feature maps to the output feature map.

[0056] The design of the above backbone network starts from the requirements of the infrared target recognition task. Designing a network model that fuses the co-feature network and the backbone feature network is to improve the model's information extraction ability for multi-scale target information, especially small-scale target information; the selection of EfficientNet and BiFPN takes into account the balance between detection accuracy and complexity, and at the same time considers that the contribution rates of feature maps with different resolutions to the output feature map are different.

[0057] Step S4: Add configuration files. The basic configuration files include models, datasets, optimization strategies, and running configurations. If you need to use the basic configuration files, you can do so by inheritance. The basic configuration files can be selected to be used or overridden according to requirements. In the case of not using the basic configuration files, all relevant configurations can be written into the same configuration file. For different tasks, the configuration files should be stored in the corresponding subfolders under the configs directory of MMdetection. Many configuration files of classic models are provided in the configs folder, which can be modified according to the usage requirements to match the custom network model.

[0058] In this embodiment, the optimization strategy includes existing improvement strategies and also improvement strategies made by developers who build the network based on mmdetection themselves. It mainly includes improvements to the Backbone module, Neck module, and Bbox_head module. In addition, there are also improvements to data augmentation strategies, parameter optimizer (optimizer) improvements, training method improvements, target bounding box screening mechanism improvements, etc.

[0059] Step S5: Network training. After the configuration file is written, use the entry script tools / train.py to perform training and validation on the model. Initialize the dataset and the model in train.py, and build an executor by calling the high-level API train_model. This executor is responsible for scheduling the training and validation steps of the network model. In train_step, call the forward_train method of the network model by setting return_loss = True to return the loss, and the data also enters the backbone network and detection head structures during this process.

[0060] Step S6: Model performance testing. Call the above-mentioned custom network model for testing in the high-level APIs single_gpu_test or multi_gpu_test. The specific steps are to complete parsing command parameters, building the dataset and data loader, building and encapsulating the network model in the entry script tools / train.py, call the high-level APIs single_gpu_test or multi_gpu_test to traverse the entire data loader, and set the parameter return_loss = False to call the forward_test function of the network model to make the model output detection results. In the present invention, the model will finally frame the position of the target in the image and output the target detection category (ship or aircraft) at the same time.

[0061] In the embodiment of this application, MMdetection decomposes the network framework into multiple components through modular combination design, and encapsulates processes such as dataset construction, model building, and training process design into modules. Based on the unified and flexible architecture of MMdetection, relevant modules can be called in combination according to requirements, and a custom network model can be constructed, which will greatly reduce the development difficulty. This method customizes a multi-scale convolutional classification network based on MMdetection, designs two feature extraction networks to extract backbone features and co-features respectively, and fuses them. Fully considering the different contributions of feature maps with different resolutions to the detection results, during the feature fusion process of the network model, BiFPN (weighted bidirectional feature pyramid network) is added to improve the detection rate of the model for infrared sea and air targets with different scales.

[0062] The method proposed in this invention has good detection effects on the test set. Especially for small-scale targets, it also has high detection accuracy. At the same time, the proposed network model takes into account both performance and efficiency, laying a foundation for subsequent application to embedded devices. The sea and air target detection results based on MMdetection are as Figure 2 shown.

[0063] As mentioned above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An infrared sea and air target recognition method based on MMdetection, characterized in that: The specific process is: Create infrared sea and air target datasets, with classified targets including aircraft and ships; Build a custom sea and sky target convolutional neural classification network based on MMdetection; After configuring the OpenMMLab environment, test it to make sure it is ready; Write a configuration file according to task requirements, the configuration file includes a classification network, a data set, an optimization strategy, and a running configuration, and perform network training in the environment; After the network training is completed, the images in the test set are tested to achieve classification and recognition of sea and air targets.

2. The infrared sea and air target recognition method based on MMdetection according to claim 1 is characterized in that: The custom sea and sky target convolutional neural classification network based on MMdetection includes: a multi-scale co-feature network, a backbone feature extraction network, a feature fusion network, and a classification and recognition network; wherein, The multi-scale co-feature network is used to extract shallow features of the image. Specifically, the original image is downsampled to obtain the initial co-feature F0, which is also used to generate the co-feature Fi in the process stage. i represents the i-th process co-feature, and each co-feature is aligned to the dimension of the backbone feature extraction network through a convolution operation. Backbone extraction network is used to extract the backbone features of the image and transform and fuse them with the shallow features extracted by the multi-scale co-feature network; The feature fusion network uses a weighted fusion mechanism to add a weight to each scale feature that needs to be fused, so as to adjust the contribution of feature maps of different scales to the output feature map; The classification and recognition network includes a category prediction network and a position prediction network. The category prediction network inputs the fused features output by the feature fusion network, and outputs the category of the target. The position prediction network inputs the fused features output by the feature fusion network, and outputs the position of the target.

3. The infrared sea and air target recognition method based on MMdetection according to claim 2 is characterized in that: The backbone extraction network is EfficientNet, and the feature fusion network is BiFPN.

4. The infrared sea and air target recognition method based on MMdetection according to claim 1 is characterized in that: The infrared sea and sky target data set comprises infrared sea surface and sky target images in multiple different scenes, at different time periods and with different resolutions.

5. The infrared sea and air target recognition method based on MMdetection according to claim 4 is characterized in that: After the target image is collected, preprocessing is also performed, and the preprocessing includes data cleaning, data integration, data transformation and data labeling; Data cleaning: delete duplicate data and unify data formats; Data integration: delete and integrate redundant data, and merge data from multiple different sources; Data transformation: scaling the data to the same scale through normalization; Data labeling: Label the sea and air targets in the data according to their categories.

6. The infrared sea and air target recognition method based on MMdetection according to claim 1 is characterized in that: When you need to use the base configuration file, do it through inheritance.

7. The infrared sea and air target recognition method based on MMdetection according to claim 1 is characterized in that: After the OpenMMLab environment is configured, a test is performed to ensure that the environment is ready. Specifically, the existing model is called and the existing weights are loaded to perform target detection reasoning on the image content. If the object in the image can be successfully detected, the environment is set up.

8. The infrared sea and air target recognition method based on MMdetection according to claim 5 is characterized in that: When establishing the data set, the collected data is divided into training set, validation set and test set in a ratio of 6:2:

2.

9. The infrared sea and air target recognition method based on MMdetection according to claim 8, characterized in that: After the configuration file is written, the model is trained and verified through the entry script tools / train.py; the data set and model are initialized in train.py, and the executor is built by calling the high-level API train_model. The executor is responsible for scheduling the training and verification steps of the network model; in train_step, the forward_train method of the network model is called by setting return_loss = True to return the loss, and the data also enters the classification network structure in this process.

10. The infrared sea and air target recognition method based on MMdetection according to claim 8, characterized in that: The model performance test is to use the test dataset to call the classification network model in the high-level API single_gpu_test or multi_gpu_test for testing.