Establishment of Tank Image Dataset Based on Deep Learning and Key Point Detection Method
By establishing a tank image data set and adopting a deep learning model, the problem of lack of tank data is solved, efficient and accurate tank key point detection is achieved, and the accuracy and speed of tank recognition and strike are improved.
Patent Information
- Application Number
- CN202211053038.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-30
AI Technical Summary
The lack of tank data leads to inaccurate positioning of key points, making it difficult to achieve high-precision strike and damage assessment.
By establishing a tank image data set based on deep learning, including tank detection model and keypoint detection model, image data augmentation processing is performed, training, verification and test sets are generated, and model training is used using Keypoint RCNN and keypoint-rcnn networks are used for model training, and homovariance uncertainty learning is introduced to optimize loss function, simplifying the model structure to reduce the amount of operations.
Efficiently generate a large number of available data sets, reduce the workload of manual labeling, improve the accuracy and speed of tank key point detection, reduce the amount of computing, and achieve fast and accurate key point recognition.
Smart Images

Figure CN115439712B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a tank image data set establishment and key point detection method based on deep learning. Background Art
[0002] Modern tanks integrate new technologies, integrate offense and defense, and have excellent mobility and firepower. They will continue to play an important role in future battlefields. Tanks have always been the most important assault weapon. In addition to having strong ground assault capabilities, amphibious tanks are also good at sea assaults, and airborne tanks can carry out airborne assaults behind enemy lines and in strategic locations. Therefore, attacks and research on tanks are still important.
[0003] As war gradually enters the era of modernization, informatization and intelligence, computer processing of battlefield information is becoming increasingly important. Modern tanks integrate new technologies, integrate offense and defense, and have excellent mobility and firepower. They will continue to play an important role in future battlefields. The attack on enemy tanks must be transformed from high firepower and high armor penetration to high precision. Therefore, the identification of tank key points is particularly important. Accurate and rapid identification of key points can guide precision strikes and is also of great help in the damage assessment of enemy and friendly tanks in the later stage.
[0004] However, since tanks can only appear in special situations such as wars and dramas, it is difficult to collect tank data at present. The lack of tank data can easily lead to inaccurate positioning of tank key points. Summary of the invention
[0005] This application provides a tank image dataset establishment and key point detection method based on deep learning, which can be used to solve the current technical problem of tank data shortage.
[0006] This application provides a tank image dataset establishment and key point detection method based on deep learning, the method includes:
[0007] Input the image to be detected into the recognition model to detect whether there is a tank, and if there is a tank, output a result image with a bounding box and key point labels; the recognition model includes a tank detection model and a tank key point detection model;
[0008] The identification model is determined by the following method:
[0009] Step 1: Establish tank image dataset, mask and key point annotation information;
[0010] Step 2, batch generate tank image datasets, and divide the tank image datasets into training set, validation set and test set;
[0011] Step 3, establish a tank detection model;
[0012] Step 4, establish a tank key point detection model;
[0013] Step 5, perform enhancement processing on the image data;
[0014] Step 6, train the tank detection model and the tank key point detection model respectively.
[0015] Optionally, establish a tank image data set, a mask, and key point annotation information, including:
[0016] Step 101, collect tank images and perform preprocessing;
[0017] Step 102, mark the key points in the tank images; the key points include wheel pairs, gun barrels, muzzles, front armor, rear armor, side armor, and top armor;
[0018] Step 103, mark the tank bounding box in the tank image, obtain the mask, and generate the minimum bounding box;
[0019] Step 104, store the obtained marking information in a json file and unify the format.
[0020] Optionally, batch generate a tank image data set and divide the tank image data set into a training set, a validation set, and a test set,
[0021] including:
[0022] Step 201, collect background pictures and adjust their sizes to a preset size;
[0023] Step 202, collect background noise images and make masks for the background noise; the background noise is the objects other than the tank in the background picture;
[0024] Step 203, randomly select one or several background noises, perform size and rotation processing, and then add them to the background picture;
[0025] Step 204, randomly select several tank pictures and perform processing on brightness, contrast, size, angle, and key point coordinates;
[0026] Step 205, generate a data set and have the json format required for training.
[0027] Optionally, establish a tank detection model, including:
[0028] The tank detection model uses Keypoint RCNN and sequentially includes a residual network, a region proposal network, and a region of interest head network from the input end to the output end;
[0029] Among them, the input end is connected to the residual network to generate a feature map. The feature map is input into the region candidate network. After anchor box processing and coordinate regression processing, candidate regions are generated. The candidate regions pass through three branches of the region of interest (ROI) head network to generate detection boxes, masks, and key point heat maps respectively;
[0030] The residual network is divided into 5 feature layers: stage1, stage2, stage3, stage4, and stage5. The feature map sizes of each feature layer are different. The length and width of the feature map in the latter stage are 0.5 times those of the previous stage. Among them, the feature map of stage5 is reduced in length and width to 0.5 of the original through max pooling with a stride of 2. The feature map is input into the region candidate network to generate anchor boxes, and regression is used to modify the anchor boxes to make them close to the annotations. The basis for selecting the feature map is as follows:
[0031]
[0032] level0 = 4 is the feature layer mapped by the current anchor box, s0 = 224 is the standard image size, and area is the area of the anchor box;
[0033] The input of the region of interest (ROI) head network is the new feature map processed by the ROIAlign layer, and the classification confidence of the object and the coordinate regression value of the detection box are output through multiple fully connected layers; the cross-entropy loss is used to calculate the classification loss, denoted as Loss cls , and the Smooth L1 Loss is used to calculate the coordinate regression loss of the detection box, denoted as The formula is as follows:
[0034]
[0035] Optionally, a tank key point detection model is established, including:
[0036] Modify the key point parameters in the keypoint-rcnn network to the content of the tank image dataset. Evaluate the matching degree between the predicted values and the true values of the custom seven types of key points through the pycocotools library. Set the confidence level to select the bounding boxes, select the most suitable one from the remaining bounding boxes through NMS, and delete the bounding boxes that overlap with other parts and the candidate bounding boxes. Set the intersection threshold to 0.3 to define the overlapping degree;
[0037] The key point branch obtains an N×28×28×C feature map through 3×3 convolution and transposed convolution, where C is the number of key points. The feature map is enlarged through bilinear interpolation, and the annotation is converted into a heat map. The cross-entropy loss is calculated using the feature map and the heat map, denoted as Loss kp ;
[0038] Regarding the problem that the performance of the multi-task network structure is greatly affected by the weights of each task loss function, introduce homoscedastic uncertainty to learn the optimal weights of different task losses; define the probability model:
[0039] P(y|f W (x))=N(f W (x),σ 2 )
[0040] where f W (x) is the output of the neural network, x is the input data, W is the weight, and σ 2 is the observation noise;
[0041] The Sigmoid activation function is:
[0042] P(y|f W (x))=Softmax(f W (x))
[0043] The maximum likelihood estimation is expressed as the following formula:
[0044]
[0045] where σ is the standard deviation of the Gaussian distribution and also serves as the noise of the model;
[0046] Maximize the likelihood distribution according to W and σ; assume that y1 is the output of the regression problem, y2 is the output of the classification problem, and σ1 and σ2 are the noises of the regression problem and the classification problem respectively, then:
[0047]
[0048] where L(W, σ1, σ2) is the loss function of the multi-task model.
[0049] Optionally, perform enhancement processing on the image data, including:
[0050] Define a function with image enhancement capabilities during the training process, and randomly change the brightness, contrast, and orientation of the image during each training iteration.
[0051] This application proposes a solution to the problem of insufficient tank image data, which can efficiently generate a large number of available data sets, reduce the requirements for manual annotation, reduce the workload, and also reduce the demand for the original tank images. In the detection part, the learning and detection of the mask part are reduced by simplifying the model, the amount of computation is reduced, and the learning speed is accelerated. And key point detection is achieved by adding a key point detection branch to the above model, and the requirements for the data set format of the same model structure are unified. Description of the Drawings
[0052] Figure 1 Flow chart of the method for establishing a tank image data set and key point detection based on deep learning provided by the embodiment of the present application;
[0053] Figure 2 Schematic diagram of the tank data set construction method provided by the embodiment of the present application;
[0054] Figure 3 Schematic diagram of the keypoint R-CNN network provided by the embodiment of the present application. Specific embodiments
[0055] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0056] The following will first be introduced in combination with the embodiments of the present application with reference to the drawings.
[0057] The method provided by the present application includes: The method includes:
[0058] Input the picture to be detected into the recognition model to detect whether there is a tank. If there is a tank, output the result picture with a bounding box and key point labels; the recognition model includes a tank detection model and a tank key point detection model.
[0059] The recognition model is determined by the following method:
[0060] Step 1, establish a tank image data set, a mask, and key point annotation information.
[0061] Step 101, collect tank images and perform preprocessing. In the embodiment of the present application, the tank images include tank images in natural scenes, high-fidelity games, and film and television works.
[0062] Step 102, label the key points in the tank images; the key points include wheel pairs, gun barrels, muzzles, front armors, rear armors, side armors, and top armors.
[0063] Step 103, label the tank bounding box in the tank image to obtain a mask and generate a minimum bounding box.
[0064] Step 104, store the obtained labeling information in a json file and unify the format.
[0065] Step 2, batch generate a tank image data set and divide the tank image data set into a training set, a validation set, and a test set.
[0066] Specifically, in step 201, collect background pictures and adjust their sizes to a preset size; the preset size can be 1920x1080.
[0067] Step 202: Collect background noise images and create a mask for the background noise; the background noise is the objects in the background image other than the tank.
[0068] Step 203: Randomly select one or several background noises, perform size and rotation processing on them, and then add them to the background image.
[0069] Step 204: Randomly select several tank images and perform processing on their brightness, contrast, size, angle, and key point coordinates.
[0070] Step 205: Generate a dataset, and have the corresponding json format required for training.
[0071] Step 3: Establish a tank detection model;
[0072] The tank detection model uses Keypoint RCNN, and sequentially includes a residual network, a region proposal network, and a region of interest head network from the input end to the output end;
[0073] Among them, the input end is connected to the residual network to generate a feature map. The feature map is input into the region proposal network. After anchor box processing and coordinate regression processing, candidate regions are generated. The candidate regions pass through three branches of the region of interest head network to generate detection boxes, masks, and key point heat maps respectively.
[0074] In the embodiment of the present application, the masked part is removed and the mask-rcnn network with transfer learning is used for training, so as to obtain an image with a tank target, realizing the rapid positioning of the tank target bounding box. This method reduces the training time, reduces the computing power requirement, and improves the stability and generalization ability of the model.
[0075] The residual network is divided into 5 feature layers: stage1, stage2, stage3, stage4, and stage5. The feature map sizes of each feature layer are different. The length and width of the feature map in the latter stage are 0.5 times that of the previous stage; among them, the feature map of stage5 is downsampled by a maximum pooling with a stride of 2, and the length and width are reduced to 0.5 of the original. From this point, upsampling is performed from top to bottom. After the feature map is upsampled and the length and width are enlarged by 2 times, it is added to the feature map after 1×1 convolution in the previous stage, and then 3×3 convolution is performed; because the feature map of stage1 is not used, the above operations generate new feature maps of stage2, stage3, stage4, and stage5, as well as the feature map at the upsampling starting point, a total of 5 feature maps. The feature map is input into the region proposal network to generate anchor boxes, and regression is used to modify the anchor boxes to make them close to the annotation; the basis for selecting the feature map is:
[0076]
[0077] level0 = 4 is the feature layer mapped by the current anchor box, s0 = 224 is the standard image size, and area is the area of the anchor box. That is, when area = 224×224, the feature map of stage4 should be selected;
[0078] The input of the region of interest heads (ROI Heads) is the new feature map processed by the ROIAlign layer, and the classification confidence of the object and the regression value of the detection box coordinates are output through multiple fully connected layers; the cross-entropy loss is used to calculate the classification loss, denoted as Loss cls , and the Smooth L1 Loss is used to calculate the regression loss of the detection box coordinates, denoted as The formula is as follows:
[0079]
[0080] Step 4, establish a tank key point detection model;
[0081] Specifically, modify the key point parameters in the keypoint-rcnn network to the content of the tank image dataset, evaluate the matching degree between the predicted values and the true values of the seven types of custom key points through the pycocotools library, set the confidence level to select the bounding boxes, select the most appropriate one from the remaining bounding boxes through NMS, delete the bounding boxes that overlap with other parts and the candidate boxes, and set the intersection threshold to 0.3 to define the overlapping degree;
[0082] The key point branch obtains an N×28×28×C feature map through 3×3 convolution and transposed convolution, where C is the number of key points. The feature map is enlarged through bilinear interpolation, and the annotation is converted into a heat map. The cross-entropy loss is calculated using the feature map and the heat map, denoted as Loss kp ;
[0083] Regarding the problem that the performance of the multi-task network structure is greatly affected by the weights of each task loss function, introduce homoscedastic uncertainty to learn the optimal weights of different task losses; define the probability model:
[0084] P(y|f W (x)) = N(f W (x), σ 2 )
[0085] where f W (x) is the output of the neural network, x is the input data, W is the weight, and σ 2 is the observation noise;
[0086] The Sigmoid activation function is:
[0087] P(y|f W(x)) = Softmax(f W (x))
[0088] The maximum likelihood estimation is expressed as the following formula:
[0089]
[0090] where σ is the standard deviation of the Gaussian distribution and also the noise of the model;
[0091] Maximize the likelihood distribution according to W and σ; assume that y1 is the output of the regression problem, y2 is the output of the classification problem, and σ1 and σ2 are the noises of the regression problem and the classification problem respectively, then:
[0092]
[0093] where L(W, σ1, σ2) is the loss function of the multi-task model.
[0094] Step 5, perform enhancement processing on the image data;
[0095] Define a function with image enhancement function during the training process, and randomly change the brightness, contrast and direction of the image during each training iteration.
[0096] The function with image enhancement function provided by this application improves the accuracy of model training.
[0097] Step 6, train the tank detection model and the tank key point detection model respectively.
[0098] It should be noted that during the execution of Step 6, the function with image enhancement function in Step 5 is adopted.
[0099] This application proposes a solution to the problem of less tank image data, which can efficiently generate a large number of available data sets, reduce the requirement of manual annotation, reduce the workload, and also reduce the demand for the original tank images. In the detection part, by simplifying the model, the learning and detection of the mask part are reduced, the computation amount is reduced, and the learning speed is accelerated. And the key point detection is realized by adding a key point detection branch to the above model at the head, and the requirements for the data set format of the same model structure are unified.
[0100] Those skilled in the art can clearly understand that the technologies in the embodiments of this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of this application, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0101] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the embodiments of the service construction device and the service loading device, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.
[0102] The embodiments of this application described above do not constitute a limitation on the protection scope of this application.
Claims
1. A method for establishing a tank image dataset based on deep learning and key point detection, characterized in that, The method includes: Input the image to be detected into the recognition model to detect whether there is a tank. If there is a tank, output the result image with bounding boxes and keypoint labels; the recognition model includes a tank detection model and a tank keypoint detection model. The recognition model is determined by the following method: Step 1, establish a tank image dataset, mask and keypoint annotation information. Step 2, batch generate the tank image dataset, and divide the tank image dataset into a training set, a validation set and a test set. Step 3, establish a tank detection model. Step 4, establish a tank keypoint detection model. Step 5, perform enhancement processing on the image data. Step 6, train the tank detection model and the tank keypoint detection model respectively. Establishing a tank detection model includes: The tank detection model adopts Keypoint RCNN, and sequentially includes a residual network, a region proposal network and a region of interest head network from the input end to the output end. Among them, the input end is connected to the residual network to generate a feature map. The feature map is input into the region proposal network, and after anchor box processing and coordinate regression processing, candidate regions are generated. The candidate regions pass through three branches of the region of interest head network to generate detection boxes, masks and keypoint heatmaps respectively. The residual network is divided into 5 feature layers: stage1, stage2, stage3, stage4, stage5. The feature map sizes of each feature layer are different, and the length and width of the feature map in the latter stage are 0.5 times that of the previous stage. Among them, the feature map of stage5 is reduced in length and width to 0.5 of the original through max pooling with a stride of 2, and the feature map is input into the region proposal network to generate anchor boxes, and the anchor boxes are modified by regression to make them close to the annotation. The basis for selecting the feature map is: level0 = 4 is the feature layer mapped by the current anchor box, s0 = 224 is the standard image size, and area is the area of the anchor box. The input of the region of interest (ROI) head network is the new feature map processed by the ROIAlign layer, and the classification confidence of the object and the regression values of the detection box coordinates are output through multiple fully connected layers. The classification loss is calculated using the cross-entropy loss, denoted as Loss cls , and the SmoothL1 Loss is used to calculate the regression loss of the detection box coordinates, denoted as The formulas are as follows:
2. The method according to claim 1, characterized in that, Establishing a tank image dataset, mask and keypoint annotation information includes: Step 101, collect tank images and perform preprocessing. Step 102, mark the keypoints in the tank images; the keypoints include wheel pairs, gun barrels, muzzles, front armor, rear armor, side armor and top armor. Step 103, mark the tank bounding box in the tank image to obtain a mask and generate the minimum bounding box. Step 104, store the obtained marking information in a json file and unify the format.
3. The method according to claim 1, wherein Batch generating the tank image dataset and dividing the tank image dataset into a training set, a validation set and a test set includes: Step 201, collect background images and adjust their sizes to the preset size. Step 202, collect background noise images and make masks for the background noise; the background noise is the objects other than the tank in the background image. Step 203, randomly select one or several background noises, perform size and rotation processing, and add them to the background image. Step 204, randomly select several tank images and perform processing on brightness, contrast, size, angle and keypoint coordinates. Step 205, generate a dataset and have the corresponding json format required for training.
4. The method according to claim 1, wherein Establishing a tank keypoint detection model includes: Modify the key point parameters in the keypoint-rcnn network to the content of the tank image dataset. Evaluate the matching degree between the predicted values and the true values of the custom seven types of key points through the pycocotools library. Set the confidence level to select the bounding boxes, and select the most appropriate ones from the remaining bounding boxes through NMS. Delete the bounding boxes that overlap with the candidate boxes in other parts, and set the intersection threshold to 0.3 to define the degree of overlap; The key point branch obtains an N×28×28×C feature map through 3×3 convolution and transposed convolution, where C is the number of key points. The feature map is enlarged through bilinear interpolation, and the annotation is converted into a heat map. The cross-entropy loss is calculated using the feature map and the heat map, denoted as Loss kp ; Regarding the problem that the performance of the multi-task network structure is greatly affected by the weights of each task loss function, introduce homoscedastic uncertainty to learn the optimal weights of different task losses; define the probability model: P(y|f W (x)) = N(f W (x), σ 2 ) where f W (x) is the output of the neural network, x is the input data, W is the weight, and σ 2 is the observation noise; The Sigmoid activation function is: P(y|f W (x)) = Softmax(f W (x)) The maximum likelihood estimation is expressed as the following formula: Among them, σ is the standard deviation of the Gaussian distribution and also serves as the noise of the model; Maximize the likelihood distribution according to W and σ; assume that y1 is the output of the regression problem and y2 is the output of the classification problem, and σ1 and σ2 are the noises of the regression problem and the classification problem respectively, then: Among them, L(W, σ1, σ2) is the loss function of the multi-task model.
5. The method according to claim 1, wherein Perform enhancement processing on the image data, including: Define a function with image enhancement function during the training process, and randomly change the brightness, contrast and direction of the image during each training iteration.
Citation Information
Patent Citations
Image recognition method, electronic equipment and computer readable storage medium
CN113343010A
Active learning for inspection tool
US20220262104A1