Aircraft target image feature point detection method and device
Through single-stage detection method and semi-supervised learning based on deep neural networks, the problems of low accuracy and low efficiency of aircraft feature point detection in traditional methods are solved, and efficient and accurate feature point detection is achieved.
Patent Information
- Application Number
- CN202510076399.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional aircraft flight test data processing methods have problems with low detection accuracy and high error detection rate, and cannot accurately extract aircraft feature points and face inefficiency when image data information is rich.
A single-stage detection method based on deep neural network is adopted to train the aircraft target image feature point detection model through semi-supervised learning, and the network structure is optimized using lightweight modules to reduce computing and storage requirements, and improve detection efficiency.
It realizes more accurate extraction of feature points of aircraft target images, improves detection efficiency, reduces labor labeling costs, and has real-time detection speed.
Smart Images

Figure CN120107652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and in particular to a method and device for detecting feature points of an aircraft target image. Background Art
[0002] By interpreting the images taken by multiple optical devices, the flight trajectory of the aircraft can be obtained by intersection and solution, which is an important basis for aircraft flight test identification and fault analysis. In traditional aircraft flight test data processing, manual interpretation, template matching, optical flow tracking, edge contour extraction and other methods are mainly used to interpret optical images. Affected by the flight status of the aircraft, complex environmental factors, and the status of the optical equipment, the traditional processing methods have the problems of low detection accuracy and high false detection rate, and are unable to accurately extract the feature points of the aircraft. In addition, in the face of increasingly rich image data information, the inefficiency of traditional processing methods has gradually become apparent. Therefore, how to accurately and efficiently extract the feature points of the aircraft is a difficult problem that needs to be solved.
[0003] In summary, the prior art has the following problem: how to accurately and efficiently extract aircraft feature points. Summary of the invention
[0004] The purpose of the present invention is to solve the problem of how to accurately and efficiently extract aircraft feature points.
[0005] To this end, on one hand, an embodiment of the present invention provides a method for detecting feature points of an aircraft target image, the method comprising the following steps:
[0006] Collecting aircraft target image data and preprocessing the collected aircraft target image data;
[0007] Using the preprocessed aircraft target image data to construct an aircraft target image dataset;
[0008] Annotating image data in the aircraft target image data set;
[0009] Use the annotated image data to build an aircraft target image feature point detection model;
[0010] Using a semi-supervised learning method to train the aircraft target image feature point detection model;
[0011] The aircraft target image feature point detection model is used to detect the target image feature points of the aircraft.
[0012] On the other hand, an embodiment of the present invention further provides an aircraft target image feature point detection device, comprising:
[0013] A processing unit, used for collecting aircraft target image data and preprocessing the collected aircraft target image data;
[0014] A collection unit, used for constructing an aircraft target image data set using the preprocessed aircraft target image data;
[0015] a labeling unit, used for labeling the image data in the aircraft target image data set;
[0016] A construction unit, used to construct an aircraft target image feature point detection model using the annotated image data;
[0017] A training unit, used for training the aircraft target image feature point detection model by using a semi-supervised learning method;
[0018] The detection unit is used to detect the target image feature points of the aircraft using the aircraft target image feature point detection model.
[0019] The above technical solution has the following beneficial effects: the present invention adopts a single-stage detection method based on a deep neural network, which can better extract the shallow positioning information features and deep semantic information features of the target image; adopts an end-to-end detection framework with real-time detection speed, and uses lightweight modules to optimize the network model structure, reducing the model calculation amount and storage requirements, further improving the detection efficiency while maintaining the accuracy, and can improve the model performance through a large amount of unlabeled data and reduce the cost of manual labeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flow chart of a method for detecting feature points of an aircraft target image provided by an embodiment of the present invention;
[0021] Figure 2 It is a structural schematic diagram of an aircraft target image feature point detection device provided by an embodiment of the present invention;
[0022] Figure 3 is a schematic diagram of the structure of the CheapConv module provided in an embodiment of the present invention;
[0023] Figure 4 is a schematic diagram of the structure of a CheapBottleNeck module provided in an embodiment of the present invention;
[0024] Figure 5 is a schematic diagram of the structure of the CheapModule module provided in an embodiment of the present invention;
[0025] Figure 6 is a schematic diagram of the structure of a backbone network provided by an embodiment of the present invention;
[0026] Figure 7 is a schematic diagram of the structure of the neck network architecture provided by an embodiment of the present invention;
[0027] Figure 8 It is a structural diagram of the prediction output end network architecture provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] In the embodiment of the present invention, Figure 1 , provides a method for detecting feature points of an aircraft target image, the method comprising the following steps:
[0030] S101: collecting aircraft target image data, and preprocessing the collected aircraft target image data;
[0031] Use optical equipment to collect aircraft target image data; the specific steps of preprocessing include:
[0032] Filter out high-quality aircraft target image data under different environmental conditions, different aircraft distances, and different aircraft attitudes;
[0033] First, the selected aircraft target image data is scaled from image size a×b to a′×b′ according to the scale γ, where a′=γa≤a o , b′=γb≤b o ,γ=min(a o / a,b o / b), a o ×b o Represents the standardized image size, and then fills the part that is less than the standardized image size with the background color.
[0034] First, count the histogram of the aircraft target image data, set the threshold T, and set the grayscale values at 1-T and T to p min and p max , and then use bilinear contrast stretching to standardize the gray value p of the image, and standardize the pixel gray value of the image data to between 0 and 1. The specific calculation expression is:
[0035] p o =(pp min ) / (p max -p min );
[0036] S102: constructing an aircraft target image data set using the preprocessed aircraft target image data;
[0037] Randomly divide the preprocessed image data into two parts: those that need to be labeled and those that do not need to be labeled;
[0038] The image data of the parts that need to be labeled and the parts that do not need to be labeled are divided into training set, verification set and test set according to a certain ratio.
[0039] S103: labeling the image data in the aircraft target image data set; including:
[0040] Marking an aircraft target on the image data to create an aircraft label;
[0041] Select feature points on the aircraft target image for annotation and create aircraft feature point labels.
[0042] The specific steps of marking include:
[0043] Use a rectangular box to mark the position of the aircraft on the image data. The rectangular box just surrounds the outline of the aircraft and creates an aircraft label.
[0044] Use annotation points within the rectangular frame to mark the positions of the aircraft feature points on the image data and create aircraft feature point labels.
[0045] S104: constructing an aircraft target image feature point detection model using the annotated image data; including:
[0046] Construct the basic module of the aircraft target image feature point detection model;
[0047] Construct the data input network architecture of the aircraft target image feature point detection model;
[0048] Construct the backbone network architecture of the aircraft target image feature point detection model;
[0049] Construct the neck network architecture of the aircraft target image feature point detection model;
[0050] Construct the prediction output network architecture of the aircraft target image feature point detection model;
[0051] Construct the loss function of the aircraft target image feature point detection model.
[0052] The basic modules of the aircraft target image feature point detection model include: CheapConv, CheapBottleNeck and CheapModule;
[0053] The specific structures are:
[0054] The specific structure of CheapConv (low-cost convolutional layer) is as follows Figure 3 As shown in the figure, first, the input feature map is subjected to a standard Conv (convolution layer) convolution operation, and then divided into two processes, the left process is subjected to a standard Conv convolution operation with a convolution kernel of 5×5 and a step size of 1, and the right process is subjected to a Linear (linear change) linear operation; finally, the features generated by the left and right processes are concat (joined) in the channel dimension, and the joined feature map is output.
[0055] The specific structure of CheapBottleNeck (low consumption bottleneck) is as follows Figure 4 As shown in the figure, the processing of the input feature map is first divided into two processes, the right process does not perform any processing; then the left process is sequentially subjected to CheapConv operation, BatchNorm batch normalization operation, and then Activation activation operation, a DWConv (depthwise separable convolution) operation with a convolution kernel of 3×3 and a step size of 2, a BatchNorm (batch normalization layer) batch normalization operation, a CheapConv operation, and a BatchNorm batch normalization operation; finally, the features generated by the left and right processes are added together to output the feature map.
[0056] The specific structure of CheapModule (low consumption module) is as follows Figure 5 As shown in the figure, the processing of the input feature map is first divided into two processes, the right process does not perform any processing; then the left process is sequentially subjected to CheapConv operation, BatchNorm batch normalization operation, Activation activation operation, CheapBottleNeck operation, and CheapConv operation; then the features generated by the left and right processes are concat-joined in the channel dimension; finally, the CheapConv operation, BatchNorm batch normalization operation, and then the activation operation is performed in the Activation (activation layer) to output the feature map.
[0057] The detection model structure includes the data input network architecture, the backbone network architecture, the neck network architecture, and the prediction output network architecture. The specific structures of each part are:
[0058] Data input network architecture: The specific structure is a standard convolution Conv (convolution layer) operation with a convolution kernel size of 6×6, a step size of 2, and a padding of 2.
[0059] Backbone network architecture: The specific structure is as follows Figure 6As shown in the figure, first, the input feature map is subjected to CheapConv operation (sequence number 1) and 3 CheapModule operations (sequence number 2); then, CheapConv operation (sequence number 3) and 6 CheapModule operations (sequence number 4), CheapConv operation (sequence number 5) and 9 CheapModule operations (sequence number 6), CheapConv operation (sequence number 7) and 3 CheapModule operations (sequence number 8) are performed in sequence; finally, SPPF spatial pyramid pooling operation (sequence number 9) is performed.
[0060] Neck network architecture: The specific structure is as follows Figure 7 As shown in the figure, firstly, the input feature map is subjected to CheapConv operation (serial number 10), then Upsample is performed (serial number 11), and the features output by layer 6 are concatenated in the channel dimension (serial number 12), CheapModule operation (serial number 13), CheapConv operation (serial number 14), Upsample is performed (serial number 15), and the features output by layer 4 are concatenated in the channel dimension (serial number 16), and CheapModule operation (serial number 17) is performed in sequence to output the features; then The features output from the 17th layer are sequentially subjected to three CheapConv operations with a step size of 2 (sequence number 18), and are concatenated with the features output from the 14th layer in the channel dimension (sequence number 19), and a CheapModule operation (sequence number 20) is performed to output the features; finally, the features output from the 20th layer are sequentially subjected to three CheapConv operations with a step size of 2 (sequence number 21), and are concatenated with the features output from the 10th layer in the channel dimension (sequence number 22), and a CheapModule operation (sequence number 23) is performed to output a feature map.
[0061] Prediction output network architecture: The specific structure is as follows Figure 8 As shown in the figure, first, a transposed convolution operation with a convolution kernel of 1×1 is performed on the input feature map, and then it is divided into two processes: target detection and feature point detection. The target detection process first passes through a standard convolution with a convolution kernel of 3×3, and then is divided into two processes, both of which perform a transposed convolution operation with a convolution kernel of 1×1, and output the target detection classification result and positioning result respectively; the feature point detection process first passes through a standard convolution with a convolution kernel of 3×3, and then is divided into two processes, both of which perform a transposed convolution operation with a convolution kernel of 1×1, and output the feature point detection positioning result and confidence respectively.
[0062] S105: Using a semi-supervised learning method to train the aircraft target image feature point detection model; the specific steps are:
[0063] Step S610: Building a training environment;
[0064] Step S620: setting training parameters;
[0065] Step S630: Initialize model parameters;
[0066] Step S640: using semi-supervised learning to train the model;
[0067] Step S650: verifying the model detection effect;
[0068] The specific steps of the initial network model parameters described in step S630 are to first use the labeled image data in the labeled aircraft target image data set to train the aircraft target image feature point detection model to obtain the initial network model parameters, and then use the parameters to initialize and assign values to the aircraft target image feature point detection teacher model and the aircraft target image feature point detection teacher model.
[0069] The semi-supervised learning training model process described in step S640 specifically includes:
[0070] Step S641: In each training round, a batch of labeled image data and unlabeled image data are randomly sampled from the aircraft target image dataset at a ratio of 1:10 each time;
[0071] Step S642: For a batch of labeled image data described in step S641, use the aircraft target image feature point detection student model to perform detection, and calculate the corresponding loss score L according to the labeled information label ;
[0072] Step S643: For a batch of unlabeled image data described in step S641, first use the aircraft target image feature point detection teacher model to generate pseudo labels, then use the aircraft target image feature point detection student model for detection, and calculate the corresponding loss score L according to the pseudo labels unlabel ;
[0073] Step S644: Calculate the comprehensive loss score L of the semi-supervised training according to the semi-supervised training loss function described in step S560, and update the student model parameters for the aircraft target image feature point detection;
[0074] Step S645: after each training round, the teacher model parameters are updated using the EMA exponential moving average method;
[0075] Step S646: Stop training when the number of training rounds reaches the maximum number of training rounds.
[0076] S106: Detecting the target image feature points of the aircraft using the aircraft target image feature point detection model.
[0077] The preprocessing of the collected aircraft target image data includes:
[0078] Filter aircraft target image data based on environmental conditions, aircraft distance, and aircraft attitude;
[0079] The selected aircraft target image data is scaled from image size a×b to a′×b′ according to scale γ, where a′=γa≤a o , b′=γb≤b o ,γ=min(a o / a,b o / b), a o ×b o Represents the standardized image size, and fills the part that is less than the standardized image size with the background color;
[0080] Count the histogram of the aircraft target image data, set the threshold to T, and set the gray value at 1-T to p min , set the gray value at T to p max , bilinear contrast stretching is used to standardize the gray value p of the image, and the pixel gray value of the image data is standardized to between 0 and 1. The specific standardization process is:
[0081] p o =(pp min ) / (p max -p min ).
[0082] The aircraft target image feature point detection model is as follows:
[0083] L=L label +λ unlabel L unlabel ,
[0084] Among them, L label represents the loss score using labeled image data; L unlabel represents the loss score using unlabeled image data; λ unlabel Represents the balance coefficient of the loss score for unlabeled image data.
[0085] Where L label , L unlabel The specific expression is:
[0086]
[0087] Where (h, w) represents the pixel position range on the aircraft target image, L cls , They represent the classification loss scores using labeled image data and unlabeled image data respectively. The specific expressions are:
[0088] Lcls =CrossEntroyLoss(X cls ,Y cls ),
[0089]
[0090] Where X cls Represents the aircraft classification detection result of the student model, Y cls Represents the image data label category, Represents the aircraft classification detection results of the teaching model.
[0091] L box , They represent the rectangular box positioning loss scores using labeled image data and unlabeled image data, respectively. The specific expressions are:
[0092] L box =1-CIoU(X box ,Y box ),
[0093]
[0094] Where X box Represents the aircraft target detection and positioning results of the student model, Y box Represents the image data aircraft target label, Represents the aircraft target detection and positioning results of the teacher model.
[0095] L kpt , They represent the feature point positioning loss scores using labeled image data and unlabeled image data respectively. The specific expressions are:
[0096] L kpt =||X kpt -Y kpt || 2
[0097]
[0098] Where X kpt Represents the detection and positioning results of the aircraft feature points of the student model, Y kpt Represents the image data aircraft feature point label, The detection and positioning results of aircraft feature points representing the teacher model.
[0099] L kpt_conf , They represent the confidence loss scores of feature points using labeled image data and unlabeled image data respectively. The specific expressions are:
[0100] L kpt_conf =BCELoss(X kpt_conf ,Y kpt_conf ),
[0101]
[0102] Where X kpt_conf Represents the confidence detection result of the aircraft target image feature points of the student model, Y kpt_conf Represents the confidence labels of feature points of aircraft target image data, The confidence detection results of the feature points of the aircraft target image representing the teacher model.
[0103] The present invention also provides a device for detecting feature points of an aircraft target image. Figure 2 As shown, including:
[0104] The processing unit 21 is used to collect the aircraft target image data and pre-process the collected aircraft target image data;
[0105] A collection unit 22, used to construct an aircraft target image data set using the preprocessed aircraft target image data;
[0106] A labeling unit 23, used for labeling the image data in the aircraft target image data set;
[0107] A construction unit 24 is used to construct an aircraft target image feature point detection model using the annotated image data;
[0108] A training unit 25, used for training the aircraft target image feature point detection model using a semi-supervised learning method;
[0109] The detection unit 26 is used to detect the target image feature points of the aircraft using the aircraft target image feature point detection model.
[0110] The processing unit 21 includes:
[0111] A screening module is used to screen the aircraft target image data according to environmental conditions, aircraft distance, and aircraft attitude;
[0112] The scaling module is used to scale the image data of the selected aircraft target from a×b to a′×b′ according to the scale γ, where a′=γa≤a o , b′=γb≤b o ,γ=min(a o / a,b o / b), a o ×b oRepresents the standardized image size, and fills the part that is less than the standardized image size with the background color;
[0113] The statistical module is used to count the histogram of the aircraft target image data. The threshold is set to T, and the gray value at 1-T is set to p min , set the gray value at T to p max , bilinear contrast stretching is used to standardize the gray value p of the image, and the pixel gray value of the image data is standardized to between 0 and 1. The specific standardization process is:
[0114] p o =(pp min ) / (p max -p min ).
[0115] Annotation unit, including:
[0116] Used to mark the aircraft target on the image data and establish an aircraft label;
[0117] It is used to select feature points on the aircraft target image for annotation and create aircraft feature point labels.
[0118] Building blocks, including:
[0119] Basic modules for building a model for detecting feature points in aircraft target images;
[0120] The data input network architecture used to build the aircraft target image feature point detection model;
[0121] Backbone network architecture for building a model for detecting feature points in aircraft target images;
[0122] Neck network architecture for building a model for detecting feature points in aircraft target images;
[0123] The prediction output network architecture used to build the aircraft target image feature point detection model;
[0124] Loss function used to build a model for detecting feature points in aircraft target images.
[0125] The aircraft target image feature point detection model is as follows:
[0126] L=L label +λ unlabel L unlabel ,
[0127] Among them, L label represents the loss score using labeled image data; L unlabel represents the loss score using unlabeled image data; λ unlabelRepresents the balance coefficient of the loss score for unlabeled image data.
[0128] Where L label , L unlabel The specific expression is:
[0129]
[0130] Where (h, w) represents the pixel position range on the aircraft target image, L cls , They represent the classification loss scores using labeled image data and unlabeled image data respectively. The specific expressions are:
[0131] L cls =CrossEntroyLoss(X cls ,Y cls ),
[0132]
[0133] Where X cls Represents the aircraft classification detection result of the student model, Y cls Represents the image data label category, Represents the aircraft classification detection results of the teaching model.
[0134] L box , They represent the rectangular box positioning loss scores using labeled image data and unlabeled image data, respectively. The specific expressions are:
[0135] L box =1-CIoU(X box ,Y box ),
[0136]
[0137] Where X box Represents the aircraft target detection and positioning results of the student model, Y box Represents the image data aircraft target label, Represents the aircraft target detection and positioning results of the teacher model.
[0138] L kpt , They represent the feature point positioning loss scores using labeled image data and unlabeled image data respectively. The specific expressions are:
[0139] L kpt =||X kpt -Y kpt || 2
[0140]
[0141] Where X kpt Represents the detection and positioning results of the aircraft feature points of the student model, Y kpt Represents the image data aircraft feature point label, The detection and positioning results of aircraft feature points representing the teacher model.
[0142] L kpt_conf , They represent the confidence loss scores of feature points using labeled image data and unlabeled image data respectively. The specific expressions are:
[0143] L kpt_conf =BCELoss(X kpt_conf ,Y kpt_conf ),
[0144]
[0145] Where X kpt_conf Represents the confidence detection result of the aircraft target image feature points of the student model, Y kpt_conf Represents the confidence labels of feature points of aircraft target image data, The confidence detection results of the feature points of the aircraft target image representing the teacher model.
[0146] The working method and principle of the device for detecting feature points of an aircraft target image have been described in detail in the embodiment of a method for detecting feature points of an aircraft target image, so they will not be repeated here.
[0147] The present invention adopts a single-stage detection method based on deep neural network, which can better extract the shallow positioning information features and deep semantic information features of the target image; adopts an end-to-end detection framework with real-time detection speed, uses lightweight modules to optimize the network model structure, reduces the model calculation amount and storage requirements, and further improves the detection efficiency while maintaining the accuracy. It can also improve the model performance through a large amount of unlabeled data and reduce the cost of manual annotation.
[0148] The above technical solution of the embodiment of the present invention is described in detail below in conjunction with specific application examples. For technical details not introduced during the implementation process, please refer to the relevant description in the previous text.
[0149] Embodiment 1:
[0150] The present invention provides a method for detecting feature points of an aircraft target image, and the specific steps include:
[0151] Step S100: using optical equipment to collect aircraft target image data;
[0152] Step S200: pre-processing the collected aircraft target image data;
[0153] Step S300: constructing an aircraft target image data set using the preprocessed aircraft target image data;
[0154] Step S400: labeling the image data in the aircraft target image data set;
[0155] Step S500: constructing an aircraft target image feature point detection model;
[0156] Step S600: using a semi-supervised learning method to train an aircraft target image feature point detection model.
[0157] Furthermore, in step S100, this embodiment uses multiple optoelectronic theodolites arranged in different positions to collect the entire video of the aircraft flight test, and transmits the video to a local database. The video sampling frequency is 100 Hz, and the video is converted into a sequence of images with an image bit number of 8 bits, a channel number of 3, and a format of bmp.
[0158] Furthermore, the specific steps of the preprocessing in step S200 of this embodiment include:
[0159] Step S210: Filter out typical image data under different environmental conditions, different aircraft distances, and different aircraft postures, and extract good quality images from the video data segment at 1 frame / s;
[0160] Step S220: first, the image size a×b selected from each photoelectric theodolite is proportionally scaled to a′×b′ according to the scale γ=640 / min(a,b), where a′=γa, b′=γb, and then the part less than 640pixel×640pixel is filled with the background color. Finally, the image data size is standardized to 640pixel×640pixel.
[0161] Step S230: First, count the histogram of the image, set the threshold T, and set the grayscale values at 1-T and T to p min and p max Preferably, the grayscale values at 5% and 95% are set to p min and p max , and then use bilinear contrast stretching to standardize the gray value p of the image, p = (pp min ) / (p max -p min ), where p 1 =p min , p 2 =p max, and finally the grayscale value of the image data is standardized to between 0 and 1. The specific calculation expression is:
[0162] p o =(pp min ) / (p max -p min );
[0163] Furthermore, the specific steps of constructing the aircraft target image dataset in step S300 of this embodiment include:
[0164] Step S310: Establishing a data set file system;
[0165] Step S320: randomly divide the pre-processed image data into two parts of image data that need to be labeled and that do not need to be labeled according to a ratio of 2:8, and store them in the images folder in the LabelData folder and the UnlabelData folder respectively;
[0166] Step S330: randomly divide the image data that need to be labeled and the image data that do not need to be labeled into a training set, a validation set, and a test set in a ratio of 6:3:1, and distribute the image data in the images folder to the train, valid, and test folders accordingly;
[0167] Furthermore, the specific steps of establishing the data set file system in step S310 of this embodiment include:
[0168] Step S311: Create a main folder for image data and name it ImageData;
[0169] Step S312: Create two subfolders under the ImageData folder, named LabelData and UnlabelData respectively;
[0170] Step S313: Create two subfolders in the LabelData folder, named images and labels respectively, and create train, valid and test subfolders in the images folder;
[0171] Step S314: Create an images subfolder under the LabelData folder, and create train, valid, and test subfolders under the images folder.
[0172] Furthermore, the specific steps of marking in step S400 of this embodiment include:
[0173] Step S410: Use the rectangular box annotation tool in the annotation software to select the upper left and lower right corners of the rectangular box to frame the outline of the aircraft image and create an aircraft label;
[0174] Step S420: Use the annotation point annotation tool in the annotation software to select feature points on the aircraft image for annotation, and create aircraft feature point labels.
[0175] Furthermore, in step S500 of this embodiment, the specific steps of constructing the aircraft target image feature point detection model include:
[0176] Step S510: constructing a basic module of an aircraft target image feature point detection model;
[0177] Step S520: constructing a data input end network architecture of an aircraft target image feature point detection model;
[0178] Step S530: constructing a backbone network architecture of an aircraft target image feature point detection model;
[0179] Step S540: constructing a neck network architecture of an aircraft target image feature point detection model;
[0180] Step S550: constructing a prediction output end network architecture of an aircraft target image feature point detection model;
[0181] Step S560: constructing a loss function of an aircraft target image feature point detection model;
[0182] Furthermore, the specific steps of constructing the basic module of the aircraft target image feature point detection model in step S510 of this embodiment include:
[0183] Step S511: Construct the CheapConv module. The specific structure is as follows: Figure 3 As shown in the figure, first, the input feature map is subjected to a standard Conv convolution operation, which is then divided into two processes, the left process is subjected to a standard Conv convolution operation with a convolution kernel of 5×5 and a step size of 1, and the right process is subjected to a Linear linear operation; finally, the features generated by the left and right processes are concat-joined in the channel dimension, and the concatenated feature map is output.
[0184] Step S512: Construct the CheapBottleNeck module. The specific structure is as follows: Figure 4As shown in the figure, the processing of the input feature map is first divided into two processes, the right process does not perform any processing; then the left process is sequentially subjected to CheapConv operation, BatchNorm batch normalization operation, Activation activation operation, DWConv depth separable convolution operation with a convolution kernel of 3×3 and a step size of 2, BatchNorm batch normalization operation, CheapConv operation, BatchNorm batch normalization operation; finally, the features generated by the left and right processes are added together to output the feature map.
[0185] Step S513: Construct the CheapModule module. The specific structure is as follows: Figure 5 As shown in the figure, the processing of the input feature map is first divided into two processes, the right process does not perform any processing; then the left process is sequentially subjected to CheapConv operation, BatchNorm batch normalization operation, Activation activation operation, CheapBottleNeck operation, and CheapConv operation; then the features generated by the left and right processes are concat-joined in the channel dimension; finally, CheapConv operation, BatchNorm batch normalization operation, Activation activation operation are sequentially performed to output the feature map.
[0186] Furthermore, the specific structure of the data input network architecture for constructing the aircraft target image feature point detection model described in step S520 of this embodiment is a standard convolution Conv operation with a convolution kernel scale of 6×6, a step size of 2, and a padding of 2.
[0187] Further, the specific structure of the backbone network architecture for constructing the aircraft target image feature point detection model in step S530 of this embodiment is as follows: Figure 6 As shown in the figure, firstly, the input feature map is subjected to CheapConv operation (sequence number 1) and 3 CheapModule operations (sequence number 2); then, CheapConv operation (sequence number 3) and 6 CheapModule operations (sequence number 4), CheapConv operation (sequence number 5) and 9 CheapModule operations (sequence number 6), CheapConv operation (sequence number 7) and 3 CheapModule operations (sequence number 8) are sequentially performed; finally, SPPF spatial pyramid pooling operation (sequence number 9) is performed;
[0188] Further, the specific structure of the neck network architecture for constructing the aircraft target image feature point detection model in step S540 of this embodiment is as follows: Figure 7As shown, first, the input feature map is subjected to CheapConv operation (sequence number 10), Upsample upsampling (sequence number 11), and the features output by the sequence number 6 layer are spliced in the channel dimension (sequence number 12), CheapModule operation (sequence number 13), CheapConv operation (sequence number 14), Upsample upsampling (sequence number 15), and the features output by the sequence number 4 layer are spliced in the channel dimension (sequence number 16), and CheapModule operation (sequence number 17) is performed in sequence to output the features; then, from the sequence number The features output from the 17th layer are sequentially subjected to 3 CheapConv operations with a step size of 2 (serial number 18), and are concatenated with the features output from the 14th layer in the channel dimension (serial number 19), and CheapModule operations (serial number 20) are performed to output the features; finally, the features output from the 20th layer are sequentially subjected to 3 CheapConv operations with a step size of 2 (serial number 21), and are concatenated with the features output from the 10th layer in the channel dimension (serial number 22), and CheapModule operations (serial number 23) are performed to output the feature map;
[0189] Further, the specific structure of the prediction output end network architecture of the aircraft target image feature point detection model constructed in step S550 of this embodiment is as follows: Figure 8 As shown in the figure, the input feature map is first subjected to a transposed convolution operation with a convolution kernel of 1×1, and then divided into two processes: the target detection process first passes through a standard convolution with a convolution kernel of 3×3, and then is divided into two processes, both of which are subjected to a transposed convolution operation with a convolution kernel of 1×1, and the target detection classification result and the positioning result are output respectively; the feature point detection process first passes through a standard convolution with a convolution kernel of 3×3, and then is divided into two processes, both of which are subjected to a transposed convolution operation with a convolution kernel of 1×1, and the feature point detection positioning result and the confidence are output respectively;
[0190] Further, the semi-supervised training loss function expression of the aircraft image target detection model in step S560 is:
[0191] L=L label +λ unlabel L unlabel
[0192] Where L label , L unlabel They represent the loss scores of using labeled image data and using unlabeled image data respectively. The specific expressions are:
[0193]
[0194] Where L cls , They represent the classification loss scores using labeled image data and unlabeled image data respectively. The specific expressions are:
[0195] L cls =CrossEntroyLoss(X cls ,Y cls )
[0196]
[0197] Where X cls Represents the aircraft classification detection result of the student model, Y cls Represents the image data label category, Represents the aircraft classification detection results of the teaching model.
[0198] L box , They represent the rectangular box positioning loss scores using labeled image data and unlabeled image data, respectively. The specific expressions are:
[0199] L box =1-CIoU(X box ,Y box )
[0200]
[0201] Where X box Represents the aircraft target detection and positioning results of the student model, Y box Represents the image data aircraft target label, Represents the aircraft target detection and positioning results of the teacher model.
[0202] L kpt , They represent the feature point positioning loss scores using labeled image data and unlabeled image data respectively. The specific expressions are:
[0203] L kpt =||X kpt -Y kpt || 2
[0204]
[0205] Where X kpt Represents the detection and positioning results of the aircraft feature points of the student model, Y kpt Represents the image data aircraft feature point label, The detection and positioning results of aircraft feature points representing the teacher model.
[0206] L kpt_conf , They represent the confidence loss scores of feature points using labeled image data and unlabeled image data respectively. The specific expressions are:
[0207] L kpt_conf =BCELoss(X kpt_conf ,Y kpt_conf )
[0208]
[0209] Where X kpt_conf Represents the confidence detection result of the aircraft target image feature points of the student model, Y kpt_conf Represents the confidence labels of feature points of aircraft target image data, The confidence detection results of the feature points of the aircraft target image representing the teacher model.
[0210] Furthermore, the specific steps of the semi-supervised learning training method described in step S600 of this embodiment are:
[0211] Step S610: Building a training environment;
[0212] Step S620: setting training parameters;
[0213] Step S630: Initialize model parameters;
[0214] Step S640: using semi-supervised learning to train the model;
[0215] Furthermore, the training environment described in step S610 of this embodiment specifically includes computer hardware using Intel(R) Core(TM) i9-9900K CPU@3.60GHz model CPU, RTX 3090 model GPU, and 64GB memory, and computer software using the Pytorch1.13 deep learning framework based on Anaconda configuration.
[0216] Furthermore, the training parameters described in step S620 of this embodiment specifically include a maximum training iteration round of 300, a batch size of 16, and an initial learning rate of 0.01.
[0217] Furthermore, the initial network model parameters described in step S630 of this embodiment specifically include the following steps: firstly using the labeled image data in the labeled aircraft target image data set to train the aircraft target image feature point detection model to obtain initial network model parameters, and then using the parameters to initialize and assign values to the aircraft target image feature point detection teacher model and the aircraft target image feature point detection teacher model.
[0218] Furthermore, the semi-supervised learning training model process described in step S640 of this embodiment specifically includes:
[0219] Step S641: In each training round, a batch of labeled image data and unlabeled image data are randomly sampled from the aircraft target image dataset at a ratio of 1:10 each time;
[0220] Step S642: For a batch of labeled image data described in step S641, use the aircraft target image feature point detection student model to perform detection, and calculate the corresponding loss score L according to the labeled information label ;
[0221] Step S643: For a batch of unlabeled image data described in step S641, first use the aircraft target image feature point detection teacher model to generate pseudo labels, then use the aircraft target image feature point detection student model for detection, and calculate the corresponding loss score L according to the pseudo labels unlabel ;
[0222] Step S644: Calculate the comprehensive loss score L of the semi-supervised training according to the semi-supervised training loss function described in step S560, and update the student model parameters for the aircraft target image feature point detection;
[0223] Step S645: after each training round, the teacher model parameters are updated using the EMA exponential moving average method;
[0224] Step S646: Stop training when the number of training rounds reaches the maximum number of training rounds.
[0225] The present invention adopts a single-stage detection method based on deep neural network, which can better extract the shallow positioning information features and deep semantic information features of the target image; adopts an end-to-end detection framework with real-time detection speed, uses lightweight modules to optimize the network model structure, reduces the model calculation amount and storage requirements, and further improves the detection efficiency while maintaining the accuracy. It can also improve the model performance through a large amount of unlabeled data and reduce the cost of manual annotation.
[0226] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of protection of the present disclosure. The attached method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.
[0227] In the above detailed description, various features are grouped together in a single embodiment to simplify the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the embodiments of the claimed subject matter require more features than are clearly stated in each claim. On the contrary, as reflected in the appended claims, the invention is in a state of having less than all the features of the disclosed individual embodiments. Therefore, the appended claims are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate preferred embodiment of the invention.
[0228] The disclosed embodiments are described above to enable any person skilled in the art to implement or use the present invention. Various modifications of these embodiments are obvious to those skilled in the art, and the general principles defined herein may also be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.
[0229] The above description includes examples of one or more embodiments. Of course, it is impossible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but it should be recognized by those skilled in the art that the various embodiments may be further combined and arranged. Therefore, the embodiments described herein are intended to cover all such changes, modifications and variations that fall within the scope of protection of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, the word is covered in a manner similar to the term "including", just as "including," is explained as a transitional word in the claims. In addition, any term "or" used in the specification of the claims is intended to mean "non-exclusive or".
[0230] Those skilled in the art may also understand that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention may be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly demonstrate the interchangeability of hardware and software, the various illustrative components, units, and steps described above have generally described their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be understood as exceeding the scope of protection of the embodiments of the present invention.
[0231] The various illustrative logic blocks or units described in the embodiments of the present invention can be implemented or operated by a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic, a discrete hardware component, or any combination of the above. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.
[0232] The steps of the method or algorithm described in the embodiments of the present invention can be directly embedded in hardware, a software module executed by a processor, or a combination of the two. The software module can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or other storage media of any form in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and can write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be arranged in an ASIC, and the ASIC can be arranged in a user terminal. Optionally, the processor and the storage medium can also be arranged in different components in the user terminal.
[0233] In one or more exemplary designs, the above functions described in the embodiments of the present invention can be implemented in hardware, software, firmware or any combination of the three. If implemented in software, these functions can be stored on a computer-readable medium, or transmitted in the form of one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media that facilitate the transfer of computer programs from one place to another. The storage medium can be any available medium that can be accessed by any general or special computer. For example, such computer-readable media can include but are not limited to RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store program codes in the form of instructions or data structures and other forms that can be read by general or special computers, or general or special processors. In addition, any connection can be appropriately defined as a computer-readable medium, for example, if the software is transmitted from a website site, server or other remote resource through a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wirelessly, such as infrared, wireless and microwave, it is also included in the defined computer-readable medium. The disk and disc include compact disk, laser disk, optical disk, DVD, floppy disk and blue-ray disk. Disks usually copy data magnetically, while discs usually copy data optically with lasers. The above combination can also be included in computer readable media.
[0234] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for detecting feature points of an aircraft target image, characterized in that: The method comprises the following steps: Collecting aircraft target image data and preprocessing the collected aircraft target image data; Using the preprocessed aircraft target image data to construct an aircraft target image dataset; Annotating image data in the aircraft target image data set; Use the annotated image data to build an aircraft target image feature point detection model; Using a semi-supervised learning method to train the aircraft target image feature point detection model; The aircraft target image feature point detection model is used to detect the target image feature points of the aircraft.
2. The method for detecting feature points of an aircraft target image according to claim 1, characterized in that: The preprocessing of the collected aircraft target image data includes: Filter aircraft target image data based on environmental conditions, aircraft distance, and aircraft attitude; The selected aircraft target image data is scaled from image size a×b to a′×b′ according to the scale γ, where a′=γa≤a o , b′=γb≤b o ,γ=min(a o / a,b o / b), a o ×b o Represents the standardized image size, and fills the part that is less than the standardized image size with the background color; Count the histogram of the aircraft target image data, set the threshold to T, and set the gray value at 1-T to p min , set the gray value at T to p max , bilinear contrast stretching is used to standardize the gray value p of the image, and the pixel gray value of the image data is standardized to between 0 and 1. The specific standardization process is: p o =(p-p min ) / (p max -p min )。 3. The method for detecting feature points of an aircraft target image according to claim 2, characterized in that: The labeling of the image data in the aircraft target image data set includes: Marking an aircraft target on the image data to create an aircraft label; Select feature points on the aircraft target image for annotation and create aircraft feature point labels.
4. The method for detecting feature points of an aircraft target image according to claim 3, characterized in that: The method of constructing an aircraft target image feature point detection model using the annotated image data includes: Construct the basic module of the aircraft target image feature point detection model; Construct the data input network architecture of the aircraft target image feature point detection model; Construct the backbone network architecture of the aircraft target image feature point detection model; Construct the neck network architecture of the aircraft target image feature point detection model; Construct the prediction output network architecture of the aircraft target image feature point detection model; Construct the loss function of the aircraft target image feature point detection model.
5. The method for detecting feature points of an aircraft target image according to claim 4, characterized in that: The aircraft target image feature point detection model is specifically: L=L label +λ unlabel L unlabel , Among them, L label represents the loss score using labeled image data; L unlabel represents the loss score using unlabeled image data; λ unlabel Represents the balance coefficient of the loss score for unlabeled image data.
6. A device for detecting feature points of an aircraft target image, characterized in that: include: A processing unit, used for collecting aircraft target image data and preprocessing the collected aircraft target image data; A collection unit, used for constructing an aircraft target image data set using the preprocessed aircraft target image data; a labeling unit, used for labeling the image data in the aircraft target image data set; A construction unit, used to construct an aircraft target image feature point detection model using the annotated image data; A training unit, used for training the aircraft target image feature point detection model by using a semi-supervised learning method; The detection unit is used to detect the target image feature points of the aircraft using the aircraft target image feature point detection model.
7. The device for detecting feature points of an aircraft target image according to claim 6, characterized in that: The processing unit comprises: A screening module is used to screen the aircraft target image data according to environmental conditions, aircraft distance, and aircraft attitude; The scaling module is used to scale the image data of the selected aircraft target from a×b to a′×b′ according to the scale γ, where a′=γa≤a o , b′=γb≤b o ,γ=min(a o / a,b o / b), a o ×b o Represents the standardized image size, and fills the part that is less than the standardized image size with the background color; The statistical module is used to count the histogram of the aircraft target image data. The threshold is set to T, and the gray value at 1-T is set to p min , set the gray value at T to p max , bilinear contrast stretching is used to standardize the gray value p of the image, and the pixel gray value of the image data is standardized to between 0 and 1. The specific standardization process is: p o =(p-p min ) / (p max -p min )。 8. The device for detecting feature points of an aircraft target image according to claim 7, characterized in that: The marking unit comprises: Used to mark the aircraft target on the image data and establish an aircraft label; It is used to select feature points on the aircraft target image for annotation and create aircraft feature point labels.
9. The device for detecting feature points of an aircraft target image according to claim 8, characterized in that: The construction unit comprises: Basic modules for building aircraft target image feature point detection models; The data input network architecture used to build the aircraft target image feature point detection model; Backbone network architecture for building a model for detecting feature points in aircraft target images; Neck network architecture for building a model for detecting feature points in aircraft target images; The prediction output network architecture used to build the aircraft target image feature point detection model; Loss function used to build a model for detecting feature points in aircraft target images.
10. The device for detecting feature points of an aircraft target image according to claim 9, characterized in that: The aircraft target image feature point detection model is specifically: L=L label +λ unlabel L unlabel , Among them, L label represents the loss score using labeled image data; L unlabel represents the loss score using unlabeled image data; λ unlabel Represents the balance coefficient of the loss score for unlabeled image data.