Unsupervised aircraft key point detection method based on isotropic constraint and skeleton generation
By introducing such denaturation constraints and skeleton generation technologies in unsupervised aircraft key point detection, the problem of detection point drift and semantic inconsistency is solved, and stable, robust and accurate key point detection is achieved, which is suitable for application scenarios for label-free data training.
Patent Information
- Application Number
- CN202510117002.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The prior art is difficult to achieve stable, robust and accurate aircraft key point detection under unsupervised conditions, especially when facing interferences such as rotation, scale changes, motion blur, and light changes, atmospheric turbulence, complex backgrounds, etc., the detection points are prone to drift and have inconsistent semantics.
The unsupervised aircraft key point detection method based on isodenatment constraints and skeleton generation is adopted. By random affine transformation of the input image, the key points of the original image and the transformed image are extracted, and the isodenatment constraints are applied, the skeleton is drawn using the key points to extract the common structural information of the aircraft.
Automatic detection without manual labeling of key point labels is realized, which improves the robustness and accuracy of detection, solves the problems of key point drift and semantic inconsistency, and enhances the generalization and robustness of the network.
Smart Images

Figure CN119942145A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of computer vision key point detection, and in particular relates to an unsupervised aircraft key point detection method based on equivariance constraints and skeleton generation. Background Art
[0002] Key point detection is very important in computer vision. It is the basis of tasks such as image matching, target tracking, 3D reconstruction, and attitude estimation. For optoelectronic tracking equipment, with the increase in frame rate and resolution, tracking tasks also put forward higher requirements for the high-precision positioning capability of flying targets. Key point detection technology is of great significance for the precise measurement of optoelectronic systems. It can obtain highly semantic attitude representation and perception, which is conducive to subsequent attitude estimation and trajectory analysis.
[0003] Traditional key point detection algorithms describe features or build models for selected local areas, and rely on classifiers or regressors to detect key points in local areas based on the calculated features or established models. Traditional algorithms have the advantages of simple structure, low computational complexity, and strong interpretability, but they require manual design of feature descriptors, are not universal, and are not robust to illumination, posture changes, and motion blur. Key point detection methods based on supervised learning use deep neural networks to automatically learn feature representations from a large amount of labeled data, and can achieve accurate key point positioning. However, its dependence on labeled data and high labeling costs limit its application in some fields. In contrast, key point detection methods based on unsupervised learning have broader application prospects. In the process of network training, the real labels of the data are not required as supervision, which undoubtedly saves a lot of time and energy. They use unsupervised learning methods in deep learning to extract features of targets in images. After data integration processing, they expect to locate those stable and distinguishable points in the image. However, in practical applications, this is very challenging. Due to the complexity of the scene and the flying target itself, the target will rotate, change scale, and blur during the movement. At the same time, it will be disturbed by factors such as lighting changes, atmospheric turbulence, and complex backgrounds, which will affect the apparent characteristics of the target, resulting in detection point drift and semantic inconsistency. Therefore, developing a stable, robust, and accurate unsupervised key point detection algorithm for actual scenes is still an urgent problem to be solved in engineering applications. Summary of the invention
[0004] To solve the above technical problems, the present invention provides an unsupervised aircraft key point detection method based on equivariance constraints and skeleton generation. Equivariance constraints are introduced into the key point detection network to solve the problem of inconsistent key point semantics. At the same time, the skeleton is drawn using key points to extract common structural information of the aircraft to solve the problem of key point drift.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The unsupervised aircraft key point detection method based on equivariance constraints and skeleton generation has the following steps:
[0007] Step 1: Input image Do a random affine transformation to get the transformed image ;
[0008] Step 2: Use the key point extraction network to extract the input image separately and the transformed image Extract key points and obtain the first and second key point sets B1={ } and B2={ },in is the number of key points. First, use the Simple Baseline key point extraction network to extract the input image and the transformed image Heat map of the first key point and the second key point heat map , for the first key point heat map Use the Soft-argmax function to get the coordinates of each key point , heat map of the second key point Use the Softmax function to get the probability map of key points , for the probability graph Apply the inverse transformation of the random affine transformation in step 1 to obtain the transformation probability map , for the transformed probability map Take the expectation in the coordinate space to get the coordinates of each key point The calculation formulas for the coordinates, probability map, and transformation probability map of key points are as follows:
[0009] ,
[0010] in, is the normalized pixel coordinate, It is the inverse transformation operation of the random affine transformation;
[0011] Step 3: For the first and second key point sets B1={ } and B2={ }Implement equivariant constraints, equivariant losses The calculation formula is as follows:
[0012] ,
[0013] in, is the number of samples in the dataset, It is The key point set of samples, It is The set of key points extracted after random affine transformation of samples;
[0014] Step 4: Use the key point set B1={ Draw a skeleton diagram . Take two key points in the first key point set and , draw a differentiable edge graph , assign a trainable weight greater than zero to each edge Get the weighted edge graph , by taking the weighted posterior graph The maximum value of each pixel in the image is used to obtain the final skeleton image. :
[0015] ,
[0016] in, is a hyperparameter that controls the thickness of the skeleton. is the normalized pixel coordinate To two key points and The Euclidean distance of the drawn edge, t is calculated of control variables;
[0017] Step 5: Input image Divide the image into several grids, randomly block 80% of the image data in the grids, and obtain an image with destroyed structural information. ;
[0018] Step 6: For images with destroyed structural information Apply weight , and with the skeleton diagram Perform splicing in the channel dimension to obtain a spliced feature map;
[0019] Step 7: Send the spliced feature map to the UNet decoder structure to obtain the reconstructed image ;
[0020] Step 8: Input image and reconstruct the image Reconstruction constraints are imposed. Reconstruction constraints are defined based on perceptual loss, which is calculated by evaluating the activation values of the image at different levels in the convolutional neural network. The calculation formula of reconstruction loss is as follows:
[0021] ,
[0022] in, is the number of samples in the dataset, It is a feature extractor, using the pre-trained VGG convolutional neural network. Indicates input images;
[0023] Step 9: Calculate the total loss function and train the network by minimizing the total loss function. , is the weight coefficient for balancing the reconstruction loss and the equivariant loss.
[0024] The beneficial effects of the present invention are:
[0025] (1) The present invention does not require manual labeling of key point labels, saving manpower, material and financial resources. At the same time, it can automatically discover the structure and pattern in the data and has strong universality.
[0026] (2) The present invention uses key points to draw the skeleton, and uses each sample in the training set to share and update the weights of each edge of the skeleton, thereby extracting the common structural information of the aircraft and capturing the complex structural dependencies between the various parts of the aircraft, thus solving the problem of key point drift.
[0027] (3) The present invention performs a random affine transformation on the input image, extracts the key points of the original image and the transformed image respectively, and imposes equivariance constraints on them, thereby enabling the network to learn various transformation relationships, making it more robust and solving the problem of inconsistent key point semantics. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a flow chart of the unsupervised aircraft key point detection method based on equivariance constraint and skeleton generation of the present invention;
[0029] Figure 2 Flowchart for equivariance constraints;
[0030] Figure 3 Draw the renderings for the skeleton;
[0031] Figure 4 This is the key point detection effect diagram. DETAILED DESCRIPTION
[0032] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0033] The present invention provides an unsupervised aircraft key point detection method based on equivariance constraint and skeleton generation, which solves the problems in the prior art that a large amount of labeled data is required for training, and detection point drift and semantic inconsistency occur in the face of rotation, scale change, motion blur, etc., or interference from illumination change, atmospheric turbulence, complex background, etc. Among them, the key point is essentially a feature, which is an abstract description of a fixed area or spatial physical relationship. It not only represents a position, but also represents the combined relationship between the context and the surrounding neighborhood, reflecting the structural information. Aircraft of the same category should have the same topological structure, and they can be represented by similar skeleton structures. Using key points to draw and generate skeleton information can not only learn the common structural features of aircraft of the same category, but also capture the complex structural dependency between the various parts of the aircraft, so that the network has stronger generalization and robustness, and can cope with the situation of key point drift. When the aircraft attitude changes, due to the existence of a robust topological structure of the skeleton, the detected key points can generally be distributed in various parts of the aircraft, but at this time the semantics of the key points are likely to be inconsistent, such as the head of an aircraft in one frame may correspond to the tail of an aircraft in another frame, because it is difficult for the network to handle such a change relationship. The present invention solves the problem of semantic inconsistency by performing random affine transformation on the input image, extracting key points of the original image and the transformed image respectively, and applying equivariance constraints to them.
[0034] like Figure 1 As shown, it is a specific process of the present invention, which includes the following steps:
[0035] Step 1: Input image Do a random affine transformation to get the transformed image ;
[0036] Affine transformation refers to the transformation of an image in two-dimensional or three-dimensional space through linear transformation (such as rotation, scaling, shearing, etc.) and translation, which maintains the parallel linear relationship and proportional relationship of the image. It does not change the straightness of the straight lines in the image and the proportional relationship of the relative positions. Affine transformation can be represented by a 2x3 matrix so that each point in the image is mapped to a new location The matrix form of the affine transformation is:
[0037] ,
[0038] in, Determines operations such as rotation, scaling, and shearing. is the amount of translation along the x, y direction. Since affine transformation is a combination of linear transformation (such as rotation, scaling, shearing) and translation, Used to control rotation, scaling and shearing operations, Used to control translation operations. By changing these six parameters, different transformation effects can be achieved. That is, random affine transformation can simulate the random attitude changes of the aircraft.
[0039] Step 2: Use the key point extraction network to extract the input image separately and the transformed image Extract key points and obtain the first and second key point sets B1={ } and B2={ },in is the number of key points. Figure 2 The following is a flowchart of equivariance constraints. First, the Simple Baseline key point extraction network is used to extract the input image separately. and the transformed image Heat map of the first key point and the second key point heat map , for the first key point heat map Use the Soft-argmax function to get the coordinates of each key point , heat map of the second key point Use the Softmax function to get the probability map of key points , for the probability graph Apply the inverse transformation of the random affine transformation in step 1 to obtain the transformation probability map , for the transformed probability map Take the expectation in the coordinate space to get the coordinates of each key point The calculation formulas for the coordinates, probability map, and transformation probability map of key points are as follows:
[0040] ,
[0041] in, is the normalized pixel coordinate, It is the inverse transformation operation of the random affine transformation;
[0042] Step 3: For the first and second key point sets B1={ } and B2={ }Implement equivariant constraints, equivariant losses The calculation formula is as follows:
[0043] ,
[0044] in, is the number of samples in the dataset, It is The key point set of samples, It is The set of key points extracted after random affine transformation of samples;
[0045] Step 4: Use the key point set B1={ Draw a skeleton diagram . Take two key points in the first key point set and , draw a differentiable edge graph , assign a trainable weight greater than zero to each edge Get the weighted edge graph , by taking the weighted posterior graph The maximum value of each pixel in the image is used to obtain the final skeleton image. :
[0046] ,
[0047] in, is a hyperparameter that controls the thickness of the skeleton. is the normalized pixel coordinate To two key points and The Euclidean distance of the drawn edge, t is calculated of the control variables. Figure 3 The skeleton is drawn as shown in the figure. For different rotations and scale changes, the network extracts the common skeletal structure information of the aircraft. The key points are stably distributed in various parts of the aircraft, and there is no key point drift.
[0048] Step 5: Input image Divide the image into several grids, randomly block 80% of the image data in the grids, and obtain an image with destroyed structural information. ;
[0049] Step 6: For images with destroyed structural information Apply weight , and with the skeleton diagram Perform splicing in the channel dimension to obtain a spliced feature map;
[0050] Step 7: Send the spliced feature map to the UNet decoder structure to obtain the reconstructed image ;
[0051] Step 8: Input image and reconstruct the image Reconstruction constraints are imposed. Reconstruction constraints are defined based on perceptual loss, which is calculated by evaluating the activation values of the image at different levels in the convolutional neural network. The calculation formula of reconstruction loss is as follows:
[0052] ,
[0053] in, is the number of samples in the dataset, It is a feature extractor, using the pre-trained VGG convolutional neural network. Indicates input images;
[0054] Step 9: Calculate the total loss function and train the network by minimizing the total loss function. , is the weight coefficient for balancing the reconstruction loss and the equivariant loss.
[0055] Through steps 1-9 of the present invention, a trained Simple Baseline key point extraction network can be obtained. By inputting an airplane picture into the network, stable and robust key point coordinates can be obtained. Figure 4 This is the effect diagram of aircraft key point detection. As shown in the figure, the semantics of the key points are consistent despite different rotations and scale changes of the aircraft. This is due to the equivariance constraint. By performing a random affine transformation on the input image, extracting the key points of the original image and the transformed image respectively and calculating the equivariance loss for them, the network can learn various transformation relationships of the aircraft's posture, making it more robust.
[0056] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An unsupervised aircraft key point detection method based on equivariance constraints and skeleton generation, characterized in that: The following steps are involved: Step 1: Input image Do a random affine transformation to get the transformed image ; Step 2: Use the key point extraction network to extract the input image separately and the transformed image Extract key points and obtain the first and second key point sets B1={ } and B2={ },in is the number of key points; Step 3: For the first and second key point sets B1={ } and B2={ }Implement equivariance constraints and calculate equivariance losses ; Step 4: Using the first key point set Draw a skeleton diagram ; Step 5: Input image Applying random masks to obtain images with destroyed structural information ; Step 6: For images with destroyed structural information Apply weight , and with the skeleton diagram Perform splicing in the channel dimension to obtain a spliced feature map; Step 7: Send the spliced feature map to the decoder to obtain the reconstructed image ; Step 8: Input image and reconstruct the image Apply reconstruction constraints and calculate reconstruction loss ; Step 9: Calculate the total loss function and train the network by minimizing the total loss function.
2. The unsupervised aircraft key point detection method based on equivariance constraint and skeleton generation according to claim 1 is characterized in that: In the step 1, the random affine transformation refers to transforming the image through a linear transformation in a two-dimensional or three-dimensional space to maintain the parallel linear relationship and proportional relationship of the image.
3. The unsupervised aircraft key point detection method based on equivariance constraint and skeleton generation according to claim 1 is characterized in that: In step 2, the Simple Baseline key point extraction network is used to extract the input image and the transformed image Heat map of the first key point and the second key point heat map , for the first key point heat map Use the Soft-argmax function to get the coordinates of each key point , for the second key point heat map Use the Softmax function to get the probability map of key points , for the probability graph Apply the inverse transformation of the random affine transformation to obtain the transformation probability map , for the transformed probability map Take the expectation in the coordinate space to get the coordinates of each key point .
4. The unsupervised aircraft key point detection method based on equivariance constraint and skeleton generation according to claim 3 is characterized in that: The calculation formulas for the coordinates, probability map, and transformation probability map of key points are as follows: , in, is the normalized pixel coordinate, It is the inverse transformation operation of the random affine transformation.
5. The unsupervised aircraft key point detection method based on equivariant constraints and skeleton generation according to claim 1 is characterized in that: The intermediate variability constraint in step 3 is given by Norm definition, equivariant loss The calculation formula is as follows: , in, is the number of samples in the training data set, It is The key point set of samples, It is The set of key points extracted after random affine transformation of samples.
6. The unsupervised aircraft key point detection method based on equivariance constraint and skeleton generation according to claim 5 is characterized in that: In step 4, two key points are selected from the first key point set. and , draw a differentiable edge graph , assign a trainable weight greater than zero to each edge Get the weighted edge graph , by taking the weighted edge graph The maximum value of each pixel in the image is used to obtain the final skeleton image. : , in, is a hyperparameter that controls the thickness of the skeleton. is the normalized pixel coordinate To two key points and The Euclidean distance of the drawn edge, t represents the calculation of the control variables.
7. The unsupervised aircraft key point detection method based on equivariance constraint and skeleton generation according to claim 6 is characterized in that: The step five includes inputting the image Divide the image into several grids, randomly block 80% of the image data in the grids, and obtain an image with destroyed structural information. .
8. The unsupervised aircraft key point detection method based on equivariance constraint and skeleton generation according to claim 7 is characterized in that: In step 7, the UNet decoder is used to reconstruct the original input image. , and obtain the reconstructed image .
9. The unsupervised aircraft key point detection method based on equivariant constraints and skeleton generation according to claim 5, characterized in that: The reconstruction constraint in step eight is defined based on the perceptual loss, which is calculated by evaluating the activation values of the image at different levels in the convolutional neural network. The calculation formula of the reconstruction loss is as follows: , in, is the number of samples in the dataset, is a feature extractor, Indicates Input image samples.
10. The unsupervised aircraft key point detection method based on equivariant constraints and skeleton generation according to claim 9, characterized in that: The total loss function , is the weight coefficient for balancing the reconstruction loss and the equivariant loss.
Citation Information
Patent Citations
Face key point detection method and system based on sparse key point calibration
CN110826501A
Self-supervision human body key point detection method based on affine transformation
CN117636389A