Aircraft pose estimation method from virtuality to reality, equipment, medium and product

By using semantic segmentation and style transfer techniques in aircraft pose estimation, the style of virtual images is transferred to the real image, which solves the problem of insufficient real data, improves the accuracy and robustness of pose estimation, and is suitable for applications in different real scenarios.

CN120014028APending Publication Date: 2025-05-16BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510100485.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, due to insufficient real data in aircraft posture estimation, the difference between virtual images and real images leads to poor performance of the model in real scenes, making it difficult to improve the accuracy of aircraft posture estimation.

Method used

The aircraft posture estimation method from virtual to real is used to obtain the mask of the aircraft image through a semantic segmentation network, and a synthetic aircraft image under different postures is generated using a rendering program. The virtual image style is transferred to the real image in combination with the style migration network, a stylized image is generated, and a pose estimation model is input for training to improve the accuracy of pose estimation.

Benefits of technology

By reducing the difference between virtual and real images, the accuracy and robustness of the aircraft pose estimation model in real scenes is significantly improved, the demand for real labeled data is reduced, and the application of the pose estimation model in different real scenes is promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014028A_ABST
    Figure CN120014028A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual-to-real aircraft pose estimation method and device, a medium and a product, and relates to the field of pose estimation, and the method comprises the steps: obtaining an aircraft image; acquiring a mask of the aircraft image by adopting a semantic segmentation network; a rendering program is adopted to obtain synthetic aircraft images under different postures and masks corresponding to the synthetic aircraft images based on the 3D model of the aircraft; inputting the aircraft image, the mask of the aircraft image, the synthetic aircraft image in different attitudes and the mask corresponding to the synthetic aircraft image into a style migration network to obtain a stylized image; inputting the stylized image into a pose estimation model, and training the pose estimation model; and estimating the attitude of the aircraft in the real aircraft image by using the trained attitude estimation model. According to the invention, the accuracy of aircraft pose estimation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of posture estimation, and in particular to a method, device, medium and product for estimating the posture of an aircraft from virtual to real. Background Art

[0002] Test flight is a very important part of aircraft development and manufacturing. This part tests whether the various indicators of the aircraft meet the requirements in various outdoor environments. Measuring the posture of the aircraft through an external visual system is of great help to the smooth progress of the test flight and the evaluation of aircraft performance. In recent years, the application of posture estimation methods based on neural networks in aerospace manufacturing and testing has received more and more attention.

[0003] Style transfer methods mainly transfer styles by aligning the statistical features of content images and stylized images. Early methods aligned the global mean and variance, etc., and some recent studies have achieved good transfer effects by introducing contrastive learning or attention mechanisms to construct the transfer relationship between the local space of content and style. Among them, the adaptive normalization method known to the inventor and the covariance feature transformation method have laid the foundation for this field. However, these two methods only align the global statistical information of the features, resulting in distortion and unwanted style textures in the generated images. In order to maintain the content, an end-to-end photo-realistic style transfer method based on wavelet transform is proposed to effectively transfer the style while maintaining the image details. Recently, the attention mechanism has been introduced into style transfer to align the local style according to the distribution of the semantic features of the content, achieving good transfer effects. The simple covariance transform network (SCTNet) proposes a contrast consistency preservation loss, which maximizes the difference between the features of the same region and other different regions through the contrast loss InforNCE to maintain the content. The domain-aware indicator proposed by DSTN (Domain-aware Style Transfer Networks, Domain-aware Style Transfer Network) transfers arbitrary styles and balances style and content through domain-aware jump connections. CAP-VSTNet (ContentAffinity PreservedVersatile Style Transfer) proposes a novel reversible residual network and unbiased linear transformation module, which maintains content affinity and achieves diversified style transfer by introducing Matting Laplacian training loss.

[0004] Although the current methods have achieved good results in the image-to-image style transfer task and can use masks to clearly identify the corresponding style and content areas, the aircraft stylization effect still has many shortcomings despite the use of masks to remove background interference.

[0005] The pose estimation methods based on deep learning are mainly divided into pose estimation methods based on 2D-3D correspondence and pose estimation methods based on template matching. The first one mainly establishes sparse or dense 2D-3D feature correspondences by voting, regression or encoding, and then solves the pose through PnP. The method based on dense feature correspondence has achieved the best accuracy, but the robustness and real-time performance need to be improved. The method based on template matching estimates the pose of the target by comparing the similarity between the representation vector of the real image and the representation vector of the template. This type of method is based on metric learning and reconstruction learning to constrain the similarity between the image representations of the same pose. This type of method has good robustness, but can only be used to estimate pose.

[0006] All pose estimation methods based on neural networks require a large number of real aircraft images with pose annotations to train the network in order to obtain good accuracy in real scenes. However, it is not only difficult to obtain a large number of aircraft images in real environments from real scenes, but also difficult to manually annotate the pose. Based on this, one way to solve the problem of insufficient real data is to use computers to generate virtual images for training pose estimation models. However, the difference between virtual images and real images causes the trained model to perform poorly in real scenes. An effective solution is to use image style transfer methods to transfer the style of real images to virtual images to reduce the difference between them. Among them, pose estimation methods based on deep learning are mainly divided into methods based on 2D-3D feature correspondence and methods based on template matching.

[0007] Early methods based on feature correspondence mainly extract sparse features of the target through voting or regression methods. In order to improve the robustness of sparse feature correspondence, a long-short range perception strategy is proposed to improve the accuracy of key point positioning in a coarse-to-fine manner. The performance of the student network is improved by aligning the local prediction distribution of the compact student network with the local prediction distribution of the teacher network. Sparse features are easy to define and solve, but their detection and matching accuracy will be significantly reduced by various changes such as occlusion and illumination. Therefore, methods based on dense 2D-3D point correspondence have received more and more research due to their better accuracy and robustness. For example, learnable and manually defined encodings are used to locate dense surface points. Shape constraints are added on the basis of dense point correspondence to improve the accuracy of dense point correspondence. The probabilistic generation model based on neural embedding not only improves robustness, but also quantifies uncertainty. Although methods based on dense point correspondence can achieve better accuracy, they have high time complexity.

[0008] The earliest pose estimation methods based on template matching were metric learning-based methods and reconstruction learning-based methods. Subsequently, edge and shape prior information was introduced to further improve robustness. It is proposed to use 2-D Lorentz distribution function to learn image similarity to capture the semantic relationship between different head pose images. The metric-based and learning-based methods are integrated, and semantic feature constraints are proposed to enhance the matching accuracy of the network in different environments. The template matching method performs well in speed and robustness, but the accuracy is still lower than that of the dense point correspondence method, and it can only estimate 3D poses.

[0009] Some pose estimation methods for aircraft have been proposed, but these methods require the collection of a large number of aircraft images with pose annotations. However, it is difficult to obtain a large number of aircraft images in real scenes, and labeling masks, detection boxes and other labels is very time-consuming and labor-intensive. In particular, there are many types of aircraft, and it is unrealistic to label different types of aircraft. One way to solve the problem of insufficient real annotated image data is to use computers to generate virtual images for training pose networks. Compared with real images, virtual images are easier to obtain the pose information of the target. However, there are often significant differences between virtual images and real images, so it is difficult to obtain good pose estimation accuracy in real scenes using virtual image training networks. In order to improve the accuracy of pose estimation in real scenes, the research focus in this field is also on how to make up for the difference between virtual and real images.

[0010] Some existing studies have noticed the impact of the difference between virtual images and real images on the accuracy of pose estimation, and proposed unsupervised pose estimation methods based on style transfer. The style of virtual images is transformed into the style of real images through style transfer methods to obtain realistic images to train pose estimation networks. For example, the unsupervised pixel-level domain adaptation method proposes to use generative adversarial networks to convert virtual domain images to real domains and then train pose estimation models, showing the potential of such methods. The deformation-aware unpaired image transformation method realizes cross-domain transfer of pose annotations and images by estimating the deformation field from the source insect image to the target insect image and generative adversarial networks. However, this method is applied to the estimation of insect 2D poses, and the output image is grayscale. The unsupervised domain adaptation pose estimation method uses the AdaIN (Adaptive Instance Normalization) style transfer method to transfer virtual images to the real domain, and aligns the output key point heat map through the mean teacher model. The above methods mainly use existing image-to-image transformation methods to transform the target image, but these style transfer methods are mainly used for image-to-image style transfer, and have not been studied for the target. In addition, the above methods are mainly evaluated in small-scale indoor environments, and their performance in complex outdoor scenes is still unknown. Moreover, current style transfer is from image to image, and there is no target-oriented style transfer method. The authenticity of previous style transfer methods is still significantly different from the real image, which has limited help in improving the performance of subsequent pose estimation models.

[0011] In summary, the application of aircraft pose estimation methods based on deep neural networks in the test flight phase of aircraft manufacturing has attracted more and more attention. However, it is very difficult to collect and annotate a large number of real images with pose labels for training the network. An alternative solution is to use virtual aircraft images to train the network. However, due to the difference between synthetic images and real images, the accuracy and robustness of the pose estimation model in real scenes are very poor, which makes it impossible to improve the accuracy of aircraft pose estimation. Therefore, it is very necessary to study Summary of the invention

[0012] The purpose of this application is to provide a method, device, medium and product for estimating aircraft posture from virtual to real, which can improve the accuracy of aircraft posture estimation.

[0013] To achieve the above objectives, this application provides the following solutions:

[0014] In a first aspect, the present application provides a method for estimating an aircraft pose from virtual to real, comprising:

[0015] Acquire aircraft images;

[0016] Using a semantic segmentation network to obtain a mask of the aircraft image;

[0017] Using a rendering program based on a 3D model of the aircraft, obtaining synthetic aircraft images in different postures and masks corresponding to the synthetic aircraft images;

[0018] Inputting the aircraft image, the mask of the aircraft image, synthetic aircraft images in different postures, and the mask corresponding to the synthetic aircraft image into a style transfer network to obtain a stylized image;

[0019] Inputting the stylized image into a pose estimation model to train the pose estimation model;

[0020] The trained pose estimation model is used to estimate the pose of the aircraft in the real aircraft image.

[0021] Optionally, the training process of the style transfer network includes:

[0022] Obtain an airplane image sample, and use the semantic segmentation network to obtain a mask of the airplane image sample;

[0023] Using a rendering program, based on a 3D model of an aircraft, to obtain rendering data; the rendering data includes: synthetic aircraft image samples in different postures and masks and posture information corresponding to the synthetic aircraft image samples;

[0024] The style transfer network is trained using the aircraft image sample, the mask of the aircraft image sample, and the rendering data.

[0025] Optionally, the training process of the pose estimation model includes:

[0026] Inputting the aircraft image sample, the mask of the aircraft image sample, the synthesized aircraft image sample, and the mask corresponding to the synthesized aircraft image sample into the style transfer network to obtain a stylized image sample;

[0027] The pose estimation model is trained using the stylized image samples and pose information in the rendering data.

[0028] Optionally, the style transfer network includes an encoder, a mask-based adaptive multi-head attention normalization module and a decoder; the structure of the encoder and the structure of the decoder are mirror-symmetrical;

[0029] The output features of the encoder are sent to the decoder through the mask-based adaptive multi-head attention normalization module to obtain a stylized image.

[0030] Optionally, VGG19 is used as the encoder.

[0031] Optionally, the style transfer network adopts a function that constrains neighboring content consistency, global style consistency, and local style consistency as a loss function.

[0032] Optionally, the loss function is expressed as:

[0033]

[0034] In the formula, represents the total loss, represents the global style loss, α represents the hyperparameter for balancing the global style loss, represents the local style loss, β represents the hyperparameter for balancing the local style loss, represents the neighborhood content consistency loss, and γ represents the hyperparameter that balances the neighborhood content consistency loss.

[0035] In a second aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-mentioned method for estimating aircraft pose from virtual to real.

[0036] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned method for estimating an aircraft pose from virtual to real.

[0037] In a fourth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for estimating aircraft posture from virtual to real.

[0038] According to the specific embodiments provided in this application, this application has the following technical effects:

[0039] The present application provides a method, device, medium and product for estimating the pose of an aircraft from virtual to real. The mask of an aircraft image is obtained by an unsupervised segmentation method (i.e., a semantic segmentation network), and then a style transfer network is used to generate a stylized image using the mask. The stylized image is input into a pose estimation model to train the pose estimation model. The pose of the aircraft in the real aircraft image is estimated using the trained pose estimation model, thereby improving the accuracy of the aircraft pose estimation. In addition, the use of an unsupervised segmentation method, a style transfer network and a pose estimation model can reduce the demand for real labeled data in actual applications and promote the application of pose estimation models in different real scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 A flowchart of a method for estimating an aircraft posture from virtual to real provided in one embodiment of the present application;

[0042] Figure 2 A schematic diagram of the framework of the aircraft pose estimation method from virtual to real;

[0043] Figure 3 A schematic diagram of the structure of a style transfer network provided in another embodiment of the present application;

[0044] Figure 4 A schematic diagram of a mask-based adaptive multi-head attention normalization module provided in an embodiment of the present application;

[0045] Figure 5 A schematic diagram of a single-head adaptive attention module provided in an embodiment of the present application;

[0046] Figure 6 An array diagram of visual comparison results of different methods provided in another embodiment of the present application;

[0047] Figure 7 A graph showing SIFID, style loss, and content loss under a hyperparameter α provided in another embodiment of the present application;

[0048] Figure 8 A graph showing SIFID, style loss, and content loss under a hyperparameter β provided in another embodiment of the present application;

[0049] Fig. 9 A graph showing SIFID, style loss, and content loss under a hyperparameter γ provided in another embodiment of the present application;

[0050] Fig.10 An array diagram showing the effect of different weights of loss on style transfer provided by another embodiment of the present application;

[0051] Fig.11 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0052] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0053] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0054] In an exemplary embodiment, the present application provides a method for estimating the posture of an aircraft from virtual to real, which is executed by a computer device, specifically, it can be executed by a computer device such as a terminal or a server alone, or it can be executed by a terminal and a server together. In the embodiment of the present application, the method is applied to a server as an example for explanation. Figure 1 As shown, the method includes:

[0055] Step 100: Acquire an aircraft image. The aircraft image may be acquired by using an optoelectronic tracking imaging system deployed around the runway.

[0056] Step 101: Use a semantic segmentation network to obtain a mask of the aircraft image. In order to train the segmentation network without relying on any manually annotated labels, an unsupervised foreground segmentation method FOUND can be used as a semantic segmentation network to obtain a mask of the aircraft image.

[0057] Step 102: Using a rendering program based on the 3D model of the aircraft, obtain synthetic aircraft images in different postures and masks corresponding to the synthetic aircraft images.

[0058] Step 103: Input the aircraft image, the mask of the aircraft image, the synthetic aircraft images in different postures, and the mask corresponding to the synthetic aircraft image into a style transfer network to obtain a stylized image.

[0059] Step 104: Input the stylized image into the pose estimation model to train the pose estimation model.

[0060] Step 105: Estimate the posture of the aircraft in the real aircraft image using the trained posture estimation model.

[0061] In another exemplary embodiment of the present application, the training process of the style transfer network includes:

[0062] Step 1: Obtain an airplane image sample and use a semantic segmentation network to obtain the mask of the airplane image sample.

[0063] Step 2: Using a rendering program, based on the 3D model of the aircraft, to obtain rendering data. The rendering data includes: synthetic aircraft image samples in different postures and masks and posture information corresponding to the synthetic aircraft image samples.

[0064] Step 3: Use the airplane image samples, the masks of the airplane image samples, and the rendering data to train the style transfer network.

[0065] In another exemplary embodiment of the present application, the training process of the pose estimation model includes:

[0066] Step 1: Input the airplane image sample, the mask of the airplane image sample, the synthesized airplane image sample, and the mask corresponding to the synthesized airplane image sample into the style transfer network to obtain a stylized image sample.

[0067] In actual application, the trained style transfer network is used to generate stylized image samples, and the background is added to the stylized image according to the mask (i.e., the mask of the aircraft image sample and the mask corresponding to the synthetic aircraft image sample). These stylized images and the corresponding masks and pose information (i.e., pose labels) constitute the stylized aircraft dataset for training pose estimation models.

[0068] In order to create a stylized aircraft dataset, it is necessary to collect real aircraft images and rendered synthetic images. Aircraft images are collected through visual tracking measurement equipment. For example, the equipment is deployed on a vehicle base to collect images of the takeoff and landing phases of the C919 aircraft test flight. 3681 real images are selected to generate stylized images. In addition, in order to evaluate the accuracy of pose estimation, the pose of the aircraft relative to the camera is calibrated using inertial navigation and GPS, and 3297 images are divided as training sets and 384 images as test sets to compare the accuracy differences between pose estimation models trained with stylized images and real images.

[0069] The synthetic aircraft images were generated by Pytorch3D rendering. The aircraft poses used for rendering were randomly sampled from the poses obtained during the actual test flight, with a total of 73,620 images. For each virtual aircraft image, 5 different real aircraft images were randomly selected to stylize it. A total of 368,100 stylized images were generated, and the pose label of each stylized image was the same as the corresponding virtual image.

[0070] Since the style transfer network only targets the aircraft area, in order to train the pose estimation model, it is necessary to add a background to the stylized image for the training of the pose estimation model. The MAT method is used to reconstruct the background of the real image after removing the aircraft, and a total of 3681 background images are generated. During training, a background image is randomly selected as the background of the stylized image.

[0071] Step 2: Use the pose information in the stylized image samples and the rendered data to train the pose estimation model.

[0072] For example, in actual applications, the previously proposed semantic segmentation network and the FOUND unsupervised segmentation model are used to obtain the aircraft mask for style transfer training. In the training experiment, only the effects of the supervised and unsupervised segmentation models on style transfer are shown, and the two segmentation models are not evaluated.

[0073] (1) Style transfer training. Since the real aircraft images are continuous video frames, the style and pose between adjacent images are basically unchanged, so one frame is extracted every 40 frames, and a total of 93 real images are used for training. Since the content and style of the virtual images are single, only 44 images are randomly selected for training. The batch size of the training is 1, iters = 3000, and the learning rate is 1e-4.

[0074] (2) Pose estimation model training. iters=100K, and all stylized images are used for training.

[0075] During the training process, the evaluation index of the style transfer network refers to the method known to the inventor, and the indexes for evaluating the style transfer network mainly use SIFID (single image Frechet Inception Distance), style loss and content loss.

[0076] SIFID is achieved by calculating the Frechet Inception Distance (FID) between the deep features of two real images and the stylized image. The lower the SIFID score, the more similar the styles of the two images are. The calculation formula of FID is:

[0077]

[0078] in, Represents the square of the mean difference between the generated and style feature maps, which measures the distance between the mean of the generated image and the real image features. Tr represents the trace of the matrix (i.e., the sum of the main diagonal elements of the matrix), which is used to measure the difference between the covariance of the generated image and the real image features. ∑ g ,Σ r are the covariance matrices of the generated image and stylized image features, is the square root of the product of the covariance matrices.

[0079] Style loss and content loss. These two items respectively calculate the difference between the generated image and the stylized image, and between the generated image and the content image to evaluate the quality of the generated image. Among them, the content loss and style loss It is expressed as:

[0080]

[0081] Among them, φ i is the i-th network layer, I c represents the content image, N l Represents the number of convolutional layers. The smaller the style loss, the better the style transfer effect. The smaller the content loss, the stronger the content maintenance.

[0082] Based on the above description, the method for estimating the aircraft posture from virtual to real provided by the present application corresponds to the following Figure 2 In order to verify the validity of the stylized data and the applicability of this framework to different pose estimation methods, in this application, the 2D-3D corresponding pose estimation method based on YOLOv8 and the template matching method can be selected for evaluation. Experiments show that the stylized data provided in this application Figure 2 The framework shown can achieve comparable accuracy to models trained on real images.

[0083] Among them, the template matching method generates a representation of the aircraft through the network. The query pose is determined by comparing the similarity between the query (real aircraft image) and the template, which is:

[0084]

[0085] In the formula, R * Represents a certain posture (rotation matrix), z i represents the representation vector of the i-th template generated by the network, z query A representation vector representing the query image.

[0086] The posture corresponding to the template with the highest similarity is the posture corresponding to the actual aircraft. In addition, the network is trained using the method of semantic and similarity joint constraints proposed above.

[0087] In another exemplary embodiment of the present application, the style transfer network used in the present application includes an encoder, a mask-based adaptive multi-head attention normalization (Mask-AdaMHAttN) module and a decoder. The structure of the encoder and the structure of the decoder are mirror-symmetrical.

[0088] The output features of the encoder are fed into the decoder through a mask-based adaptive multi-head attention normalization module to obtain a stylized image.

[0089] For example, Figure 3As shown in the figure, when the encoder is VGG19, the 2×2 maximum pooling layer is used as a downsampling layer in the encoder. Correspondingly, the decoder uses bilinear interpolation for upsampling, and the decoder uses a bilinear interpolation layer with scale=2 to gradually restore the size. The convolution layer of the encoder is divided into four stages by the maximum pooling layer, and the output feature map size of each stage is 64, 128, 256 and 512 respectively. The three Mask-AdaMHAttN modules migrate style features at different scales to the content feature map.

[0090] The style and content features output by the encoder's convolutional layer 4_1 are fed into the decoder through the Mask-AdaMHAttN module. In addition, two skip connections are added after the convolutional layers 2_2 and 3_4, respectively, to migrate the features of the stylized image to the feature map of the content image through the Mask-AdaMHAttN module. The masks of the stylized image and the content image are simultaneously fed into the Mask-AdaMHAttN module to select the feature region of the target. Finally, the migrated features are passed to the decoder through the splicing operation, as follows:

[0091] M fuse =conv(cat(M cs ,M out )).

[0092] Where M fuse represents the fused features, M cs represents the output of the Mask-AdaMHAttN module, M out Represents the output of the previous layer, cat represents the concatenation operation in the channel dimension, and conv represents the convolutional layer.

[0093] The decoder finally outputs a stylized image of the same size as the input content image.

[0094] Furthermore, AdaAttN introduces an attention module to align the feature distribution point by point. However, the attention feature map is still calculated globally, which will introduce unnecessary background style for the style transfer of the aircraft and is not economical. In addition, AdaAttN only uses one attention map to transfer the style, which will miss the features of the aircraft area that is not paid attention to. Therefore, this application adopts Figure 4 The Mask-AdaMHAttN module shown in Figure 2. The multi-head attention module in the Mask-AdaMHAttN module contains N Figure 5 The single-head adaptive attention module shown. The input of each single-head adaptive attention module is the selected style content feature, and it generates query (Q), key (K) and value (V) through a linear layer. The final output is the mean and variance. Figure 5 middle, represents the features belonging to the aircraft selected according to the style mask, Represents the features belonging to the aircraft selected according to the content mask. 2 represents the square of the mean, S i represents the standard deviation of the output of the i-th attention head, M i represents the mean of the output of the i-th attention head.

[0095] The Mask-AdaMHAttN module introduces Mask to transfer the style of the aircraft area, and introduces a multi-head attention module to focus on the features of different aircraft areas through multiple attentions. The specific operations are as follows:

[0096] The content feature map and style feature map are represented as F c and F s , and assuming that they are equal in size, C,H M ,W M Respectively represent the number of channels, width and height of the feature map. The Mask of the content and styled image is represented as W B ,H B Indicates the size of the input Mask, which is consistent with the input image size. Based on this, first c ,Mask s Plastic and F c and F s The same size. Then the feature vector of the corresponding position in the feature map is extracted according to the mask. i and j both represent the coordinates of the aircraft in the mask, and retain the position index corresponding to each feature vector. The feature vectors of content and style are concatenated into a feature vector group N c and N s Represent the total number of extracted content and style vectors respectively.

[0097] After obtaining the feature graph containing only aircraft features, we calculate the query (Q), key (K), and value (V), and we have:

[0098] Q=linear(F c )

[0099] K=linear(F s ).

[0100] V=linear(F s )

[0101] Where linear is a linear layer operating in the channel dimension, The attention map A is calculated as:

[0102] A=Softmax(Q*K T ).

[0103] In the formula,

[0104] According to the attention map, the point-by-point mean and variance are calculated:

[0105] M=A*V.

[0106]

[0107] In the formula, represents the point-wise mean, represents the point-wise variance.

[0108] The above is the calculation for a single attention map, which has limited representation capabilities. To overcome this problem, we can further introduce multi-head attention and calculate the point-by-point weighted mean M through multiple attention maps. multi and variance S multi ,have:

[0109] M multi =linear(Concat(M 0 ,M 1 ,…,M mh )).

[0110] S multi =linear(Concat(S 0 ,S 1 ,…,S mh )).

[0111] Where M i ,S i (i=0,1,...,mh) represents the calculated i-th weighted mean and variance, and mh represents the number of multi-head attention graphs. Then, use M multi and S multi Perform point-by-point weighting on the content feature map to generate the migration feature map F cs :

[0112] F cs =S multi Norm(F c )+M multi .

[0113] According to the position index of Mask, F cs Insert it into the corresponding position in the input feature map, and do not operate on other feature vectors in non-aircraft areas.

[0114] Through the above process, each point of the aircraft in the content feature map will be matched with the most similar semantic point, thereby ensuring the correspondence of local styles.

[0115] Based on the above description, the calculation process of the mask-based adaptive multi-head attention normalization module is shown in Table 1.

[0116] Table 1 Calculation flow chart of the mask-based adaptive multi-head attention normalization module

[0117]

[0118] In another exemplary embodiment of the present application, in order to obtain a more realistic stylized effect, the synthetic and real aircraft images are highly semantically correlated, and it is not only desired to transfer the global style of the aircraft, but also to have the local semantics and style correspond to each other. Based on this, in the present application, the style transfer network adopts a loss function that simultaneously constrains the consistency of neighboring content, the consistency of the overall style, and the consistency of the local style. That is, the loss function of the style transfer network includes three parts: global style loss, local style loss, and neighborhood content consistency loss.

[0119] (1) Global style loss.

[0120] The global style loss does not directly constrain the style similarity between the generated image and the stylized image, but rather aligns the generated image (the stylized virtual image) I g With stylized image I s The mean and variance between the feature maps are used to ensure the overall style of the generated image. Based on this, we have:

[0121]

[0122] in, represents the global style loss, μ is the mean function, σ is the variance, and They represent the generated feature map and style feature map output by the i-th network layer of the encoder respectively. N represents the number of selected network layers.

[0123] In order to ensure the overall style effect, the mean and variance between the content and style feature maps output by the encoder's convolutional layer 1_1, convolutional layer 2_1, convolutional layer 3_1, and convolutional layer 4_1 are calculated.

[0124] (2) Local style loss.

[0125] The goal of the local style loss is to constrain the points in the generated image to be the same in style as the points in the stylized image with similar semantics. Although the aircraft in the virtual image and the real image have strong semantic correlation both in the whole and in the part, they are not point-by-point corresponding due to the different postures of the aircraft. In this case, it is necessary to search for the most similar points in the matching generated image and the stylized image, and constrain the similarity between the two to be the largest. In actual training, the above goal is generally achieved by constraining the features output by the encoder, where:

[0126] First, the similarity between the features is characterized by calculating the cosine distance between the features point by point, which is:

[0127]

[0128] For the convenience of expression, x is used in the formula i and j Respectively represent the generated feature maps and style feature map The i-th feature point and the j-th feature point in d ij Represents x i and j The cosine distance between the two. i and j The smaller the cosine distance between them, the more similar they are. Further normalizing the cosine distance, we have:

[0129]

[0130] represents the normalized cosine distance, d ik Indicates the distance between the i-th feature point and the k-th feature, min k represents the minimized distance d ik , ε represents a fixed constant, which is prevented from being 0, ε=1e -5 .

[0131] Convert the distance into probability w through the exponential function ij ,have:

[0132]

[0133] Among them, h>0 is the bandwidth parameter.

[0134] Finally, in order to prevent instability caused by scale changes, further normalization is performed to transform it into a scale-invariant version:

[0135]

[0136] After the above operations, CX ij Describes xi and j The similarity between ik represents the probability between the i-th feature point and the k-th feature. In order to optimize the network parameters as a loss function, x is converted to i and j The similarity between them is converted to:

[0137]

[0138] in, Represents the local style loss. CX(F g ,F s )∈[0,1], the larger its value is, the higher the similarity between the two.

[0139] Since CX(F g ,F s ) calculates the difference between the most similar points between images, thereby increasing the similarity between the generated image and the stylized image, thereby achieving style transfer.

[0140] (3) Neighborhood content consistency loss.

[0141] In the case of randomly pairing content and stylized images, local style loss can select the most similar ones from the global features for alignment, but local style loss cannot effectively guarantee the invariance of neighborhood content. Therefore, neighborhood content consistency loss can be introduced to constrain the similarity between features between adjacent positions.

[0142] First, generate the feature map F g Randomly select N eigenvectors from Where x = 1, 2, ..., N, each vector is used as the center vector in a neighborhood, and 8 nearest neighbor vectors are selected from its surroundings, denoted as Where y = 1, 2, ..., 8. Similarly, at the same position from the content feature map F c Sampling to obtain feature vector and Calculate the difference between the center vector and the nearest neighbor vector:

[0143]

[0144] in, and The difference between the feature vectors at the x and y positions of the generated feature and the content feature, respectively, is called the difference vector. The difference vector at the same position is defined as a positive pair between the stylized feature and the content feature, otherwise it is a negative pair. The difference vector is mapped and normalized to the unit sphere through two MLP layers. The InfoNCE contrast loss is used to shorten the distance between positive pairs and increase the distance between negative pairs. The neighborhood content consistency loss is:

[0145]

[0146] in, represents the neighborhood content consistency loss, τ represents the temperature hyperparameter, and the default value is 0.07. and The difference between the feature vectors representing the generated features and the content features.

[0147] Based on the above description, the loss function is expressed as:

[0148]

[0149] In the formula, represents the total loss, and α, β, and γ represent the hyperparameters that balance various losses.

[0150] In another exemplary embodiment of the present application, the posture measurement evaluation index mainly uses the posture error to evaluate the performance of the posture estimation model. The posture error is the difference between the predicted posture matrix R and the actual posture matrix The calculated angular error between

[0151]

[0152] The solution obtained Referring to previous methods, the template matching method evaluation experiment reported the proportions of posture errors less than 5°, 10°, 15° and 20° (corresponding to Acc5, ACC10, Acc15, Acc20 in Table 3), while the 2D-3D matching pose measurement method evaluation experiment only reported the proportion of posture errors less than 20°.

[0153] In another exemplary embodiment of the present application, in addition to using the mask generated by the FOUND method, the mask obtained by the supervised segmentation model can also be used for style transfer for comparison. Figure 6 It can be seen that the stylization differences of these different masks are small. In addition, compared with other style transfer methods, the method proposed in this application has the best overall stylization effect.

[0154] In order to reduce background interference, all methods use Mask to remove the background area of ​​the input real image. In addition, CAP-VST (Content Affinity Preserved Versatile Style Transfer), DSTN (Domain-aware Style Transfer Networks) and WCT2 (whitening and coloring transform) are all Mask-based versions to transfer the aircraft area. Based on this, the results of the comparison in different dimensions are as follows:

[0155] (1) Qualitative comparison. Figure 6 The stylized images of various methods are shown in Figure 1. Although CAP-VST, WCT2, StyA2K (styleAll-to-key) and SCTNet (SCTNet Simple Covariance Transformation) have transferred some real aircraft styles, they still retain significant virtual styles, and the overall visual effect is still unreal. StyA2K and CCPL obviously blur the local detail features, while AdaAttN (AdaptiveAttentionNormalization) and DSTN cause a lot of content loss, and the visual effect deviates significantly from the real and virtual. The style transfer method based on the diffusion model is mainly used for artistic style transfer. Figure 6 The paper shows a realistic airplane image based on the diffusion model StyleID (Style Injection in Diffusion), and the visual effect is significantly inferior to the method proposed in this application. The visual effect of the style transfer using unsupervised segmentation Mask in this application is slightly worse than that using supervised segmentation Mask, but it is still significantly better than other methods.

[0156] (2) Quantitative comparison. Referring to previous methods, the SIFID index, style loss and content loss are used to quantitatively evaluate the style transfer effect, as shown in Table 2. The SIFID of the unsupervised version of the method proposed in this application is slightly higher than that of the supervised version, but it is still much lower than the known methods. The style loss of the proposed transfer method is also the smallest, which shows that the method proposed in this application can effectively reduce the difference between the generated image and the real image. The content loss of the method proposed in this application is greater than that of WCT2 and CAP-VST. These two methods do better preserve the content details, but they also retain too much style of the content image, and the authenticity effect is poor. In summary, the method proposed in this application not only better preserves the characteristics of the content image, but also achieves a more realistic stylization effect.

[0157] Table 2 Quantitative comparison of this application and other style transfer methods

[0158]

[0159] In another exemplary embodiment of the present application, the aircraft posture estimation effect is evaluated.

[0160] (1) Evaluation of template matching methods.

[0161] Based on Figure 2 The framework shown in Figure 3 evaluates the matching accuracy of the template matching pose estimation model trained with different stylized datasets, as shown in Table 3. In Table 3, Acc5, Acc10, Acc15, and Acc20 represent the proportions of rotation errors less than 5, 10, 15, and 20 degrees, respectively. Although the stylized datasets generated by the previous methods can achieve matching accuracy far better than the synthetic datasets, they are significantly lower than the test results of the real data. The method of the method has achieved matching accuracy that exceeds the real data regardless of whether it uses supervised or unsupervised masks.

[0162] The best accuracy can be achieved, on the one hand, because the difference between synthetic images and real images is significantly reduced, and on the other hand, because the template matching model relies on the network to generate differentiated representations for accurate matching, and more realistic images involved in training improve the model's representation ability and help generate more discriminative representations.

[0163] Table 3 Test results of template matching methods trained on different stylized datasets

[0164]

[0165] In another exemplary embodiment of the present application, a style transfer ablation experiment is performed. In this embodiment, the visual effect of style transfer without a certain loss function training is explored. SCTNet is the key to keeping the content features unchanged. Without CCPL loss, a large amount of content information will be lost. In the absence of neighboring local style loss, the style of the generated image is closer to the style of the synthesized image. The global style loss can reduce the inconsistency of local styles. Quantitative experimental results show that when the three losses are trained simultaneously, SIFID and style loss obtain the minimum value.

[0166] like Figures 7 to 9 As shown in the figure, assuming that the content loss scale is 0.1, the weights of two loss functions are fixed, and the weight of the other loss function is set from small to large to {1, 5, 10, 15, 20} to explore the impact of loss functions with different weights on style transfer indicators. Fig.10 The effect of different weights of loss on style transfer is shown. An increase in the difference in the ratio between the global style loss or the local style loss and the content loss will lead to stylization failure, while an increase in the weight of the content loss will significantly weaken the style transfer to the content image.

[0167] The hyperparameters α, β, and γ represent the weights of three different loss terms. Fig.10 As shown in the figure, the effect of changing the weight of one of the losses from small to large on the style transfer effect is observed, while the coefficients of the other two loss functions remain fixed. When the weights of the global style loss and the local style loss increase, the information in the content image will be lost. This is mainly because the excessive weights of these two style losses will limit the maintenance of the content loss on the content image features. On the contrary, the increase in the weight of the content loss will keep the original style more and more unchanged, which limits the role of the style loss. Therefore, it is necessary to balance the weights of these three losses to obtain the best style transfer effect.

[0168] The SIFID, style loss, content loss and visual effects of the stylized image are combined to determine α, β and γ. The values ​​of these three parameters are set to {1, 5, 10, 15, 20} respectively, with a total of 125 parameter combinations. Figure 7-Figure 9 The curves of SIFID, style loss, and content loss are shown as the coefficients increase when the other two parameters are fixed. The increase of α and β will not only increase the style loss and SIFID, but also lose the content information, which is consistent with the stylized image. The increase of γ will reduce the content loss, while both SIFID and style loss will increase. Considering the visual effect of style transfer and the values ​​of SIFID, style loss, and content loss, α, β, and γ are finally set to 5, 5, and 15.

[0169] In summary, this application proposes a new unsupervised pose estimation model training framework based on style transfer method. The framework first obtains the mask of the real aircraft image through an unsupervised segmentation method, and then the proposed style transfer network uses the mask to generate a realistic aircraft dataset to train the pose estimation model, which does not rely on any manually annotated pose, mask and other labels. A style transfer method is also proposed, which maintains information of different scales through three jump connections from the encoder to the decoder, and the proposed mask-based adaptive multi-head attention normalization module is used on three paths to transfer style features to the content feature map. In addition, the training loss simultaneously constrains the consistency of local style and the consistency of global style to ensure the stylized effect. Experiments show that the style transfer method proposed in this application can obtain a more realistic stylized effect than previous methods, and under the proposed unsupervised training framework, the pose estimation model trained with the stylized dataset generated by the proposed method has significantly improved accuracy than the pose estimation model trained with the stylized dataset generated by other style transfer methods, which can significantly reduce the demand for real labeled data and promote the application of neural network pose estimation methods in different real scenes.

[0170] Furthermore, compared with the prior art, the present application has the following advantages:

[0171] (1) This application designs an unsupervised aircraft pose estimation training framework based on the style transfer method. This framework only requires rendering data and unlabeled real images to complete the training of the pose estimation model, without relying on any manually labeled labels.

[0172] Among them, this framework only needs real aircraft images without their annotation information to complete the training of the pose estimation model. First, multiple synthetic images of aircraft are obtained through rendering, and their poses, masks, etc. can also be obtained together. At the same time, the masks of the aircraft in the real images are obtained through an unsupervised segmentation model. Then, the style transfer network is trained using the rendered and real images and their masks, and then the trained network is used to generate stylized virtual images. Finally, the stylized virtual images and their annotation information are used to train the pose estimation model.

[0173] (2) This application proposes a novel style transfer method from virtual to real for aircraft targets. The proposed style transfer method transfers multi-scale style information through a mask-based adaptive multi-scale attention normalization module and ensures semantic correspondence through the loss of local and global style consistency.

[0174] Moreover, the proposed style transfer method is targeted at aircraft targets, which can minimize the interference of the background and obtain more realistic stylized images. The design of an unsupervised pose estimation training framework has significant advantages in accuracy, applicability and scalability, which can significantly reduce the demand for labeled data in real scenes and promote the application of pose estimation methods in a wider range of scenarios.

[0175] (3) The accuracy of the pose estimation model trained based on the proposed framework in this application is also significantly higher than previous style transfer methods.

[0176] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Fig.11 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store aircraft posture estimation data from virtual to real. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for estimating the aircraft posture from virtual to real is implemented.

[0177] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0178] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0179] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0180] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0181] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0182] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0183] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0184] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for estimating aircraft posture from virtual to real, characterized in that: The method for estimating the aircraft posture from virtual to real includes: Acquire aircraft images; Using a semantic segmentation network to obtain a mask of the aircraft image; Using a rendering program based on a 3D model of the aircraft, obtaining synthetic aircraft images in different postures and masks corresponding to the synthetic aircraft images; Inputting the aircraft image, the mask of the aircraft image, synthetic aircraft images in different postures, and the mask corresponding to the synthetic aircraft image into a style transfer network to obtain a stylized image; Inputting the stylized image into a pose estimation model to train the pose estimation model; The trained pose estimation model is used to estimate the pose of the aircraft in the real aircraft image.

2. The method for estimating an aircraft posture from virtual to real according to claim 1, characterized in that: The training process of the style transfer network includes: Obtain an airplane image sample, and use the semantic segmentation network to obtain a mask of the airplane image sample; Using a rendering program, based on a 3D model of an aircraft, to obtain rendering data; the rendering data includes: synthetic aircraft image samples in different postures and masks and posture information corresponding to the synthetic aircraft image samples; The style transfer network is trained using the aircraft image sample, the mask of the aircraft image sample, and the rendering data.

3. The method for estimating an aircraft posture from virtual to real according to claim 2, characterized in that: The training process of the pose estimation model includes: Inputting the aircraft image sample, the mask of the aircraft image sample, the synthesized aircraft image sample, and the mask corresponding to the synthesized aircraft image sample into the style transfer network to obtain a stylized image sample; The pose estimation model is trained using the stylized image samples and pose information in the rendering data.

4. The method for estimating aircraft posture from virtual to real according to claim 1, characterized in that: The style transfer network includes an encoder, a mask-based adaptive multi-head attention normalization module and a decoder; The structure of the encoder and the structure of the decoder are mirror-symmetrical; The output features of the encoder are sent to the decoder through the mask-based adaptive multi-head attention normalization module to obtain a stylized image.

5. The method for estimating aircraft posture from virtual to real according to claim 4, characterized in that: VGG19 is used as the encoder.

6. The method for estimating aircraft posture from virtual to real according to claim 1, characterized in that: The style transfer network uses a function that constrains neighboring content consistency, global style consistency, and local style consistency as a loss function.

7. The method for estimating an aircraft posture from virtual to real according to claim 6, characterized in that: The loss function is expressed as: L total =αL g +βL cx +γL ccpl ; Where, L total represents the total loss, L g represents the global style loss, α represents the hyperparameter for balancing the global style loss, and L cx represents the local style loss, β represents the hyperparameter for balancing the local style loss, and L ccpl represents the neighborhood content consistency loss, and γ represents the hyperparameter that balances the neighborhood content consistency loss.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for estimating an aircraft pose from virtual to real according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for estimating an aircraft posture from virtual to real according to any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for estimating an aircraft posture from virtual to real according to any one of claims 1 to 7 is implemented.