Object 3D pose estimation method based on unsupervised domain adaptation
Through the unsupervised domain adaptation method, iterative optimization is used to use aircraft three-dimensional models and real images to generate multi-scale attitude prototypes, solving the problem of limited attitude estimation methods in runway safety applications in the prior art, and realizing accurate attitude estimation and automatic safety hazard detection under limited labeling data.
Patent Information
- Application Number
- CN202210720132.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-06-23
AI Technical Summary
The prior art has limited attitude estimation methods for aircraft and other targets in runway safety applications, making it difficult to achieve accurate attitude estimation with limited labeling data.
A target 3D pose estimation method based on unsupervised domain adaptation is adopted. By acquiring the aircraft three-dimensional model and real images, the model is pre-trained and iteratively optimized using the backbone network to generate a multi-scale pose prototype for pose estimation.
Under the condition of limited labeling data, accurate estimation of the target attitude of the runway area is achieved, and safety hazards such as runway deviation can be automatically detected, improving the safety protection capabilities of the airport runway.
Smart Images

Figure CN115098944B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target pose estimation, and in particular to a target 3D pose estimation method based on unsupervised domain adaptation. Background Art
[0002] Safety is the lifeline of civil aviation. According to the accident distribution data of different flight phases between 1999 and 2019, more than 50% of civil aviation accidents occurred during the take-off and landing of aircraft. As one of the most important ground places for aircraft take-off and landing, the risk index of safety accidents on the runway is very high. Among them, runway deviation is one of the important factors causing runway safety accidents. Effective monitoring of runway deviation events is an important means to ensure civil aviation safety and improve flight operation efficiency. Commonly used scene monitoring sensor equipment generally senses the target orientation information through airborne equipment such as magnetic compasses and gyroscopes, and then judges in advance abnormal situations such as deviation and departure from the runway of the monitored aircraft. However, the magnetic compass is easily interfered by electromagnetic fields, and the gyroscope will cause error accumulation when the magnetic compass is interfered with.
[0003] Image-based target pose estimation methods are not affected by external environments such as electromagnetic interference. Existing image-based target pose estimation methods are mainly divided into two categories: (1) Key point-based pose estimation methods, which first estimate the target key points, then use the mapping relationship between the three-dimensional model and the two-dimensional key points, use the PnP method to calculate the rotation matrix, and solve the target Euler angle based on the rotation matrix; (2) Regression-based methods, that is, using image information, through deep neural networks and learned deep models to directly regress the target 3D pose information.
[0004] Existing methods have the following problems: (1) Keypoint-based pose estimation methods rely heavily on the accuracy of keypoint detection and the accuracy of 3D models; (2) Regression-based pose estimation methods have high accuracy but require a large amount of training data, and high-precision target pose truth is difficult to obtain. Existing public datasets are usually applied to specific scenarios, such as face pose estimation for face recognition, human pose estimation for pedestrian monitoring, and vehicle pose estimation for autonomous driving applications. At present, there are very limited pose estimation methods for targets such as aircraft for runway safety applications, which makes it difficult to obtain labeled data. Summary of the invention
[0005] In view of the defects in the prior art, the present invention provides a target 3D pose estimation method based on unsupervised domain adaptation, which can achieve accurate estimation of the target pose in the runway area under the condition of limited labeled data.
[0006] The present invention provides a method for estimating a target 3D posture based on unsupervised domain adaptation, comprising the following steps:
[0007] S1, obtaining a three-dimensional model of an aircraft, projecting the three-dimensional model of the aircraft onto an image, obtaining a composite image and a posture label corresponding to the composite image; inputting the composite image and the posture label corresponding to the composite image as training data into a backbone network for model pre-training, and obtaining an initial model;
[0008] S2, acquiring a real image, and setting a mixed image to include the synthetic image and the real image;
[0009] S3, inputting the real image into the initial model to obtain a pseudo pose label corresponding to the real image;
[0010] S4, obtaining a multi-scale posture prototype based on the posture label corresponding to the synthetic image, the pseudo posture label corresponding to the real image and the posture label corresponding to the mixed image by statistical calculation; the posture label corresponding to the mixed image includes the posture label corresponding to the synthetic image and the pseudo posture label corresponding to the real image;
[0011] S5, setting the input image to include the synthetic image, the real image and the mixed image; using the input image and the label corresponding to the input image to train the initial model to obtain an optimized initial model; the label corresponding to the input image includes the posture label corresponding to the synthetic image, the pseudo posture label corresponding to the real image and the posture label corresponding to the mixed image;
[0012] S6, looping steps S3-S5 for a preset number of times to obtain an optimized model;
[0013] S7, inputting the real image into the optimization model to obtain a posture estimation result.
[0014] Preferably, the step S1 specifically includes:
[0015] Sample the Euler angle in the preset posture space to obtain the sampling value of the posture space;
[0016] estimating a target rotation matrix in combination with the sampled values;
[0017] Obtain the three-dimensional model of the aircraft, and estimate the mapping matrix from the three-dimensional model of the aircraft to the two-dimensional image based on the assumption that the center point of the aircraft projected onto the image is consistent with the center point of the image;
[0018] Rotate the three-dimensional aircraft model according to the target rotation matrix using the angle value of the preset attitude space, and project the rotated three-dimensional aircraft model onto the two-dimensional image according to the mapping matrix to obtain a composite image and an attitude label corresponding to the composite image;
[0019] The composite image and the posture label corresponding to the composite image are input into the backbone network as training data for model pre-training to obtain an initial model.
[0020] Preferably, the step S1 further comprises: rendering the composite image using a Blender Render algorithm.
[0021] Preferably, the step S3 specifically includes:
[0022] Inputting the real image into the initial model to obtain a pseudo-pose label corresponding to the real image, performing feature extraction on the pseudo-pose label to obtain a feature vector of the real image;
[0023] Inputting the feature vector of the real image into an unordered multi-classifier to obtain a first probability density value;
[0024] Inputting the feature vector of the real image into a sequential binary classifier to obtain a probability density value of a multi-scale representation;
[0025] The pseudo-pose labels corresponding to the real image are screened based on the first probability density value and the probability density value of the multi-scale representation to obtain cleaned pseudo-pose labels.
[0026] Preferably, the step S4 specifically includes:
[0027] Using the pseudo-pose label corresponding to the real image to count the pose prototype in the real image domain, using the pose label corresponding to the synthetic image to count the pose prototype in the synthetic image domain, and using the pose label corresponding to the mixed image to count the pose prototype in the mixed image domain;
[0028] The image domain is set to include a real image domain, a synthetic image domain and a mixed image domain;
[0029] All images belonging to the k-th angle interval in the image domain are counted, and the average feature vectors of all the images are calculated to obtain the posture prototype corresponding to the k-th angle interval, thereby obtaining the multi-scale posture prototype in the image domain.
[0030] Preferably, in step S5, using the input image and the label corresponding to the input image to train the initial model to obtain the optimized initial model specifically includes:
[0031] Using the input image and the label corresponding to the input image to train the initial model, to obtain a posture value of the input image;
[0032] Calculating an objective function using the multi-scale posture prototype and the posture value of the input image;
[0033] The objective function is optimized using gradient descent and the initial model is updated to obtain an optimized initial model.
[0034] Preferably, the objective function includes: cross entropy loss of discrete posture representation, mean square error of continuous posture representation and prototype distances between the original domain, target domain and hybrid domain at different scales and angle intervals.
[0035] Preferably, the screening of the pseudo-gesture labels corresponding to the real image based on the first probability density value and the probability density value of the multi-scale representation includes:
[0036] The angle estimation value of the first probability density value is converted into a hot single-hot code, and then the hot single-hot code is converted into a Gaussian space to obtain a corresponding posture soft label, the cosine distance between the posture soft label and the probability density value is calculated, and the pseudo posture label is screened according to the cosine distance.
[0037] Preferably, the screening of the pseudo-gesture labels corresponding to the real image based on the first probability density value and the probability density value of the multi-scale representation further includes:
[0038] Converting the posture estimation value output by the sequential binary classifier into a zero-one vector to obtain a continuous estimation value zero-one vector;
[0039] The Euclidean distance between the continuous estimated value zero-one vector and the zero-one vector output by the sequential binary classifier is calculated, and the pseudo-gesture label is screened according to the Euclidean distance.
[0040] The beneficial effects of the present invention are:
[0041] Under the condition of limited annotated data, it is possible to accurately estimate the target posture in the runway area and automatically detect potential safety hazards such as runway deviation. The present invention is applied to the runway safety monitoring system, which can effectively warn / alert the deviation target, greatly improve the airport runway safety protection capability, and enhance the level of smart air traffic control construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the specific embodiments or the description of the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.
[0043] Figure 1 Schematic diagram of a flow chart of an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0046] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0047] It should be further understood that the term "and / or" used in the present specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0048] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0049] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should have the common meanings understood by those skilled in the art to which the invention belongs.
[0050] like Figure 1 As shown, an embodiment of the present invention provides a method for estimating a target 3D posture based on unsupervised domain adaptation, comprising the following steps:
[0051] S1, obtaining a three-dimensional model of the aircraft, projecting the three-dimensional model of the aircraft onto an image, obtaining a synthetic image and a posture label corresponding to the synthetic image; inputting the synthetic image and the posture label corresponding to the synthetic image as training data into a backbone network for model pre-training, and obtaining an initial model;
[0052] S2, obtaining a real image, and setting the mixed image to include a synthetic image and a real image;
[0053] S3, input the real image into the initial model to obtain the pseudo pose label corresponding to the real image;
[0054] S4, obtaining a multi-scale posture prototype based on the posture label corresponding to the synthetic image, the pseudo posture label corresponding to the real image and the posture label corresponding to the mixed image by statistical calculation; the posture label corresponding to the mixed image includes the posture label corresponding to the synthetic image and the pseudo posture label corresponding to the real image;
[0055] S5, setting the input image to include a synthetic image, a real image, and a mixed image; using the input image and the label corresponding to the input image to train the initial model to obtain an optimized initial model; the label corresponding to the input image includes a posture label corresponding to the synthetic image, a pseudo posture label corresponding to the real image, and a posture label corresponding to the mixed image;
[0056] S6, looping steps S3-S5 for a preset number of times to obtain an optimized model;
[0057] S7, input the real image into the optimization model to obtain the pose estimation result.
[0058] Among them, the backbone network includes 5 convolutional layers, 5 maximum pooling layers and 3 fully connected layers, which are used to extract high-level semantic features of the image.
[0059] In the embodiment of the present invention, the model is iteratively trained using synthetic images and their corresponding posture labels, synthetic images and their corresponding pseudo-pose labels, and mixed images and their corresponding posture labels, so that accurate estimation of the target posture in the runway area can be achieved under the condition of limited annotated data.
[0060] Step S1 specifically includes:
[0061] Sample the Euler angle in the preset posture space to obtain the sampling value of the posture space;
[0062] Estimate the target rotation matrix based on the sampled values;
[0063] Obtain the three-dimensional model of the aircraft, and estimate the mapping matrix from the three-dimensional model of the aircraft to the two-dimensional image based on the assumption that the center point of the aircraft projected onto the image is consistent with the center point of the image;
[0064] The 3D model of the aircraft is rotated according to the target rotation matrix by using the angle value of the preset attitude space, and the rotated 3D model of the aircraft is projected onto the 2D image according to the mapping matrix to obtain a composite image and an attitude label corresponding to the composite image;
[0065] The synthetic image and the posture label corresponding to the synthetic image are input into the backbone network as training data for model pre-training to obtain the initial model.
[0066] In the embodiment of the present invention, a reference coordinate system is first established based on the aircraft structure. The horizontal line connecting the left and right wings is set as the x-axis, the horizontal line connecting the tail to the nose (the aircraft's forward direction) is set as the z-axis, and the vertical straight line y-axis is set as the vertical straight line. represents the target 3D pose, where: represents the pitch angle (Pitch), which rotates around the x-axis; θ represents the yaw angle (Yaw), which rotates around the y-axis; ω represents the roll angle (Roll), which rotates around the z-axis. Define the scale factor δ, assume that the angle interval [-V, V] is evenly divided into intervals of δ, and define the Euler c ={-V, -V+δ, ..., 0, ..., V-δ, V} is the attitude characterization basis based on the scale factor δ. In theory, the scale factor δ directly affects the size of the attitude characterization basis. When the scale factor δ is large, the corresponding angle range is large and the attitude description accuracy is low; when the scale factor δ is small, the model becomes larger and the computational complexity increases due to the increase in the number of estimated parameters.
[0067] The present invention is based on a multi-scale discrete posture representation method, which can comprehensively balance accuracy and efficiency. For the Euler angle α∈[-V, V], l=3 scales are used to represent the target posture respectively, and the coefficient ratio between adjacent scales is s=4, then the scale coefficient corresponding to the i∈[1,l]th scale is δ i =2V / s i ,use represents the posture representation basis corresponding to the i∈[1,l]th scale, where Corresponding to the angle interval on the i∈[1,l]th scale. Based on the above assumptions, the continuous space attitude value can be calculated by the following formula
[0068]
[0069] in Represents the estimated probability value of the j-th angle interval on the i-th scale.
[0070] Next, the Euler angles are sampled in the attitude space, and the sampling range includes -π to π in the Yaw direction, -π / 2 to π / 2 in the Pitch direction, and -π / 2 to π / 2 in the Roll direction. Since the main indicator of runway deviation is to determine whether the orientation of the aircraft in the Yaw direction is consistent with the runway direction, in the embodiment of the present invention, the sensitivity to attitude changes in the Yaw direction is higher than that in the other two directions, and the sampling interval in the Yaw direction is smaller than that in the Pitch direction and the Roll direction. In addition, assuming that the direction in which the target slides along the runway is (α, β, γ), an angle tolerance vector (θα ,θ β ,θ γ ), when the target yaw direction is toward (α-θ α ,α+θ α ) range, the pitch direction is in the range of (β-θ β ,β+θ β ) range, the roll direction is in the range of (γ-θ γ ,γ+θ γ ), set a lower sampling interval.
[0071] Acquiring the aircraft three-dimensional model may include acquiring the aircraft three-dimensional model from public 3D model data sets ShapeNetCore and ModelNet, and collecting aircraft CAD models from the Internet including 3D Resource Network (http: / / www.3dyw.com), Model Cloud (http: / / www.moxingyun.com), 3D Academy (http: / / www.3dxy.com) and other websites.
[0072] The target rotation matrix is estimated by combining the sampling values of the above-mentioned attitude space, and the mapping matrix from the aircraft to the image is estimated based on the assumption that the center point of the aircraft is projected onto the image and is consistent with the center point of the image, thereby establishing a mapping relationship between the 3D model and the 2D image. The angular value of the preset attitude space is used to rotate the three-dimensional model of the aircraft according to the target rotation matrix, and the rotated three-dimensional model of the aircraft is projected onto the two-dimensional image according to the mapping matrix to obtain a synthetic image and the attitude label corresponding to the synthetic image. Finally, the synthetic image and the attitude label corresponding to the synthetic image are input into the backbone network as training data for model pre-training to obtain the initial model.
[0073] Step S1 also includes: using the Blender Render algorithm to render the synthesized image, and a highly realistic synthesized image can be obtained. The embodiment of the present invention can also appropriately change the scaling ratio of the projection matrix according to the ratio of the target area area to the image size in the real scene to ensure the diversity of the synthesized image and generate a multi-scale synthesized image.
[0074] Step S3 specifically includes:
[0075] Input the real image into the initial model to obtain the pseudo-pose label corresponding to the real image, perform feature extraction on the pseudo-pose label to obtain the feature vector of the real image;
[0076] Input the feature vector of the real image into the unordered multi-classifier to obtain a first probability density value;
[0077] Input the feature vector of the real image into the sequential binary classifier to obtain the probability density value of the multi-scale representation;
[0078] The pseudo-pose labels corresponding to the real image are screened based on the first probability density value and the probability density value of the multi-scale representation to obtain the cleaned pseudo-pose labels.
[0079] Among them, the unordered multi-classifier includes a fully connected layer and a softmax layer, where the fully connected layer converts the input feature vector to the scale space, and the output vector dimension can be expressed as in represents the dimension of the posture representation basis on the i-th layer, Represents the number of all posture intervals in the i-th layer. The softmax layer further converts the output of the fully connected layer into the corresponding probability density value in the multi-scale posture space to represent the probability distribution of different posture intervals.
[0080] The unordered multi-classifier assumes that the posture labels in the same Euler direction are independent, but in fact, the target posture should change continuously and stably. In this embodiment of the present invention, it is assumed that the two Euler angles in the same Euler direction are θ 1 and θ 2 , according to the continuity of posture change, θ 1 and θ 2 The 1 ≥θ 2 or θ 1 <θ 2 , that is, there is a sequential constraint relationship between different angles in the same Euler direction. The posture sequence estimation method based on sequential binary classification in the embodiment of the present invention utilizes the continuity of posture changes and estimates the relative position of the input target in the posture space by comparing the sizes between different postures, so as to simplify the posture estimation problem. The number of sequential binary classifiers and the dimension of the corresponding multi-scale posture representation basis Consistent. Assumption Representation of posture The kth element in , then the corresponding order binary classifier is used to output whether the estimated angle value is greater than
[0081] For the sequential binary classifier, when the estimated angle α is close to the interval split point When α may belong to the interval It may also belong to the interval There is ambiguity at this time. The embodiment of the present invention introduces the overlap coefficient ε and uses the overlap interval and To improve the robustness of the model near the interval division points.
[0082] In the embodiment of the present invention, the accuracy of pseudo-pose labels can be improved by designing a disordered multi-classifier for the global pose space and a sequential binary classifier focusing on the local pose space to screen pseudo-pose labels.
[0083] Screening the pseudo-pose label corresponding to the real image based on the first probability density value and the probability density value of the multi-scale representation includes:
[0084] The angle estimation value of the first probability density value is converted into a hot single-hot encoding, and then the hot single-hot encoding is converted into a Gaussian space to obtain the corresponding posture soft label, and the cosine distance between the posture soft label and the probability density value is calculated, and the pseudo posture label is screened according to the cosine distance. The larger the cosine distance value, the higher the credibility of the pseudo posture label.
[0085] Screening the pseudo-pose label corresponding to the real image based on the first probability density value and the probability density value of the multi-scale representation also includes:
[0086] Convert the posture estimation value output by the sequential binary classifier into the form of a zero-one vector to obtain a continuous estimation value zero-one vector;
[0087] Calculate the Euclidean distance between the continuous estimated value zero-one vector and the zero-one vector output by the sequential binary classifier, and filter the pseudo-pose labels according to the Euclidean distance. The smaller the Euclidean distance, the higher the credibility of the pseudo-pose label.
[0088] In the embodiment of the present invention, the pseudo-pose labels corresponding to the real image are screened based on the first probability density value and the probability density value of the multi-scale representation, and erroneous labels can be deleted to improve the credibility of the pseudo-pose labels, thereby improving the performance of the model.
[0089] Step S4 specifically includes:
[0090] The pseudo-pose labels corresponding to the real images are used to count the pose prototypes in the real image domain, the pose labels corresponding to the synthetic images are used to count the pose prototypes in the synthetic image domain, and the pose labels corresponding to the mixed images are used to count the pose prototypes in the mixed image domain.
[0091] The image domain is set to include a real image domain, a synthetic image domain and a mixed image domain;
[0092] All images belonging to the kth angle interval in the image domain are counted, and the average eigenvector of all images is calculated to obtain the posture prototype corresponding to the kth angle interval; similarly, the posture prototypes of all angle intervals are obtained, thereby obtaining a multi-scale posture prototype in the image domain.
[0093] In the embodiment of the present invention, step S4 can better characterize the target posture characteristics at different scale levels and improve the robustness of the model.
[0094] In step S5, the initial model is trained using the input image and the label corresponding to the input image to obtain the optimized initial model, which specifically includes:
[0095] The initial model is trained using the input image and the label corresponding to the input image to obtain the posture value of the input image;
[0096] Calculate the objective function using the multi-scale pose prototype and the pose value of the input image;
[0097] Gradient descent is used to optimize the objective function and update the initial model to obtain the optimized initial model.
[0098] During training, a step-by-step refinement method is used for iterative optimization. When updating parameters, the gradient descent method is used to optimize the following objective functions: (1) cross entropy loss based on discrete pose representation; (2) mean square error based on continuous pose representation; (3) prototype distances between the original domain, target domain, and hybrid domain at different scales and angles.
[0099] The target 3D posture estimation method based on unsupervised domain adaptation proposed in the present invention can accurately estimate the target posture in the runway area under the condition of limited annotated data, and realize automatic detection of safety hazards such as runway deviation. The present invention is applied to the runway safety monitoring system, which can effectively warn / alert the deviation target, greatly improve the airport runway safety protection capability, and improve the level of smart air traffic control construction.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.
Claims
1. A method for target 3D pose estimation based on unsupervised domain adaptation, It is characterized in that The steps include: S1, obtaining a three-dimensional model of an aircraft, projecting the three-dimensional model of the aircraft onto an image, obtaining a composite image and a posture label corresponding to the composite image; inputting the composite image and the posture label corresponding to the composite image as training data into a backbone network for model pre-training, and obtaining an initial model; S2, acquiring a real image, and setting a mixed image to include the synthetic image and the real image; S3, inputting the real image into the initial model to obtain a pseudo pose label corresponding to the real image; S4, obtaining a multi-scale posture prototype based on the posture label corresponding to the synthetic image, the pseudo posture label corresponding to the real image and the posture label corresponding to the mixed image by statistical calculation; the posture label corresponding to the mixed image includes the posture label corresponding to the synthetic image and the pseudo posture label corresponding to the real image; S5, setting the input image to include the synthetic image, the real image and the mixed image; using the input image and the label corresponding to the input image to train the initial model to obtain an optimized initial model; the label corresponding to the input image includes the posture label corresponding to the synthetic image, the pseudo posture label corresponding to the real image and the posture label corresponding to the mixed image; S6, looping steps S3-S5 for a preset number of times to obtain an optimized model; S7, inputting the real image into the optimization model to obtain a posture estimation result; The step S3 specifically includes: Inputting the real image into the initial model to obtain a pseudo-pose label corresponding to the real image, performing feature extraction on the pseudo-pose label to obtain a feature vector of the real image; Inputting the feature vector of the real image into an unordered multi-classifier to obtain a first probability density value; Inputting the feature vector of the real image into a sequential binary classifier to obtain a probability density value of a multi-scale representation; The pseudo-pose labels corresponding to the real image are screened based on the first probability density value and the probability density value of the multi-scale representation to obtain cleaned pseudo-pose labels.
2. The method for estimating target 3D pose based on unsupervised domain adaptation according to claim 1, It is characterized in that The step S1 specifically includes: Sample the Euler angle in the preset posture space to obtain the sampling value of the posture space; estimating a target rotation matrix in combination with the sampled values; Obtain the three-dimensional model of the aircraft, and estimate the mapping matrix from the three-dimensional model of the aircraft to the two-dimensional image based on the assumption that the center point of the aircraft projected onto the image is consistent with the center point of the image; Rotate the three-dimensional aircraft model according to the target rotation matrix using the angle value of the preset attitude space, and project the rotated three-dimensional aircraft model onto the two-dimensional image according to the mapping matrix to obtain a composite image and an attitude label corresponding to the composite image; The composite image and the posture label corresponding to the composite image are input into the backbone network as training data for model pre-training to obtain an initial model.
3. The method for estimating target 3D pose based on unsupervised domain adaptation according to claim 1, It is characterized in that The step S1 further includes: rendering the composite image using a Blender Render algorithm.
4. The method for estimating target 3D pose based on unsupervised domain adaptation according to claim 1, It is characterized in that The step S4 specifically includes: Using the pseudo-pose label corresponding to the real image to count the pose prototype in the real image domain, using the pose label corresponding to the synthetic image to count the pose prototype in the synthetic image domain, and using the pose label corresponding to the mixed image to count the pose prototype in the mixed image domain; The image domain is set to include a real image domain, a synthetic image domain and a mixed image domain; All images belonging to the k-th angle interval in the image domain are counted, and the average feature vectors of all the images are calculated to obtain the posture prototype corresponding to the k-th angle interval, thereby obtaining the multi-scale posture prototype in the image domain.
5. The method for estimating target 3D pose based on unsupervised domain adaptation according to claim 1, It is characterized in that In step S5, the initial model is trained using the input image and the label corresponding to the input image to obtain the optimized initial model, which specifically includes: Using the input image and the label corresponding to the input image to train the initial model, to obtain a posture value of the input image; Calculating an objective function using the multi-scale posture prototype and the posture value of the input image; The objective function is optimized using gradient descent and the initial model is updated to obtain an optimized initial model.
6. The method for estimating target 3D pose based on unsupervised domain adaptation according to claim 5, It is characterized in that The objective function includes: cross entropy loss of discrete posture representation, mean square error of continuous posture representation and prototype distances between original domain, target domain and hybrid domain at different scales and angle intervals.
7. The method for estimating target 3D pose based on unsupervised domain adaptation according to claim 1, It is characterized in that The screening of the pseudo-pose label corresponding to the real image based on the first probability density value and the probability density value of the multi-scale representation includes: The angle estimation value of the first probability density value is converted into a hot single-hot code, and then the hot single-hot code is converted into a Gaussian space to obtain a corresponding posture soft label, the cosine distance between the posture soft label and the probability density value is calculated, and the pseudo posture label is screened according to the cosine distance.
8. The method for estimating target 3D pose based on unsupervised domain adaptation according to claim 1, It is characterized in that The filtering of the pseudo-pose label corresponding to the real image based on the first probability density value and the probability density value of the multi-scale representation further includes: Converting the posture estimation value output by the sequential binary classifier into a zero-one vector to obtain a continuous estimation value zero-one vector; The Euclidean distance between the continuous estimated value zero-one vector and the zero-one vector output by the sequential binary classifier is calculated, and the pseudo-gesture label is screened according to the Euclidean distance.
Citation Information
Patent Citations
Domain randomization-based attitude estimation model training method and device
CN111784772A