Cross-domain target attitude estimation adaptive method from virtual domain to real domain

By constructing an intermediate domain data set between the virtual domain and the real domain, and using a dual-branch network and feature fusion module to extract features, combining the pose estimation generator and the least squares fitting algorithm to solve the pose parameters, the problem of target pose estimation in complex and variable real environments is solved, and efficient cross-domain knowledge transfer and generalization capabilities are achieved.

CN119941849APending Publication Date: 2025-05-06SHANXI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411927491.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to achieve efficient target pose estimation in complex and changeable real environments, and deep learning methods rely on a large amount of labeled data, resulting in poor generalization capabilities.

Method used

A cross-domain target pose estimation adaptive method from virtual domain to real domain is proposed. By constructing an intermediate domain data set, features are extracted using a dual-branch network and feature fusion module, pose estimation generator and least squares fitting algorithm are combined to solve pose parameters, and adaptive is achieved through the domain discriminator.

Benefits of technology

Reliance on large amounts of labeled data in the real environment is reduced, domain differences are narrowed, and the generalization ability of the target pose estimation model is improved, and knowledge transfer between virtual domains and real domains is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941849A_ABST
    Figure CN119941849A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and particularly relates to a cross-domain target attitude estimation adaptive method from a virtual domain to a real domain. In order to solve the problems that attitude estimation data is difficult to collect and the labeling cost is high in a complex real environment and aim at improving the attitude estimation precision in the real environment, the method fully utilizes the advantages that synthetic data is easy to collect and label, any possible situation and scene in the real world can be provided, and the attitude estimation precision in the real environment can be improved. According to the method, the generalization and robustness of the model are better improved while the data quantity and quality are improved, a domain self-adaption technology is introduced, knowledge learned in synthetic domain data is self-adapted to a real environment, model migration is completed, and finally self-adaption learning from a virtual environment to the real environment is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a cross-domain target posture estimation adaptive method from a virtual domain to a real domain. Background Art

[0002] With the continuous development of computer vision technology, target pose estimation has become increasingly important in many application scenarios. The target pose estimation task aims to determine the position and orientation of the target object in three-dimensional space, which is extremely critical for achieving accurate positioning and operation. For example, in robot navigation, accurate target pose estimation can help the robot avoid obstacles and reach the destination safely; in industrial automation, robots need to accurately locate and grasp objects to complete complex assembly tasks.

[0003] However, target pose estimation methods based on traditional features usually rely on manually designed features, such as SIFT and SURF. These methods perform well in certain specific scenarios, but often do not work well in complex and changing environments. Deep learning-based methods train models with a large amount of labeled data and can better handle complex scenarios, but their performance is highly dependent on the quality and quantity of training data. Target pose estimation methods based on cross-domain technology have obvious domain differences and weak adaptive capabilities, which makes it difficult for the model to quickly adapt to new environments and has poor generalization capabilities.

[0004] Therefore, the present invention proposes a cross-domain target pose estimation adaptive method from virtual domain to real domain. This method can reduce the dependence of deep learning methods on a large amount of labeled data in the real environment while narrowing the domain differences, thereby achieving effective transfer of knowledge between virtual domain and real domain, and improving the generalization ability of the target pose estimation model, which is of great significance. Summary of the invention

[0005] In order to solve the problem of difficult data collection and expensive labeling in complex real environments, this paper is committed to improving the accuracy of posture estimation in real environments. It can effectively alleviate the problem of data quantity and quality while making up for the inability of existing posture estimation algorithms to balance accuracy and efficiency. A cross-domain target posture estimation method is proposed to complete adaptive learning from virtual environments to real environments.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention provides a cross-domain target posture estimation adaptive method from a virtual domain to a real domain, comprising the following steps:

[0008] Step 1: Build the source domain dataset D s , consisting of n sets of “RGB image-depth image-annotation information” pairs; construct the target domain dataset Dt , consisting of m sets of “RGB image-depth image” pairs, s represents the virtual environment, and t represents the real environment;

[0009] Step 2: Based on the source domain dataset D s and the target domain dataset D t Constructing the intermediate domain dataset D m , by N m The group is composed of "RGB image-depth image" pairs;

[0010] Step 3: The source domain dataset D s , target domain dataset D t and the intermediate domain dataset D m Input cross-domain target posture estimation model; the cross-domain target posture estimation model includes three parts: feature extractor f, target posture estimation module and adaptive module; the feature extractor f adopts a dual-branch network to extract the source domain dataset D s , target domain dataset D t , intermediate domain dataset D m The RGB image I rgb and depth image I depth The characteristic is expressed as F r and F d , using the feature fusion module to fuse F r and F d The characteristic after is F f , whose feature dimension is BCN p , N p is the number of sampled feature points. The target posture estimation module uses the posture estimation generator g to obtain the correspondence of the points, and then uses the least squares fitting algorithm to solve the posture parameters (R, t), where R is the rotation parameter and t is the translation parameter. The adaptive module introduces a domain discriminator to achieve cross-domain.

[0011] Further, the step 1 is specifically as follows:

[0012] Step 1.1, use a 3D scanner to scan the 3D model corresponding to the object in the real world, the format of which is .ply;

[0013] Step 1.2: In the simulation software, import the .ply 3D model, build a scene in the virtual environment, fix the light intensity, write a script to obtain the image and the corresponding annotation information, and obtain the "RGB image-depth image-annotation information" pair, that is, the source domain dataset D s , i=1,2,3,...,n, where n represents the number of data sets in the virtual environment. Represents the RGB image and corresponding depth image of virtual scene i, b iRepresents the pose estimation annotation information in virtual scene i;

[0014] Step 1.3: Fix the target object, build a scene in a real environment, fix the lighting, use a camera to shoot the object for a week, and use the 3D reconstruction method to reconstruct the 3D model of the target object to obtain the "RGB image-depth image" pair, that is, the target domain dataset D t , j = 1, 2, 3, ..., m, where m represents the number of data sets in the real environment. Represent the RGB image and corresponding depth image of the real scene j respectively;

[0015] Step 1.4: Change the target object to be detected and the shooting scene, and repeat the above three steps until a complete data set D is obtained. s and D t .

[0016] Further, the step 2 is specifically as follows:

[0017] The source domain dataset D s and the target domain dataset D t RGB image After transformation MLP_E is converted to the feature space, the source domain and target domain color features are After fusing the RGB information, the image is restored through inverse transformation MLP_D, which is the RGB image in the intermediate domain. As shown below:

[0018]

[0019] Among them, l∈[0,1] refers to the influence degree of the source domain RGB image;

[0020] The source domain dataset D s and the target domain dataset D t Depth image Use weighted fusion to obtain the depth image of the intermediate domain

[0021]

[0022] Among them, m∈[0,1] refers to the influence degree of the source domain depth image;

[0023] At this point, the intermediate domain dataset D is completed m The construction of N m The group is composed of "RGB image-depth image" pairs, which are in the form of k=1,2,3,...,N m , k represents the kth scene in the middle domain, N mIt is the smaller value between the source domain and the target domain data volume.

[0024] Furthermore, in step 3, the least squares fitting algorithm is used to solve the posture parameters (R, t), specifically:

[0025] For the source domain dataset D s , the feature after feature extractor f is F f s , use the pose estimation generator g to obtain the semantic information, center point information and key point information of the target object, and then obtain two point sets of a target object through cluster voting. On this basis, the least squares fitting algorithm is used to solve the pose parameters (R, t). The least squares fitting algorithm is calculated by minimizing the following loss:

[0026]

[0027] in, is the point set of K key points detected in the camera coordinate system, is the set of points they correspond to in the object coordinate system.

[0028] Furthermore, the adaptation module in step 3 includes two processes: adaptation from the source domain to the intermediate domain and adaptation from the intermediate domain to the target domain. The specific process is as follows:

[0029] For the adaptive process from source domain to intermediate domain:

[0030] The source domain dataset D s , intermediate domain dataset D m The data is preprocessed and fed into the feature extractor f;

[0031] After passing through the feature extractor f, the source domain aggregate feature F is obtained f s and the intermediate domain aggregation feature F f m , F f s They are fed to the pose generator g and the domain discriminator d respectively to obtain the pose estimation result and domain label of the source domain. f m Feed it to the domain discriminator d to obtain the domain label of the intermediate domain;

[0032] For the adaptive process from the intermediate domain to the target domain:

[0033] The intermediate domain dataset D m , target domain dataset D t The data is preprocessed and fed into the feature extractor f;

[0034] After passing through the feature extractor f, the intermediate domain aggregate feature F is obtained f m and the target domain aggregate feature F f t , F f m With F f t Feed it to the domain discriminator d to obtain the domain labels of the intermediate domain and the target domain.

[0035] Furthermore, the method also includes training a cross-domain target pose estimation model, wherein the total loss during the training process is divided into two parts: a task loss of the pose estimation part and a domain loss of the adaptation part;

[0036] The task loss of the attitude estimation part includes the following three parts:

[0037]

[0038] L seg =(-t(1-(c i ·l i )) γ log(c i ·l i ))

[0039]

[0040] Among them, i is a random point, p i Represents the coordinate information and feature information of point i, N is the number of vertices on the surface of the object, M is the number of key points of the selected target, of i j , They represent the predicted value and true value of the translation offset from the i-th random point to the j-th key point, Δx i , Δx i * They represent the predicted value and true value of the offset from the i-th random point to the center of the object, I is the target object instance, II is the indicator function, and when p i ∈I is equal to 1 when it holds, otherwise it is equal to 0; t is the balance parameter, γ is the focusing parameter, c i , l i Respectively represent the confidence that the i-th point belongs to the corresponding object category and the one-hot representation of the real object category label;

[0041] The task loss of the pose estimation part is:

[0042] L task =v1L keypoints +v2L seg +v3L center

[0043] Among them, v1, v2, and v3 are the weight factors of the three sub-task losses of key point information, semantic information, and center point information respectively;

[0044] The domain loss of the adaptive part includes the domain loss of the adaptive process from the source domain to the intermediate domain and the domain loss of the adaptive process from the intermediate domain to the target domain:

[0045] For the domain loss of the source domain to intermediate domain adaptation process: Assume that the feature vector set of the source domain is The set of eigenvectors of the intermediate domain is n and z are the data volumes of the source domain and the intermediate domain, respectively. as well as Represents the mean and variance of the two domains, which is used to describe the data distribution of the two domains and introduces the inter-domain alignment loss L align_1 The feature alignment is achieved as shown below:

[0046]

[0047] in, s , m are the covariance matrices of the source domain and the intermediate domain, respectively. F represents the Frobenius norm, and α and b are weight factors.

[0048] Design a domain discriminator d and use the domain adversarial loss L adv_1 Optimize as shown below:

[0049] L adv_1 =L bce (d(F f s ),0)+L bce (d(F f m ),1)

[0050] Among them, L bce represents the binary cross entropy loss function, 0 means the sample comes from the source domain, 1 means the sample comes from the intermediate domain;

[0051] The domain loss of the source domain to intermediate domain adaptation process is:

[0052] L domain_1 =L align_1 +L adv_1 ;

[0053] For the domain loss of the intermediate domain to target domain adaptation process: Assume that the feature vector set of the intermediate domain is The feature vector set of the target domain is z and m are the data sizes of the intermediate domain and the target domain, respectively. as well as Represents the mean and variance of the two domains, which is used to describe the data distribution of the two domains and introduces the inter-domain alignment loss L align_2 The feature alignment is achieved as shown below:

[0054]

[0055] in, s , m are the covariance matrices of the intermediate domain and the target domain, respectively. F represents the Frobenius norm, and α and b are weight factors.

[0056] Design a domain discriminator d and use the domain adversarial loss L adv_2 Optimize as shown below:

[0057] L adv_2 =L bce (d(F f m ),0)+L bce (d(F f t ),1)

[0058] Among them, L bce represents the binary cross entropy loss function, 0 means the sample is from the intermediate domain, and 1 means the sample is from the target domain;

[0059] The domain loss of the adaptive process from the intermediate domain to the target domain is:

[0060] L domain_2 =L align_2 +L adv_2 ;

[0061] The total loss of the cross-domain target pose estimation model is as follows:

[0062] L total =ω1L task +ω2(L domain_1 +L domain_2 )

[0063] Among them, ω1 and ω2 are the balance factors of the loss of the attitude estimation part and the adaptive part respectively;

[0064] The corresponding loss is calculated, and the parameters of the cross-domain target posture estimation model are optimized according to the loss value until convergence, that is, the training of the cross-domain target posture estimation model is completed.

[0065] Compared with the prior art, the present invention has the following advantages:

[0066] 1. Make full use of the advantages of synthetic data that are easy to collect and label, and provide any possible situations and scenarios in the real world, while improving the quantity and quality of data and better improving the generalization and robustness of the model.

[0067] 2. Considering that the actual application goal of the algorithm is to improve the estimation accuracy in real scenarios, domain adaptation technology is introduced to adapt the knowledge learned in synthetic domain data to the real environment to complete the migration of the model.

[0068] 3. The method of the present invention is easy to implement, and its application value is mainly reflected in the following aspects: 1) In the field of intelligent security, this technology can be used to simulate various security threat scenarios, improve the recognition accuracy and response speed of the monitoring system to abnormal behaviors, and make security monitoring more intelligent and efficient. 2) In the medical field, this technology can be used to synthesize medical imaging data, simulate the imaging features of different pathological states and different imaging devices, and effectively solve the problems of scarcity and difficulty in labeling medical data while protecting patient privacy. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is the overall flow chart of the present invention;

[0070] Figure 2 The data acquisition flow chart of the present invention;

[0071] Figure 3 This is a framework diagram of the model proposed in the present invention. DETAILED DESCRIPTION

[0072] In order to further illustrate the technical solution of the present invention, the present invention is further described below through embodiments.

[0073] like Figures 1 to 3 As shown, a cross-domain target posture estimation adaptive method from a virtual domain to a real domain in this embodiment includes the following steps:

[0074] Step 1: Build the source domain dataset D s , consisting of n sets of “RGB image-depth image-annotation information” pairs; construct the target domain dataset D t , which consists of m sets of "RGB image-depth image" pairs, s represents the virtual environment, and t represents the real environment, as follows:

[0075] Step 1.1, use EinScan-SP 3D scanner to scan the 3D model corresponding to the object in the real world, the format of which is .ply;

[0076] Step 1.2: In the simulation software blender, import the .ply 3D model, build a scene in the virtual environment, fix the light intensity, write a script to obtain the image and the corresponding annotation information, and obtain the "RGB image-depth image-annotation information" pair, that is, the source domain dataset. i=1,2,3,...,n, where n represents the number of data sets in the virtual environment. denote the RGB image and the corresponding depth image of the virtual scene i, respectively, and b i Represents the pose estimation annotation information in virtual scene i;

[0077] Step 1.3, fix the target object, build a scene in a real environment, fix the lighting, use a realscene camera to shoot the object for a week, use the 3D reconstruction method to reconstruct the 3D model of the target object, and obtain the "RGB image-depth image" pair, that is, the target domain dataset j = 1, 2, 3, ..., m, where m represents the number of data sets in the real environment. Represent the RGB image and corresponding depth image of the real scene j respectively;

[0078] Step 1.4: Change the target object to be detected and the shooting scene, and repeat the above three steps until a complete data set D is obtained. s and D t .

[0079] Step 2: Based on the source domain dataset D s and the target domain dataset D t Constructing the intermediate domain dataset D m , by N m The group "RGB image-depth image" pair is composed as follows:

[0080] The source domain dataset D s and the target domain dataset D t RGB image After transformation MLP_E is converted to the feature space, the source domain and target domain color features are After fusing the RGB information, the image is restored through inverse transformation MLP_D, which is the RGB image in the intermediate domain. As shown below:

[0081]

[0082] Among them, l∈[0,1] refers to the influence degree of the source domain RGB image;

[0083] The source domain dataset D s and the target domain dataset D t Depth image Use weighted fusion to obtain the depth image of the intermediate domain

[0084]

[0085] Among them, m∈[0,1] refers to the influence degree of the source domain depth image;

[0086] At this point, the intermediate domain dataset D is completed m The construction of N m The group is composed of "RGB image-depth image" pairs, which are in the form of k=1,2,3,...,N m , k represents the kth scene in the middle domain, N m It is the smaller value between the source domain and the target domain data volume.

[0087] Step 3: The source domain dataset D s , target domain dataset D t and the intermediate domain dataset D m Input cross-domain target posture estimation model; the cross-domain target posture estimation model includes three parts: feature extractor f, target posture estimation module and adaptive module; the feature extractor f adopts a dual-branch network to extract the source domain dataset D s , target domain dataset D t , intermediate domain dataset D m The RGB image I rgb and depth image I depth The characteristic is expressed as F r and F d , using the feature fusion module to fuse F r and F d The characteristic after is F f , whose feature dimension is BCN p , N p is the number of sampled feature points. The target posture estimation module uses the posture estimation generator g to obtain the corresponding relationship of the points, and then uses the least squares fitting algorithm to solve the posture parameters (R, t), where R is the rotation parameter and t is the translation parameter. The adaptive module introduces a domain discriminator to achieve cross-domain;

[0088] The least squares fitting algorithm is used to solve the posture parameters (R, t), specifically:

[0089] For the source domain dataset D s , the feature after feature extractor f is F f s, use the pose estimation generator g to obtain the semantic information, center point information and key point information of the target object, and then obtain two point sets of a target object through cluster voting. On this basis, the least squares fitting algorithm is used to solve the pose parameters (R, t). The least squares fitting algorithm is calculated by minimizing the following loss:

[0090]

[0091] in, is the point set of K key points detected in the camera coordinate system, is the set of points they correspond to in the object coordinate system.

[0092] The adaptation module includes two processes: adaptation from the source domain to the intermediate domain and adaptation from the intermediate domain to the target domain. The specific process is as follows:

[0093] For the adaptive process from source domain to intermediate domain:

[0094] The source domain dataset D s , intermediate domain dataset D m The data is preprocessed and fed into the feature extractor f;

[0095] After passing through the feature extractor f, the source domain aggregate feature F is obtained f s and the intermediate domain aggregation feature F f m , F f s They are fed to the pose generator g and the domain discriminator d respectively to obtain the pose estimation result and domain label of the source domain. f m Feed it to the domain discriminator d to obtain the domain label of the intermediate domain;

[0096] For the adaptive process from the intermediate domain to the target domain:

[0097] The intermediate domain dataset D m , target domain dataset D t The data is preprocessed and fed into the feature extractor f;

[0098] After passing through the feature extractor f, the intermediate domain aggregate feature F is obtained f m and the target domain aggregate feature F f t , F f m With F f t Feed it to the domain discriminator d to obtain the domain labels of the intermediate domain and the target domain;

[0099] The method also includes training a cross-domain target pose estimation model, wherein the total loss during the training process is divided into two parts: a task loss of the pose estimation part and a domain loss of the adaptation part;

[0100] The task loss of the attitude estimation part includes the following three parts:

[0101]

[0102] L seg =(-t(1-(c i ·l i )) γ log(c i ·l i ))

[0103]

[0104] Among them, i is a random point, p i Represents the coordinate information and feature information of point i, N is the number of vertices on the surface of the object, M is the number of key points of the selected target, of i j , They represent the predicted value and true value of the translation offset from the i-th random point to the j-th key point, Δx i , Δx i * They represent the predicted value and true value of the offset from the i-th random point to the center of the object, I is the target object instance, II is the indicator function, and when p i ∈I is equal to 1 when it holds, otherwise it is equal to 0; t is the balance parameter, γ is the focusing parameter, c i , l i Respectively represent the confidence that the i-th point belongs to the corresponding object category and the one-hot representation of the real object category label;

[0105] The task loss of the pose estimation part is:

[0106] L task =v1L keypoints +v2L seg +v3L center

[0107] Among them, v1, v2, and v3 are the weight factors of the three sub-task losses of key point information, semantic information, and center point information respectively;

[0108] The domain loss of the adaptive part includes the domain loss of the adaptive process from the source domain to the intermediate domain and the domain loss of the adaptive process from the intermediate domain to the target domain:

[0109] For the domain loss of the source domain to intermediate domain adaptation process: Assume that the feature vector set of the source domain is The set of eigenvectors of the intermediate domain is n and z are the data volumes of the source domain and the intermediate domain, respectively. as well as Represents the mean and variance of the two domains, which is used to describe the data distribution of the two domains and introduces the inter-domain alignment loss L align_1 The feature alignment is achieved as shown below:

[0110]

[0111] in, s , m are the covariance matrices of the source domain and the intermediate domain, respectively. F represents the Frobenius norm, and α and b are weight factors.

[0112] Design a domain discriminator d and use the domain adversarial loss L adv_1 Optimize as shown below:

[0113] L adv_1 =L bce (d(F f s ),0)+L bce (d(F f m ),1)

[0114] Among them, L bce represents the binary cross entropy loss function, 0 means the sample comes from the source domain, 1 means the sample comes from the intermediate domain;

[0115] The domain loss of the source domain to intermediate domain adaptation process is:

[0116] L domain_1 =L align_1 +L adv_1 ;

[0117] For the domain loss of the intermediate domain to target domain adaptation process: Assume that the feature vector set of the intermediate domain is The feature vector set of the target domain is z and m are the data sizes of the intermediate domain and the target domain, respectively. as well as Represents the mean and variance of the two domains, which is used to describe the data distribution of the two domains and introduces the inter-domain alignment loss L align_2 The feature alignment is achieved as shown below:

[0118]

[0119] in, s , mare the covariance matrices of the intermediate domain and the target domain, respectively. F represents the Frobenius norm, and α and b are weight factors.

[0120] Design a domain discriminator d and use the domain adversarial loss L adv_2 Optimize as shown below:

[0121] L adv_2 =L bce (d(F f m ),0)+L bce (d(F f t ),1)

[0122] Among them, L bce represents the binary cross entropy loss function, 0 means the sample is from the intermediate domain, and 1 means the sample is from the target domain;

[0123] The domain loss of the adaptive process from the intermediate domain to the target domain is:

[0124] L domain_2 =L align_2 +L adv_2 ;

[0125] The total loss of the cross-domain target pose estimation model is as follows:

[0126] L total =ω1L task +ω2(L domain_1 +L domain_2 )

[0127] Among them, ω1 and ω2 are the balance factors of the loss of the attitude estimation part and the adaptive part respectively;

[0128] The corresponding loss is calculated, and the parameters of the cross-domain target posture estimation model are optimized according to the loss value until convergence, that is, the training of the cross-domain target posture estimation model is completed.

[0129] The above shows and describes the main features and advantages of the present invention. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and scope of the equivalent elements of the claims are included in the present invention.

[0130] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A cross-domain target pose estimation adaptive method from a virtual domain to a real domain, characterized in that: The following steps are involved: Step 1: Build the source domain dataset D s , consisting of n sets of "RGB image-depth image-annotation information" pairs; construct the target domain dataset D t , consisting of m sets of "RGB image-depth image" pairs, s represents the virtual environment, and t represents the real environment; Step 2: Based on the source domain dataset D s and the target domain dataset D t Constructing the intermediate domain dataset D m , by N m The group is composed of "RGB image-depth image" pairs; Step 3: The source domain dataset D s , target domain dataset D t and the intermediate domain dataset D m Input cross-domain target posture estimation model; the cross-domain target posture estimation model includes three parts: feature extractor f, target posture estimation module and adaptive module; the feature extractor f adopts a dual-branch network to extract the source domain dataset D s , target domain dataset D t , intermediate domain dataset D m The RGB image I rgb and depth image I depth The characteristic is expressed as F r and F d , using the feature fusion module to fuse F r and F d The characteristic after is F f , whose feature dimension is BCN p , N p is the number of sampled feature points. The target posture estimation module uses the posture estimation generator g to obtain the correspondence of the points, and then uses the least squares fitting algorithm to solve the posture parameters (R, t), where R is the rotation parameter and t is the translation parameter. The adaptive module introduces a domain discriminator to achieve cross-domain.

2. According to claim 1, a cross-domain target posture estimation adaptive method from a virtual domain to a real domain is characterized in that: The step 1 is specifically as follows: Step 1.1, use a 3D scanner to scan the 3D model corresponding to the object in the real world, the format of which is .ply; Step 1.2: In the simulation software, import the .ply 3D model, build a scene in the virtual environment, fix the light intensity, write a script to obtain the image and the corresponding annotation information, and obtain the "RGB image-depth image-annotation information" pair, that is, the source domain dataset D s , D s ={(I rgbs i , I depths i , b i )}, i = 1, 2, 3, ..., n, n represents the number of data sets in the virtual environment, Represents the RGB image and corresponding depth image of virtual scene i, b i Represents the pose estimation annotation information in virtual scene i; Step 1.3, fix the target object, build a scene in a real environment, fix the lighting, use a camera to shoot the object for a week, use the 3D reconstruction method to reconstruct the 3D model of the target object, and obtain the "RGB image-depth image" pair, that is, the target domain dataset D t , m represents the number of data sets in the real environment, Represent the RGB image and corresponding depth image of the real scene j respectively; Step 1.4: Change the target object to be detected and the shooting scene, and repeat the above three steps until a complete data set D is obtained. s and D t .

3. The method for adaptively estimating cross-domain target posture from a virtual domain to a real domain according to claim 1, characterized in that: The step 2 is specifically as follows: The source domain dataset D s and the target domain dataset D t RGB image After transformation MLP_E is converted to the feature space, the source domain and target domain color features are After fusing the RGB information, the image is restored through inverse transformation MLP_D, which is the RGB image in the intermediate domain. As shown below: Among them, l∈[0,1] refers to the influence degree of the source domain RGB image; The source domain dataset D s and the target domain dataset D t Depth image Use weighted fusion to obtain the depth image of the intermediate domain Among them, m∈[0,1] refers to the influence degree of the source domain depth image; At this point, the intermediate domain dataset D is completed m The construction of N m The group consists of "RGB image-depth image" pairs, which are in the form of k represents the kth scene in the middle domain, N m It is the smaller value between the source domain and the target domain data volume.

4. The method for adaptively estimating cross-domain target posture from a virtual domain to a real domain according to claim 1, characterized in that: In step 3, the least squares fitting algorithm is used to solve the posture parameters (R, t), specifically: For the source domain dataset D s , the feature after the feature extractor f is F f s , use the pose estimation generator g to obtain the semantic information, center point information and key point information of the target object, and then obtain two point sets of a target object through cluster voting. On this basis, the least squares fitting algorithm is used to solve the pose parameters (R, t). The least squares fitting algorithm is calculated by minimizing the following loss: in, is the point set of K key points detected in the camera coordinate system, is the set of points they correspond to in the object coordinate system.

5. The method for adaptively estimating cross-domain target posture from a virtual domain to a real domain according to claim 1, characterized in that: The adaptive module in step 3 includes two processes: adaptation from the source domain to the intermediate domain and adaptation from the intermediate domain to the target domain. The specific process is as follows: For the adaptive process from source domain to intermediate domain: The source domain dataset D s , intermediate domain dataset D m The data is preprocessed and fed into the feature extractor f; After passing through the feature extractor f, the source domain aggregate feature F is obtained f s and the intermediate domain aggregation feature F f m , F f s They are fed to the pose generator g and the domain discriminator d respectively to obtain the pose estimation result and domain label of the source domain. f m Feed it to the domain discriminator d to obtain the domain label of the intermediate domain; For the adaptive process from the intermediate domain to the target domain: The intermediate domain dataset D m , target domain dataset D t The data is preprocessed and fed into the feature extractor f; After passing through the feature extractor f, the intermediate domain aggregate feature F is obtained f m and the target domain aggregate feature F f t , F f m With F f t Feed it to the domain discriminator d to obtain the domain labels of the intermediate domain and the target domain.

6. The method for adaptively estimating cross-domain target posture from a virtual domain to a real domain according to claim 5, characterized in that: The method also includes training a cross-domain target pose estimation model, wherein the total loss during the training process is divided into two parts: a task loss of the pose estimation part and a domain loss of the adaptation part; The task loss of the attitude estimation part includes the following three parts: L seg =(-t(1-(c i ·l i )) γ log(c i ·l i )) Among them, i is a random point, p i Represents the coordinate information and feature information of point i, N is the number of vertices on the surface of the object, M is the number of key points of the selected target, They represent the predicted value and true value of the translation offset from the i-th random point to the j-th key point, Δx i , Δx i * They represent the predicted value and true value of the offset from the i-th random point to the center of the object, I is the target object instance, II is the indicator function, and when p i ∈I is equal to 1 when it holds, otherwise it is equal to 0; t is the balance parameter, γ is the focusing parameter, c i , l i Respectively represent the confidence that the i-th point belongs to the corresponding object category and the one-hot representation of the real object category label; The task loss of the pose estimation part is: L task =v1L keypoints +v2L seg +v3L center Among them, v1, v2, and v3 are the weight factors of the three sub-task losses of key point information, semantic information, and center point information respectively; The domain loss of the adaptive part includes the domain loss of the adaptive process from the source domain to the intermediate domain and the domain loss of the adaptive process from the intermediate domain to the target domain: For the domain loss of the source domain to intermediate domain adaptation process: Assume that the feature vector set of the source domain is The set of eigenvectors of the intermediate domain is n and z are the data volumes of the source domain and the intermediate domain, respectively. as well as Represents the mean and variance of the two domains, which is used to describe the data distribution of the two domains and introduces the inter-domain alignment loss L align_1 The feature alignment is achieved as shown below: L align_1 =a(||m s -m m ||2 2 )+b(|s s 2 -s m 2 |)+(1-a-b)(||∑ s -∑ m || F 2 ) Among them, s and m are the covariance matrices of the source domain and the intermediate domain respectively, F represents the Frobenius norm, and α and b are weight factors; Design a domain discriminator d and use the domain adversarial loss L adv_1 Optimize as shown below: L adv_1 =L bce (d(F f s ),0)+L bce (d(F f m ),1) Among them, L bce represents the binary cross entropy loss function, 0 means the sample comes from the source domain, 1 means the sample comes from the intermediate domain; The domain loss of the source domain to intermediate domain adaptation process is: L domain_1 =L align_1 +L adv_1 ; For the domain loss of the intermediate domain to target domain adaptation process: Assume that the feature vector set of the intermediate domain is X m ={x m1 ,x m2 ,...,x mz }, the feature vector set of the target domain is X t ={x t1 ,x t2 ,...,x tm }, z and m are the data volumes of the intermediate domain and the target domain respectively, using as well as Represents the mean and variance of the two domains, which is used to describe the data distribution of the two domains and introduces the inter-domain alignment loss L align_2 The feature alignment is achieved as shown below: L align_2 =a(||m m -m t ||2 2 )+b(|s m 2 -s t 2 |)+(1-a-b)(||∑ m -∑ t || F 2 ) in, s , m are the covariance matrices of the intermediate domain and the target domain, respectively. F represents the Frobenius norm, and α and b are weight factors. Design a domain discriminator d and use the domain adversarial loss L adv_2 Optimize as shown below: L adv_2 =L bce (d(F f m ),0)+L bce (d(F f t ),1) Among them, L bce represents the binary cross entropy loss function, 0 means the sample is from the intermediate domain, and 1 means the sample is from the target domain; The domain loss of the adaptive process from the intermediate domain to the target domain is: L domain_2 =L align_2 +L adv_2 ; The total loss of the cross-domain target pose estimation model is as follows: L total =ω1L task +ω2(L domain_1 +L domain_2 ) Among them, ω1 and ω2 are the balance factors of the loss of the attitude estimation part and the adaptive part respectively; The corresponding loss is calculated, and the parameters of the cross-domain target posture estimation model are optimized according to the loss value until convergence, that is, the training of the cross-domain target posture estimation model is completed.

Citation Information

Patent Citations

  • Robust cross-domain attitude estimation method and system based on image mixing mechanism

    CN113642684A

  • Video segmentation field adaptive method and system based on cross-domain environment motion alignment

    CN118015623A

  • Bidirectional fusion 6D object pose estimation method

    CN118799393A