Consistency constraint cross-domain multi-view target recognition method and device
By constructing a feature embedding space of the gravitational field and a consistency regularization term, the problems of noisy pseudo-labels and insufficient semantic consistency in cross-domain multi-view target detection are solved, achieving higher retrieval accuracy and stability.
Patent Information
- Application Number
- CN202310149726.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2043-02-17
AI Technical Summary
Existing multi-view target detection algorithms are insufficient in eliminating the adverse effects of noise and pseudo-labels and mitigating the weakening of semantic consistency in data from different modalities, resulting in insufficient accuracy in cross-domain multi-view target retrieval.
We design a feature embedding space based on a gravitational field, construct a feature gravitational field through pivot feature selection and feature movement mechanisms to guide cross-domain samples to appropriate positions, introduce a consistency regularization term, use KL divergence to construct the consistency of instance and prototype similarity, and combine two consistency similarity measurement strategies to achieve semantic alignment across instance and prototype spaces.
It improves the accuracy of cross-domain multi-view target retrieval, enhances the stability of feature distribution and semantic alignment capability, reduces the negative impact of pseudo-labels, and improves retrieval performance.
Smart Images

Figure CN116129205B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of multi-view target retrieval, domain adaptation, and semantic alignment, and particularly to a cross-domain multi-view target recognition method and apparatus with consistency constraints. Background Technology
[0002] With the development and progress of scientific research, target retrieval technology has been widely used in daily life. [1] With the rapid growth of data, the efficient and accurate retrieval of required data has attracted widespread attention. Most existing retrieval methods require labeled data to explore the multi-view features of target objects, but labeled data is not readily available in practical applications. Inspired by the advantages of labeled images and the development of domain adaptation algorithms, it is feasible to transfer learned knowledge from images to multi-view targets. One of the most popular tasks is to search for relevant multi-view targets using images in unlabeled databases, i.e., cross-domain multi-view target retrieval.
[0003] This task is challenging due to the domain differences between different modalities. Therefore, researchers have proposed many methods to address the domain adaptation problem, such as: difference-based approaches. [2-4] Based on adversarial [5-6] Based on prototype [7-9] However, the performance of these methods is inevitably affected by noisy pseudo-labels. For example, in prototype-based methods, the prototype of the source domain can be obtained using its own label information, while the target prototype needs to be computed using pseudo-labels assigned by the source classifier. Since there are obvious spurious pseudo-labels in the inter-domain gaps, the prototypes obtained from these labels can mislead the domain adaptation process. Therefore, it is crucial to suppress the negative information generated by these noisy pseudo-labels.
[0004] On the other hand, difference-based and adversarial-based methods cannot effectively reduce domain differences because they ignore the semantic information involved in the instances. Prototype-based methods overcome this shortcoming, such as MSTN. [7] SC-IFA [9] Methods such as these. Although these methods have achieved good results in semantic alignment, they only measure and align semantics in one space and cannot fully maintain semantic consistency, thus weakening domain alignment.
[0005] In summary, existing multi-view target detection algorithms have the following two drawbacks and shortcomings:
[0006] 1. How to eliminate the adverse effects of noise-induced false labels on domain adaptation;
[0007] 2. How to mitigate the weakening of semantic consistency caused by aligning different modalities in a single space. Summary of the Invention
[0008] This invention provides a consistency-constrained cross-domain multi-view target recognition method and apparatus. It designs a feature embedding space based on a gravitational field, constructing a feature gravitational field through the selection of pivot features and a feature movement mechanism. This guides cross-domain samples to move to appropriate positions in the feature distribution space, improving the stability of the feature distribution. In the prototype update part, a similarity measure between instances and prototypes is introduced, along with a consistency regularization constraint. By using KL (Kullback-Leibler) divergence to construct the consistency of instance-prototype similarity, the negative impact of noise pseudo-labels is better suppressed, and semantic alignment is effectively guided, improving the accuracy of cross-domain multi-view target retrieval. Through two consistent similarity measurement strategies, instances are mapped to a category space, achieving semantic alignment across instance and prototype spaces. This alleviates the insufficiency of measuring in a single space for semantic information mining. Details are described below.
[0009] A cross-domain, multi-view target recognition method with consistency constraints, the method comprising the following steps:
[0010] We design a feature embedding space based on a gravitational field. By selecting pivot features and using a feature movement mechanism, we construct a feature gravitational field to guide cross-domain samples to move to a suitable position in the feature distribution space.
[0011] Prototype-based learning updates the source domain prototype using labeled source domain features and updates the target domain prototype using pseudo-labeled target domain features, thereby achieving the learning of semantic representations.
[0012] Instead of predicting the probability of a category, the similarity between instances and prototypes is used to associate instances with all prototypes, neutralizing the impact of erroneous pseudo-labels on the calculation of the target prototype; a consistency regularization term is used to constrain the similarity measure, and KL divergence is used to build consistency between instance and prototype similarity.
[0013] Using a single instance as a benchmark, the similarity between prototypes from two different domains is measured, encouraging consistent similarity between single instances and cross-domain prototypes belonging to the same category. Instances are used to guide the alignment of prototypes and reduce the distance between cross-domain prototypes.
[0014] Instance pairs are constructed, and based on these instance pairs, consistent similarity between the instance pairs and individual prototypes is explored. This reduces the intra-class distance between instances at both the feature and semantic levels, thereby enhancing the performance of cross-domain multi-view target retrieval.
[0015] A cross-domain multi-view target recognition device with consistency constraints, the device comprising: a processor and a memory, the memory storing program instructions, the processor calling the program instructions stored in the memory to cause the device to execute any of the steps of the method described above.
[0016] The beneficial effects of the technical solution provided by this invention are:
[0017] 1. This invention designs a feature embedding space based on a gravitational field. By selecting fulcrum features and a feature movement mechanism, a feature gravitational field is constructed to guide cross-domain samples to move to a suitable position in the feature distribution space, thus making up for the instability of traditional feature distribution and the fuzziness of class boundaries.
[0018] 2. This invention uses source domain samples and labels to calculate the source prototype, and the target prototype is calculated using target samples and pseudo-labels. To reduce the negative impact of using pseudo-labels alone to predict the target prototype on retrieval performance, this invention uses the similarity between instances and prototypes instead of probabilistic prediction of categories. By mining the semantic association between cross-domain instances and class prototypes, the similarity between instances and prototypes is used to obtain association information.
[0019] 3. Inspired by consistency regularization in semi-supervised learning, this invention introduces a consistency regularization term to constrain the similarity measurement between instances and prototypes, and constructs the consistency of instance-prototype similarity by using symmetric KL (Kullback-Leibler) divergence. Through two consistent similarity measurement strategies, instances are mapped to a category space, achieving semantic alignment across instance and prototype spaces, thus alleviating the insufficiency of measuring semantic information in a single space.
[0020] 4. In the first strategy, the similarity between prototypes from two different domains is measured based on a single instance that distinguishes the source domain and the target domain. This encourages a consistent similarity between a single instance and a cross-domain prototype belonging to the same category. The aim is to use instances to guide the alignment of prototypes and reduce the distance between cross-domain prototypes. Starting from the instance level, this makes the mapping from instance to prototype more accurate.
[0021] 5. In the second strategy, instance pairs are constructed by calculating the distance between source features and target features. Based on the instance pairs, the consistency similarity between the instance pairs and individual prototypes is explored. By measuring the pairs of instances with the same prototype, the intra-class distance between instances is reduced at the feature level and semantic level, thereby enhancing the performance of cross-domain multi-view target retrieval. Attached Figure Description
[0022] Figure 1 A flowchart of a cross-domain multi-view target recognition method with consistency constraints;
[0023] Figure 2 A network structure diagram for cross-domain multi-view target recognition with consistency constraints. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below.
[0025] Example 1
[0026] A cross-domain, multi-view target recognition method with consistency constraints, see [link to relevant documentation]. Figure 1 The method includes the following steps:
[0027] 101: Utilizing convolutional neural networks to extract image features (source domain) and multi-view features of the target object (target domain);
[0028] 102: Based on the adversarial learning strategy, the global alignment of the source and target domains is guided by training classifiers and discriminators;
[0029] 103: Design a feature embedding space based on a gravitational field. By selecting pivot features and a feature movement mechanism, a feature gravitational field is constructed to guide cross-domain samples to move to a suitable position in the feature distribution space, thereby improving the stability of the feature distribution.
[0030] 104: Based on the prototype learning method, the source domain prototype is updated using labeled source domain features, and the target domain prototype is updated using pseudo-labeled target domain features, thereby achieving the learning of semantic representation;
[0031] Compared to adversarial training, prototype-based domain adaptation methods, while considering the semantic information contained in instances, inevitably suffer from the influence of noise-based labels on the target prototypes obtained using pseudo-labels due to the significant differences between the source and target domains and the lack of label information in the target domain. Therefore, this method mines the similarity information between instances and prototypes, encouraging consistent similarity between them and reducing the impact of noise-based pseudo-labels on domain adaptation.
[0032] 105: Unlike previous methods that only used pseudo-labels to predict target prototypes, this method uses the similarity between instances and prototypes instead of probabilistic predictions of categories. It leverages similarity to associate instances with all prototypes, neutralizing the impact of erroneous pseudo-labels on target prototype computation. Then, a consistency regularization term is used to constrain the similarity measure, and consistency between instance and prototype similarity is constructed using KL (Kullback-Leibler) divergence. Through these two consistent similarity measurement strategies, the accuracy of instance mapping to the category space is improved, achieving semantic alignment across instance and prototype spaces, thus mitigating the inadequacy of measuring in a single space for semantic information mining.
[0033] 106: In the first strategy, similarity between prototypes from two different domains is measured based on a single instance. Consistent similarity between individual instances and cross-domain prototypes belonging to the same category is encouraged. Instances are used to guide prototype alignment and reduce the distance between cross-domain prototypes, making the instance-to-prototype mapping more accurate.
[0034] 107: In the second strategy, instance pairs are constructed, and based on these pairs, consistent similarity between the instance pairs and individual prototypes is explored. This reduces the intra-class distance between instances at both the feature and semantic levels, enhancing the performance of cross-domain multi-view target retrieval.
[0035] In summary, the embodiments of the present invention propose a novel method for cross-domain multi-view target retrieval and design a novel network structure, thereby improving the performance of multi-view target retrieval.
[0036] Example 2
[0037] The scheme in Example 1 will be further described below with specific examples and calculation formulas:
[0038] 201: Utilizing neural networks to extract image features (source domain) and multi-view features of the target object (target domain);
[0039] Step 201 primarily includes: using AlexNet as the CNN architecture to extract features from the source and target domains. In this network structure, the first five layers (conv1-conv5) are convolutional layers, and the remaining three layers are fully connected layers. To better learn visual feature representations, in the fully connected layers... Then add a bottleneck layer.
[0040] 202: Achieve global alignment of the source and target domains through adversarial training;
[0041] The encoder F, shared by the source and target domains, extracts feature representations. For the source domain data, a classifier C is trained using cross-entropy (CE) loss.
[0042]
[0043] Among them, L CE For cross-entropy loss, F(x) s ) is the feature extractor, D S For source domain data, x s For the source domain sample, y s For source domain tags, This is the set of all source domain samples.
[0044] Using L D A domain discriminator D is trained to distinguish features between the source and target domains. Simultaneously, F is trained to generate features for the deception discriminator.
[0045]
[0046] Where F(x) t ) is the feature extractor, x t For the target domain sample, D T For target domain data, For the source domain sample set, The target domain sample set.
[0047] Global alignment can be achieved by minimizing the differences between the source and target domains and the classifier error.
[0048] 203: This invention designs a field-based feature embedding space. By selecting fulcrum features and using a feature classification mechanism, a feature gravitational field is constructed to guide cross-domain samples to a suitable position in the feature distribution space, thereby overcoming the instability and fuzzy class boundaries of traditional feature distributions and improving feature representation capabilities.
[0049] Based on category K, the source domain data S and target domain data T are divided to construct a pivot feature set.
[0050]
[0051]
[0052] Where, x s For the source domain sample, y t The true label corresponding to the source domain sample; x t For the target domain sample, The pseudo-labels assigned by the classifier to the target domain samples.
[0053] Fulcrum Features Filtering: Pivot Feature Set Including cross-domain instances with the same class label and different categories from the same domain, a new feature is selected as the pivot feature if it minimizes the average distance to other objects.
[0054]
[0055] Where d(·,·) is the Euclidean distance between the features of the two candidate pivots. This represents the total number of samples in the source and target domains. During training, the selection of pivot features is updated iteratively.
[0056] 204: Construct a feature-based gravitational field by using iteratively updated features as the fulcrum of the field and a feature classification mechanism;
[0057] This method defines two axes, which are cross-domain but of the same category and are positive. Same domain but different categories are negative The characteristic gravitational field forms a relatively effective reference space at the sample level, and at the same time, it can learn the global characteristic distribution across domains.
[0058] During the construction of the characteristic gravitational field, the fulcrum set is sampled:
[0059]
[0060] Using a manifold mixing algorithm, a large number of pivot points are generated through linear interpolation. The mixing coefficients are obtained using a Dirichlet distribution.
[0061] w~Dir(α),|α|>2 (7)
[0062] The mixing process is repeated to generate fulcrums, thereby constructing a gravitational field between the fulcrum features:
[0063]
[0064] Among them, w k Let be the mixing coefficient. Then, under the influence of the two axes, the sample will stably move in the gravitational field towards the direction with the greatest influence of the domain-invariant embedding.
[0065] 205: After constructing the gravitational field, define how features move within it:
[0066] When learning domain-invariant features, features of the same category (whether from the source or target domain) should cluster together. Therefore, samples in a gravitational field should move towards the same category, i.e., the positive direction as defined above, and away from the negative direction. Thus, the samples... There should be good interaction between the sample and the field, so that the sample can move to a reasonable position in the feature space.
[0067] Using L pos Constraining the positive movement of the sample:
[0068]
[0069] Conversely, for samples moving away from the negative direction, use L... neg Apply constraints:
[0070]
[0071] in, As a positive pivot point in a gravitational field, As the negative pivot point in the gravitational field, These are all the pivot points in the field.
[0072] 206: Based on the prototype learning method, the source domain prototype is updated using labeled source domain features, and the target domain prototype is updated using pseudo-labeled target domain features, thereby achieving the learning of semantic representation;
[0073]
[0074] Where K is the total number of categories, and Let L be the k-th prototype of the source and target domains, φ(·) be the distance metric function, and L be the k-th prototype of the source and target domains. ST (x s ,y s ,x t () represents semantic transfer loss.
[0075] 207: This approach uses the similarity between instances and prototypes instead of probabilistic category predictions to mine semantic associations between cross-domain instances and class prototypes. It employs consistency regularization to constrain similarity metrics and builds consistency between instance and prototype similarities using KL (Kullback-Leibler) divergence. This reduces the negative impact of using pseudo-labels alone to predict categories on retrieval performance and better guides cross-domain semantic alignment.
[0076] This invention proposes two consistent similarity measurement strategies. In the first strategy, an instance is used to measure the similarity between prototypes from two different domains. Consistent similarity between a single instance and cross-domain prototypes belonging to the same category is encouraged. This strategy aims to use instances to guide prototype alignment and reduce the distance between cross-domain prototypes. In the second strategy, it is suggested to use an instance pair to measure consistency with a single prototype. Instances within this pair are measured using the same prototype, thereby reducing the intra-class distance between different instances at both the feature and semantic levels. The combination of these two strategies improves the accuracy of instance mapping to the category space, achieving semantic alignment across instance and prototype spaces.
[0077] 208: Measure the similarity between prototypes from two different domains based on a single instance. Use KL (Kullback-Leibler) divergence to constrain the similarity between a single instance and multiple prototypes, encouraging consistent similarity between a single instance and cross-domain prototypes belonging to the same category.
[0078] The specific steps in this strategy include:
[0079] Randomly select sample x i It does not distinguish whether it comes from the source domain or the target domain.
[0080] During training, the selected x i To retrieve samples, embodiments of this invention use the source prototype. The target prototype and the query sample x are respectively measured. i Similarity:
[0081]
[0082]
[0083] in, and This is the loss function. L sim This can be viewed as a type of cross-entropy loss, which trains the network to distinguish whether the prototype and the selected sample belong to the same category. S is the source domain, T is the target domain, and K is the total number of categories. For sample x i The distance metric function from the source prototype, where τ is a temperature parameter. For sample x i Distance metric function relative to the target prototype.
[0084] A consistency regularization term is introduced to encourage samples to have consistent similarity with cross-domain prototypes (of the same class), guiding the alignment of cross-domain prototypes and improving semantic consistency. Sample x... i The similarity between the source prototype and the target prototype is described as follows:
[0085]
[0086]
[0087] Where M(j) represents sample x i In the source prototype set The probability of a match, N(j) represents the probability of a match for sample x. i In the target prototype set The probability of a match.
[0088] The convergence of two similarities is constrained by using symmetric KL (Kullback-Leibler) divergence:
[0089]
[0090] 209: Construct instance pairs and, based on these pairs, perform a consistent similarity measurement between instances belonging to the same pair and a single prototype, using KL (Kullback-Leibler) divergence as a constraint. This prototype-based approach shortens the intra-class distance between different instances, enhancing semantic consistency.
[0091] The specific steps in this strategy include:
[0092] Selected instance pairs Each instance pair contains source and target domain instances. The prototype Z does not distinguish between source and target domains. k Perform similarity measurements with instances in the instance pair and apply consistency constraints. Reduce intra-class distance of instances by focusing on the prototype.
[0093] Among them, instance pairs We apply nearest neighbor information to mine instance pair relationships. We calculate the Euclidean distance between the source and target features, select the k nearest neighbors, and compare the labels.
[0094]
[0095] in, For source tags, For target pseudo-labels.
[0096] Two samples with a label comparison score S=1 are considered positive pairs, and the distance between positive pairs should be close to each other. Conversely, samples with S=0 should be pushed apart.
[0097] Then, the similarity to the same prototype is measured using the source and target samples in the instance pair, respectively. The losses in equations (4) and (5) can then be written as:
[0098]
[0099]
[0100] Where N is the number of instance pairs. To distinguish between the two strategies mentioned above, Z is used. k This represents a class prototype that does not distinguish between source and target domains, and considers Z to be... k To query the prototype.
[0101] Source samples in instance pairs and target sample With query prototype Z k The similarity is represented as follows:
[0102]
[0103]
[0104] Use KL (Kullback-Leibler) divergence to constrain the convergence of two similarities:
[0105]
[0106] Example 3
[0107] The feasibility of the schemes in Examples 1 and 2 is verified by specific experiments, as detailed below:
[0108] This invention compares the results with some classic domain adaptation methods on the general multi-view object retrieval dataset MI3DOR-2. MI3DOR-2 contains 40 classes, a total of 19,694 two-dimensional images and 3,982 multi-view object models. Among them, 19,294 images and 3,182 models are divided into the training set, while another 400 images and 800 models are divided into the test set.
[0109] The specific evaluation metrics are as follows: Nearest Neighbor (NN), First Tier (FT), Second Tier (ST), F-measure (F), Discounted Cumulative Gain (DCG), and Average Normalized Corrected Retrieval Rank (ANMRR).
[0110] Except for AMNRR, higher values for the other metrics indicate better retrieval performance.
[0111]
[0112] The experimental data above shows that the consistency-constrained cross-domain multi-view target recognition method proposed in this invention improves all evaluation metrics compared to previous methods. Specifically, the improvements in the first five metrics are 5.8%-26.5%, 3.8%-32.4%, 6.9%-31.9%, 5.6%-32.4%, and 5.9%-32.4%, respectively, while the AMNRR decreases by 9.1%-32.5%, indicating that this invention has superior cross-domain multi-view target retrieval performance.
[0113] Example 4
[0114] A cross-domain multi-view target recognition device with consistency constraints, the device comprising: a processor and a memory. The processor and memory store program instructions, and the processor invokes the program instructions stored in the memory to cause the device to perform the following method steps:
[0115] We design a feature embedding space based on a gravitational field. By selecting pivot features and using a feature movement mechanism, we construct a feature gravitational field to guide cross-domain samples to move to a suitable position in the feature distribution space.
[0116] Prototype-based learning updates the source domain prototype using labeled source domain features and updates the target domain prototype using pseudo-labeled target domain features, thereby achieving the learning of semantic representations.
[0117] Instead of predicting the probability of a category, the similarity between instances and prototypes is used to associate instances with all prototypes, neutralizing the impact of erroneous pseudo-labels on the calculation of the target prototype; a consistency regularization term is used to constrain the similarity measure, and KL divergence is used to build consistency between instance and prototype similarity.
[0118] Using a single instance as a benchmark, the similarity between prototypes from two different domains is measured, encouraging consistent similarity between single instances and cross-domain prototypes belonging to the same category. Instances are used to guide the alignment of prototypes and reduce the distance between cross-domain prototypes.
[0119] Instance pairs are constructed, and based on these instance pairs, consistent similarity between the instance pairs and individual prototypes is explored. This reduces the intra-class distance between instances at both the feature and semantic levels, thereby enhancing the performance of cross-domain multi-view target retrieval.
[0120] The selection of the pivot feature is as follows:
[0121] Pivot feature set Including cross-domain instances with the same class label and different categories from the same domain, a new feature is selected as the pivot feature if it minimizes the average distance to other objects.
[0122]
[0123] Where d(·,·) is the Euclidean distance between the features of the two candidate pivots. This represents the total number of samples in the source and target domains.
[0124] The characteristic gravitational field is constructed as follows:
[0125]
[0126] Among them, w k This is the mixing coefficient.
[0127] Among them, guiding cross-domain samples to a suitable position in the feature distribution space is as follows:
[0128] Using L pos Constraining the positive movement of the sample:
[0129]
[0130] For samples moving away from the negative direction, use L neg Apply constraints:
[0131]
[0132] in, As a positive pivot point in a gravitational field, As the negative pivot point in the gravitational field, These are all the pivot points in the field.
[0133] It should be noted that the device descriptions in the above embodiments correspond to the method descriptions in the embodiments, and the embodiments of the present invention will not be repeated here.
[0134] The execution entities of the aforementioned processor and memory can be devices with computing functions such as computers, microcontrollers, and single-chip microcomputers. In specific implementations, the embodiments of the present invention do not limit the execution entities and can select them according to the needs of actual applications.
[0135] Data signals are transmitted between the memory and the processor via a bus, which will not be elaborated upon in this embodiment of the invention.
[0136] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium, the storage medium including a stored program, which, when the program is running, controls the device where the storage medium is located to execute the method steps in the above embodiments.
[0137] The computer-readable storage medium includes, but is not limited to, flash memory, hard disk, solid-state drive, etc.
[0138] It should be noted that the description of the readable storage medium in the above embodiments corresponds to the description of the method in the embodiments, and the embodiments of the present invention will not be repeated here.
[0139] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated.
[0140] A computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in or transmitted through a computer-readable storage medium. A computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic or semiconductor, etc.
[0141] References:
[0142] [1]Zhen Chen,Jie Liu,Meilu Zhu,Peter Y.M.Woo,and Yixuan Yuan.Instanceimportance-aware graph convolutional network for 3d medical diagnosis.MedicalImage Anal.,78:102421,2022.
[0143] [2]Hongliang Yan,Yukang Ding,Peihua Li,Qilong Wang,Yong Xu,andWangmeng Zuo.Mind the class weight bias:Weighted maximum mean discrepancy forunsupervised domain adaptation.In IEEE Conference on Computer Vision andPattern Recognition,pages 945–954,2017.
[0144] [3]Jian Shen,Yanru Qu,Weinan Zhang,and Yong Yu.Wasserstein distanceguided representation learning for domain adaptation.In Proceedings of theThirty-Second AAAI Conference on Artificial Intelligence,(AAAI-18),pages4058–4065,2018.
[0145] [4]Guoliang Kang,Lu Jiang,Yi Yang,and AlexanderG.Hauptmann.Contrastive adaptation network for unsupervised domainadaptation.In IEEE Conference on Computer Vision and Pattern Recognition,pages 4893–4902,2019.
[0146] [5]Yaroslav Ganin and Victor S.Lempitsky.Unsupervised domainadaptation by backpropagation.In Proceedings of the 32nd InternationalConference on Machine Learning,volume 37of JMLR Workshop and ConferenceProceedings,pages 1180–1189,2015.
[0147] [6]Mingsheng Long,Han Zhu,Jianmin Wang,and Michael I.Jordan.Deeptransfer learning with joint adaptation networks.In Proceedings of the 34thInternational Conference on Machine Learning,volume 70of Proceedings ofMachine Learning Research,pages 2208–2217,2017.
[0148] [7]Shaoan Xie,Zibin Zheng,Liang Chen,and Chuan Chen.Learning semanticrepresentations for unsupervised domain adaptation.In Proceedings of the 35thInternational Conference on Machine Learning,volume 80of Proceedings ofMachine Learning Research,pages 5419–5428,2018.
[0149] [8] Yuting Su, Yuqian Li, Dan Song, Weizhi Nie, Wenhui Li, and An-AnLiu. Consistent domain structure learning and domain alignment for 2d image-based 3d objects retrieval. In Proceedings of the TwentyNinth InternationalJoint Conference on Artificial Intelligence, IJCAI, pages 883–889, 2020.
[0150] [9] Heyu Zhou, Weizhi Nie, Dan Song, Nian Hu, Xuanya Li, and An-AnLiu. Semantic consistency guided instance feature alignment for 2d image-based3d shape retrieval. In MM'20: The 28th ACM International Conference on Multimedia, pages 925–933, 2020.
[0151] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0152] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A cross-domain multi-view target recognition method with consistency constraints, characterized in that, The method is executed by a processor and includes the following steps: Convolutional neural networks are used to extract source domain image features and multi-view features of target objects in the target domain. Based on the adversarial learning strategy, the global alignment of the source and target domains is guided by training classifiers and discriminators. We design a feature embedding space based on a gravitational field. By selecting pivot features and using a feature movement mechanism, we construct a feature gravitational field to guide cross-domain samples to move to a suitable position in the feature distribution space. Prototype-based learning updates the source domain prototype using labeled source domain features and updates the target domain prototype using pseudo-labeled target domain features, thereby achieving the learning of semantic representations. Instead of predicting the probability of a category, the similarity between instances and prototypes is used to associate instances with all prototypes, neutralizing the impact of erroneous pseudo-labels on the calculation of the target prototype; a consistency regularization term is used to constrain the similarity measure, and KL divergence is used to build consistency between instance and prototype similarity. Using a single instance as a benchmark, the similarity between prototypes from two different domains is measured, encouraging consistent similarity between single instances and cross-domain prototypes belonging to the same category. Instances are used to guide the alignment of prototypes and reduce the distance between cross-domain prototypes. Construct instance pairs and use them as a benchmark to explore consistent similarity between instance pairs and individual prototypes, thereby reducing intra-class distance between instances at the feature level and semantic level and enhancing cross-domain multi-view target retrieval performance. The constructed characteristic gravitational field is as follows: ; in, The mixing coefficient; The method for guiding cross-domain samples to a suitable position in the feature distribution space is as follows: use Constraining the positive movement of the sample: ; For samples moving away from the negative direction, use Apply constraints: where, As a positive pivot point in a gravitational field, As the negative pivot point in the gravitational field, All the pivots in the field; 。 2. The cross-domain multi-view target recognition method with consistency constraints according to claim 1, characterized in that, The selection of the fulcrum features is as follows: Pivot feature set This includes cross-domain instances with the same class label and different categories from the same domain. A new feature is selected as the pivot feature if it minimizes the average distance to other objects. Let Euclidean distance be the feature distance between two candidate pivots. This represents the total number of samples in the source and target domains. 。 3. A cross-domain multi-view target recognition device with consistency constraints, characterized in that, The device includes a processor and a memory, the memory storing program instructions, the processor calling the program instructions stored in the memory to cause the device to perform the steps of the method according to any one of claims 1-2.
Citation Information
Patent Citations
Multi-source cross-domain expression recognition method and device and storage medium
CN114612961A
Cross-domain multi-view target website retrieval method and device based on residual semantic consistency
CN115640418A