Domain generalization method, device and model product

Through the combination of style memory and semantic memory, the style characteristics related to the field and the semantic characteristics are learned and the semantic characteristics are decoupled, which solves the problem of insufficient field generalization ability in the existing technology and achieves better generalization ability in the target field.

CN113887238BActive Publication Date: 2025-05-27JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111151006.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-05-27
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

The existing domain generalization method is difficult to effectively generalize to target areas with large differences from the source domain data distribution, and feature learning based on adversarial training cannot accurately represent semantic information.

Method used

Through the style features related to the field of style memory learning, the semantic memory learns the field-independent semantic features, and decoupling of style features and semantic features, extracting the semantic features that are independent of the sample.

Benefits of technology

The model's ability to generalize to any unknown target domain is improved, so that the model can better generalize to target domains with large data distribution differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113887238B_ABST
    Figure CN113887238B_ABST
Patent Text Reader

Abstract

The present disclosure proposes a domain generalization method, device and model product, which relates to the field of computer vision technology. The present disclosure uses style memory to learn domain-related style features, uses semantic memory to learn domain-independent semantic features, and then decouples style features and semantic features. According to the semantic features, the category of samples in the source domain is predicted by classification, and the parameters of the domain generalization model are updated according to the learning loss of style features, the learning loss of semantic features, the decoupling loss, and the classification loss, so that the model is capable of extracting semantic features of samples that are independent of the domain, and improving the generalization ability of the model for any unknown target domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular to a domain generalization method, device and model product. Background Art

[0002] Domain generalization technology is one of the basic topics in the field of computer vision. The problem it studies is to learn a model with strong generalization ability from several data sets with different data distributions (source domains) in order to achieve better prediction results on unknown test sets (target domains). Since the target domain data is invisible during the training process, domain generalization is a very challenging and practical scenario.

[0003] The domain generalization method based on data enhancement generates some "pseudo" target domain data through a generative adversarial network, and then expands this part of the generated data into the training set for training, thereby improving the generalization ability of the model. Generally, these "pseudo" target domain data are obtained by perturbing the source domain data along the direction of domain distribution changes. This method can often only generate data between multiple source domain data distributions or adjacent data, and often cannot cover the target domain that is significantly different from the original domain distribution, so the generalization ability is limited.

[0004] Domain generalization methods based on feature learning attempt to learn domain-independent features through adversarial training. The learned domain-independent features cannot accurately represent semantic information and thus cannot be well generalized to the target domain. Summary of the invention

[0005] A technical problem to be solved by the embodiments of the present disclosure is to improve the domain generalization capability.

[0006] The disclosed embodiment utilizes style memory to learn domain-related style features, utilizes semantic memory to learn domain-independent semantic features, and then decouples style features from semantic features, thereby enabling the model to extract domain-independent semantic features of samples and improving the model's generalization ability for any unknown target domain.

[0007] Some embodiments of the present disclosure provide a domain generalization method, including:

[0008] Using samples from the source domain and based on style memory information, learn domain-related style features;

[0009] Using samples from the source domain and semantic memory information, we learn domain-independent semantic features.

[0010] Decouple style features from semantic features;

[0011] According to the semantic features, the category of the sample in the source domain is predicted by classification;

[0012] According to the learning loss of style features, the learning loss of semantic features, the decoupling loss, and the classification loss, the parameters of the domain generalization model are updated. After learning, the domain generalization model can extract domain-independent semantic features of samples in the target domain.

[0013] In some embodiments, the domain generalization model includes:

[0014] The encoder includes: a feature encoder for extracting sample features, a semantic encoder cascaded with the feature encoder for extracting semantic features, and a style encoder cascaded with the feature encoder for extracting style features;

[0015] The memory encoder is consistent with the structure of the encoder. During the learning process, the parameters of the memory encoder are updated according to the accumulation of the parameters of the encoder over time.

[0016] A classifier is cascaded with the semantic encoder of the encoder.

[0017] In some embodiments, the style memory information includes a style memory feature library corresponding to each source domain, which is used to store the style memory features of each sample of each source domain, wherein the style memory features of each sample of the source domain are obtained through a feature encoder and a style encoder in the memory encoder.

[0018] In some embodiments, the learning of domain-related style features includes: performing comparative learning of domain-related style features, so that the style features of samples in any source domain have high similarity with the style memory features in the style memory feature library of the same source domain, and have low similarity with the style memory features in the style memory feature library of different source domains.

[0019] In some embodiments, the performing contrastive learning of domain-related style features includes: constructing a contrastive learning loss function using a softmax function to determine the learning loss of the style features, and performing contrastive learning of domain-related style features using the contrastive learning loss function.

[0020] In some embodiments, the semantic memory information includes a semantic memory feature library, which is used to store the semantic memory features of the variants of each sample in each source domain, wherein the semantic memory features of the variants of each sample in the source domain are obtained through a feature encoder and a semantic encoder in a memory encoder, and the variant of each sample is similar to each sample.

[0021] In some embodiments, the variant of each sample is obtained by performing data augmentation on each sample, or the variant of each sample is selected from samples in any source domain that belong to the same category as each sample.

[0022] In some embodiments, the learning of domain-independent semantic features includes:

[0023] Calculate the first semantic similarity distribution between the semantic feature of each sample in the source domain and each semantic memory feature in the semantic memory feature library;

[0024] Calculate the second semantic similarity distribution between the semantic memory feature of the variant of each sample in the source domain and each semantic memory feature in the semantic memory feature library;

[0025] According to the first semantic similarity distribution and the second semantic similarity distribution, a cross entropy loss function is constructed to determine the learning loss of the semantic feature, so that the first semantic similarity distribution and the second semantic similarity distribution tend to be consistent.

[0026] In some embodiments, decoupling the style feature from the semantic feature comprises: decoupling the style feature from the semantic feature using orthogonal constraints.

[0027] In some embodiments, decoupling the style feature and the semantic feature by using the orthogonal constraint includes:

[0028] Construct a style feature matrix according to the style features of each sample in the source domain;

[0029] Construct a semantic feature matrix based on the semantic features of each sample in the source domain;

[0030] According to the transpose of one of the style feature matrix and the semantic feature matrix and the other one, an orthogonal constraint is constructed to determine the decoupling loss so that the style feature matrix and the semantic feature matrix tend to be orthogonal.

[0031] In some embodiments, predicting the category of the sample in the source domain through classification according to the semantic feature includes: inputting the semantic feature of the sample in the source domain into the classifier, and predicting and outputting the category of the sample.

[0032] In some embodiments, a cross entropy loss function is constructed based on the predicted categories and category labels of each sample in each source field to determine the classification loss.

[0033] In some embodiments, during the learning process, the parameters of the encoder are updated using a gradient descent method, or the parameters of the memory encoder at a current moment are jointly updated based on the parameters of the memory encoder at a previous moment and the parameters of the encoder at a previous moment.

[0034] In some embodiments, it also includes: after the domain generalization model is learned, the feature encoder and the semantic encoder of the encoder are used to extract the semantic features of the samples of the target domain, and the classifier is used to classify the samples according to the semantic features of the samples of the target domain to determine the corresponding categories of the samples of the target domain.

[0035] In some embodiments, the samples in the source domain include image samples used as training data, and the samples in the target domain include image samples in actual working scenarios.

[0036] Some embodiments of the present disclosure provide a domain generalization model product for domain generalization, including:

[0037] The encoder includes: a feature encoder for extracting sample features, a semantic encoder cascaded with the feature encoder for extracting semantic features, and a style encoder cascaded with the feature encoder for extracting style features;

[0038] The memory encoder is consistent with the structure of the encoder. During the learning process, the parameters of the memory encoder are updated according to the accumulation of the parameters of the encoder over time.

[0039] A classifier is cascaded with the semantic encoder of the encoder.

[0040] In some embodiments, the feature encoder in the encoder and the memory encoder includes ResNet, the semantic encoder or the style encoder in the encoder and the memory encoder includes a multilayer perceptron, and the classifier includes a multilayer perceptron.

[0041] Some embodiments of the present disclosure provide a domain generalization device, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the domain generalization method of each embodiment based on instructions stored in the memory.

[0042] Some embodiments of the present disclosure provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the domain generalization method of each embodiment. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. The present disclosure can be more clearly understood according to the following detailed description with reference to the drawings.

[0044] Obviously, the drawings described below are only some embodiments of the present disclosure, and a person skilled in the art can obtain other drawings based on these drawings without creative work.

[0045] Figure 1A schematic diagram showing the structure of a domain generalization model based on style memory and semantic memory mechanisms according to some embodiments of the present disclosure.

[0046] Figure 2 A flowchart illustrating a domain generalization method according to some embodiments of the present disclosure is shown.

[0047] Figure 3 It is a schematic diagram of the structure of the domain generalization device of some embodiments of the present disclosure. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure.

[0049] Unless otherwise specified, descriptions such as “first” and “second” in the present disclosure are used to distinguish different objects and are not used to indicate meanings such as size or time sequence.

[0050] As mentioned above, domain generalization technology is one of the basic topics in the field of computer vision. The problem it studies is to learn a model with strong generalization ability from several data sets with different data distributions (source domains) in order to achieve better prediction results on unknown test sets (target domains).

[0051] The following is a formal definition of domain generalization: given D source domain data Each source field D d Contains N d Samples where x d,i Indicates the source domain D d The i-th sample in y d,i ∈{1,2,…,n c} is x d,i The category labels (Labels), n c The goal of domain generalization is to learn a good model from multiple source domain data that can be generalized to unknown target domain data D. T superior.

[0052] Figure 1 A schematic diagram showing the structure of a domain generalization model based on style memory and semantic memory mechanisms according to some embodiments of the present disclosure.

[0053] like Figure 1As shown in the figure, the domain generalization model includes: encoder, memory encoder, and classifier. The encoder includes: feature encoder for extracting sample features, semantic encoder cascaded with feature encoder for extracting semantic features, and style encoder cascaded with feature encoder for extracting style features. In order to learn semantic features and style features at the instance level, a memory encoder is also introduced. The structure of the memory encoder is consistent with that of the encoder, but the memory encoder is a time-domain enhanced version of the encoder. During the learning process, the parameters of the memory encoder are updated according to the accumulation of the encoder parameters over time. That is, the memory encoder contains information at historical moments, not just determined by the information at the current moment, so it has "memory". The classifier is cascaded with the semantic encoder of the encoder.

[0054] The feature encoder in the encoder and memory encoder includes, for example, ResNet, or other networks that can extract sample features. The semantic encoder or style encoder in the encoder and memory encoder includes, for example, a multi-layer or single-layer MLP (multi-layer perceptron), a fully connected network. The classifier includes, for example, a multi-layer or single-layer MLP, or other networks that can implement classification functions.

[0055] Let’s define some symbols. For a sample x d,i , after feature encoder E f (x d,i ,θ e,f ) Extract features to get z d,i , and then passes through the semantic encoder E c (z d,i ,θ e,c ) to obtain the semantic feature c d,i , and at the same time pass through the style encoder E s (Z d,i ,θ e,s ) to get the style feature s d,i , where θ e,f ,θ e,c ,θ e,s The points represent feature encoder parameters, semantic encoder parameters, and style encoder parameters respectively. The final classifier Classify samples according to semantic features. Represents the classifier parameters. The above feature encoder, semantic encoder and style encoder form a complete encoder E e ={E f ,E c ,E s}, and its corresponding parameters are That is, the encoder parameters include feature encoder parameters, semantic encoder parameters, and style encoder parameters. m With encoder E e The structure of is exactly the same. In order to facilitate the distinction, the memory encoder E m ={E m,f ,E m,c ,E m,s}, the corresponding parameters are Among them, E m,f ,E m,c ,E m,s Respectively represent the feature encoder, semantic encoder and style encoder in the memory encoder, θ m,f ,θ m,c ,θ m,s They represent the feature encoder parameters, semantic encoder parameters, and style encoder parameters in the memory encoder respectively.

[0056] The learning of domain generalization models mainly includes three components: learning of style features that are invariant within a domain, learning of semantic features that are invariant between domains, and decoupling of semantic features from style features. Among them, learning of style features that are invariant within a domain: maintain a style memory feature library for each source domain, and then use contrastive learning to bring the style features of samples belonging to the same domain closer, while pushing away the style features of samples belonging to different domains. Learning of semantic features that are invariant between domains: maintain a semantic memory feature library, and learn the semantic information of samples through a new contrastive learning paradigm of the "jury system". Decoupling of semantic features from style features: orthogonal constraints can be used to achieve decoupling of semantic features from style features.

[0057] Figure 2 A flowchart of a domain generalization method according to some embodiments of the present disclosure is shown. The learning process of the domain generalization model is specifically described through the domain generalization method.

[0058] like Figure 2 As shown, the domain generalization method of this embodiment includes: steps 210 to 250.

[0059] In step 210, domain-related style features are learned using samples of the source domain according to the style memory information, or in other words, domain-invariant (or shared) style features are learned.

[0060] The style memory information includes a style memory feature library corresponding to each source domain, which is used to store the style memory features of each sample in each source domain, also known as a style feature enqueue. The style memory features of each sample in the source domain are obtained through the feature encoder and the style encoder in the memory encoder.

[0061] For the sample x d,i , extract its style memory features through the feature encoder and style encoder in the memory encoder Then store it in the source field D d Style memory feature library middle, The source domain D is stored in d Style memory features of different samples Where B is the length of the style memory feature library. D source domains correspond to D style memory feature libraries.

[0062] The learning of the domain-related style features includes: performing comparative learning of the domain-related style features, for example, constructing a comparative learning loss function using a softmax function to determine the learning loss of the style features, and performing comparative learning of the domain-related style features using the comparative learning loss function, so that the style features of samples in any source domain have a high similarity with the style memory features in the style memory feature library of the same source domain, and have a low similarity with the style memory features in the style memory feature library of different source domains.

[0063] In order to learn the style features shared by the domain, for the sample x d,i , extract its style feature s through the feature encoder and style encoder in the encoder d,i =E s (E f (x d,i ,θ e,f ),θ e,s ), and let s d,i The style memory features in the style memory feature library of the same source domain have high similarity, and the style memory features in the style memory feature library of different source domains have low similarity. A contrastive learning loss function is constructed to achieve this goal. The contrastive learning loss function is used to determine the learning loss of style features, also known as style contrastive loss.

[0064]

[0065] Among them, Z s =B·∑ d N d , B represents the length of the style memory feature library, ∑ d N d represents the total number of samples in each source field, Z s Represents the total number of style memory features of all style memory feature libraries corresponding to all source domains, is a normalized number, = is the temperature factor, is a hyperparameter, Represents x 1 and x 2 The cosine similarity of exp represents the exponential function with the natural constant e as the base. The contrastive learning loss function is constructed based on the softmax function, which converts the style feature of each source domain into and B×(D-1) source domain style feature pairs d′≠d. As learning or training progresses, the style feature s d,i The features in the style memory feature library of the field have high similarity, while the features in the style memory feature library of other fields have low similarity, so that the invariant (or shared) style information in the field can be learned.

[0066] In step 220, domain-independent semantic features are learned using samples from the source domain according to semantic memory information, or in other words, domain-invariant (or shared) semantic features are learned.

[0067] Semantic features are factors that determine sample categories. They are domain-independent, invariant, and shared across domains.

[0068] Define the variants of the sample. The variant of each sample is obtained by performing data augmentation on each sample, or the variant of each sample is selected from samples in any source domain that belong to the same category as each sample. Definition For x d,i A variant of is from any source field but x d,i randomly selected from samples belonging to the same category, or is x d,i After data enhancement, the variant of each sample is similar to the sample. d,i and The semantic features of the samples have high similarity, so that the real semantic information of this type of samples can be learned. To this end, a new contrastive learning paradigm ("jury system") is proposed to achieve this goal.

[0069] First, a semantic memory feature library is constructed as semantic memory information. The semantic memory feature library is used to store the semantic memory features of the variants of each sample in each source domain, also known as a semantic feature enqueue. When the length of the semantic memory feature library is fixed, a first-in-first-out storage method for semantic memory features can be used. The semantic memory features of the variants of each sample in the source domain are obtained through the feature encoder and the semantic encoder in the memory encoder. For example, for Semantic memory features The semantic memory features of each sample variant are successively Store it in the semantic memory feature library, and remove the semantic memory features of the historical moment from the semantic memory feature library, and finally obtain a semantic memory feature library with a library length of B

[0070] Next, the semantic features of each sample in the source domain are calculated. For example, for x d,i , the semantic feature c is obtained through the feature encoder and semantic encoder in the encoder d,i =E c (E f (x d,i ,θ e,f ),θ e,c ).

[0071] Next, we learn domain-independent semantic features and regard the semantic memory features of each sample in the semantic memory feature library as a “jury” to evaluate c d,i and The similarity specifically includes: calculating the first semantic similarity distribution between the semantic features of each sample in the source domain and each semantic memory feature in the semantic memory feature library; calculating the second semantic similarity distribution between the semantic memory features of the variants of each sample in the source domain and each semantic memory feature in the semantic memory feature library; constructing a cross entropy loss function according to the first semantic similarity distribution and the second semantic similarity distribution to determine the learning loss of the semantic feature, so that the first semantic similarity distribution and the second semantic similarity distribution tend to be consistent. The following is described in conjunction with the formula.

[0072] x d,i The first semantic similarity distribution with all sample features in the semantic memory feature library in The definition is as follows:

[0073]

[0074] The second semantic similarity distribution with all sample features in the semantic memory feature library in The definition is as follows:

[0075]

[0076] If c d,i and If the real distribution only contains the semantic information that is invariant between domains, then their semantic similarity with the samples in the semantic memory feature library is close, that is, the two distributions p(x d,i θ e,c ,θe,f ,V c )and To this end, the following loss function is proposed to penalize the cross entropy between the two distributions, namely the cross entropy loss function, which is used to determine the learning loss of semantic features.

[0077]

[0078] Where Z c =∑ d N d is a normalized number, indicating the total number of samples in each source domain. It should be noted that this embodiment does not directly constrain c d,i and similarity, such as their inner product Instead, align c d,i and The similarity distribution of the features in the semantic memory feature library is similar to that of the traditional contrastive learning, which directly brings c d,i and There is a big difference.

[0079] In step 230 , the style features and the semantic features are decoupled, for example, using orthogonal constraints.

[0080] Only learning the above-mentioned semantic features that are invariant between domains cannot benefit from learning the style features that are invariant within a domain, because the semantic features and the style features are not connected. Therefore, this embodiment proposes to make the semantic features and the style features orthogonal in the feature space, so that the semantic features do not contain style information, and the style features do not contain semantic information, so as to achieve the decoupling of semantics and style.

[0081] Using orthogonal constraints, decoupling the style features and semantic features includes: constructing a style feature matrix H according to the style features of each sample in the source domain s ; Construct the semantic feature matrix H according to the semantic features of each sample in the source domain c ; According to the transposition of one of the style feature matrix and the semantic feature matrix and the other matrix, an orthogonal constraint, also known as orthogonal loss, is constructed to determine the decoupling loss so that the style feature matrix and the semantic feature matrix tend to be orthogonal. As an example, the orthogonal constraint is as follows:

[0082]

[0083] Among them, H c and H s Each row is a semantic feature c d,i and style features d,i The matrix composed of is the squared Frobenius norm.

[0084] In step 240 , the category of the sample in the source domain is predicted by classification according to the semantic features.

[0085] For example, the semantic features of the samples in the source domain are input into the classifier, and the category of the samples is predicted and output.

[0086] Among them, according to the predicted categories and category labels of each sample in each source field, a cross entropy loss function is constructed to determine the classification loss. The formula is as follows:

[0087]

[0088] Among them, the classifier According to the sample x d,i The semantic features of d,i To predict sample x d,i The category of M = ∑ d N d Represents the total number of samples in each source field, y d,i is the sample x d,i The corresponding category labels, Represents classifier parameters.

[0089] In step 250, the parameters of the domain generalization model are updated according to the learning loss of the style features, the learning loss of the semantic features, the decoupling loss, and the classification loss. The domain generalization model after learning is able to extract domain-independent semantic features of samples in the target domain.

[0090] That is, all loss functions are added together to train the domain generalization model: L = L cls +L s +L c +L o , update the parameters of the domain generalization model, the parameters of the domain generalization model include etc. are the parameters of each encoder and classifier.

[0091] During the learning process, the encoder parameters Use the gradient descent method to update (momentum update) and memorize the encoder parameters The encoder parameters are dynamically updated in the following way: the parameters of the memory encoder at the current moment are jointly updated according to the parameters of the memory encoder at the previous moment and the parameters of the encoder at the previous moment. The formula is expressed as: Where α∈[0,1) is the momentum parameter.

[0092] After the domain generalization model is learned, in the testing phase or model application phase, the feature encoder E f and semantic encoder E c Extract the semantic features of the samples in the target domain, and use the classifier C to classify the samples according to the semantic features of the samples in the target domain to determine the corresponding categories of the samples in the target domain. After the test is passed, the model can be put into practical use.

[0093] The above embodiment uses style memory to learn domain-related style features and uses semantic memory to learn domain-independent semantic features, deeply explores style information in different fields and semantic information of different categories, and then decouples style features and semantic features, so that the model is capable of extracting domain-independent semantic features of samples, thereby improving the model's generalization ability for any unknown target domain.

[0094] The domain generalization technology proposed in the present disclosure can be applied to various domain generalization scenarios of computer vision, for example, to image content understanding related products in various actual operation scenarios (real business scenarios). At this time, the samples in the source domain include image samples used as training data, and the samples in the target domain include image samples in actual operation scenarios. After generalizing the model based on the samples in the source domain, the model can achieve better prediction results in the unknown target domain, thereby making up for the difference between the training data (which can be considered as the source domain) and the test data in the actual operation scenario (which can be considered as the target domain), and enhancing the generalization ability of the model in the actual operation scenario. For example, in the violent sorting products in the logistics scenario, the use of this visual domain generalization technology can make up for the difference in violent sorting data collected by different sorting centers, and enhance the sorting performance of the online violent sorting recognition model in different sorting centers and different video acquisition environments.

[0095] Figure 3 FIG. 1 is a schematic diagram of the structure of a domain generalization device in some embodiments of the present disclosure. Figure 3 As shown, the domain generalization device 300 of this embodiment includes: a memory 310 and a processor 320 coupled to the memory 310 , and the processor 320 is configured to execute the domain generalization method in any of the aforementioned embodiments based on instructions stored in the memory 310 .

[0096] For example, using samples from the source domain, domain-related style features are learned based on style memory information; using samples from the source domain, domain-independent semantic features are learned based on semantic memory information; style features and semantic features are decoupled; based on semantic features, the categories of samples from the source domain are predicted through classification; based on the learning loss of style features, the learning loss of semantic features, the decoupling loss, and the classification loss, the parameters of the domain generalization model are updated. After learning, the domain generalization model can extract domain-independent semantic features of samples from the target domain.

[0097] The memory 310 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, an application program, a boot loader, and other programs.

[0098] The domain generalization device 300 may also include an input / output interface 330, a network interface 340, a storage interface 350, etc. These interfaces 330, 340, 350 and the memory 310 and the processor 320 may be connected, for example, via a bus 360. The input / output interface 330 provides a connection interface for input / output devices such as a display, a mouse, a keyboard, and a touch screen. The network interface 340 provides a connection interface for various networked devices. The storage interface 350 provides a connection interface for external storage devices such as SD cards and USB flash drives.

[0099] Some embodiments of the present disclosure provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the domain generalization method in any of the aforementioned embodiments.

[0100] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more non-transient computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer program code.

[0101] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0102] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0104] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A domain generalization method, characterized in that, it includes: Using the samples of the source domain, according to the style memory information, through the learning of style features shared within the domain, learning the domain-related style features; Using the samples of the source domain, according to the semantic memory information, through the learning of semantic features shared across domains, learning the domain-agnostic semantic features; Decoupling the style features and semantic features by making the semantic features and style features orthogonal in the feature space; Predicting the category of the samples in the source domain according to the semantic features; Updating the parameters of the domain generalization model according to the learning loss of the style features, the learning loss of the semantic features, the decoupling loss, and the classification loss. After learning, the domain generalization model can extract the domain-agnostic semantic features of the samples in the target domain.

2. The method according to claim 1, characterized in that, the domain generalization model includes: An encoder, including: a feature encoder for extracting sample features, a semantic encoder cascaded with the feature encoder for extracting semantic features, and a style encoder cascaded with the feature encoder for extracting style features; A memory encoder, having the same structure as the encoder. During the learning process, the parameters of the memory encoder are updated according to the accumulation of the parameters of the encoder over time; A classifier, cascaded with the semantic encoder of the encoder.

3. The method according to claim 2, characterized in that, The style memory information includes a style memory feature library corresponding to each source domain, which is used to store the style memory features of each sample in each source domain, wherein, the style memory feature of each sample in the source domain is obtained through the feature encoder and the style encoder in the memory encoder.

4. The method according to claim 3, characterized in that, The learning of the domain-related style features includes: Performing contrastive learning of the domain-related style features, so that the style features of any sample in the source domain have a high similarity with the style memory features in the style memory feature library of the same source domain and a low similarity with the style memory features in the style memory feature libraries of different source domains.

5. The method according to claim 4, characterized in that, The contrastive learning of the domain-related style features includes: Using the softmax function to construct a contrastive learning loss function for determining the learning loss of the style features, and using the contrastive learning loss function to perform contrastive learning of the domain-related style features.

6. The method according to claim 2, characterized in that, The semantic memory information includes a semantic memory feature library, which is used to store the semantic memory features of the variants of each sample in each source domain, wherein, the semantic memory feature of the variant of each sample in the source domain is obtained through the feature encoder and the semantic encoder in the memory encoder, and the variant of each sample is similar to each sample.

7. The method according to claim 6, characterized in that, The variant of each sample is obtained by performing data augmentation on each sample, or the variant of each sample is selected from the samples in any source domain that belong to the same category as each sample.

8. The method according to claim 6, wherein, the learning of domain - independent semantic features includes: calculating a first semantic similarity distribution between the semantic features of each sample in the source domain and each semantic memory feature in the semantic memory feature library; calculating a second semantic similarity distribution between the semantic memory features of each variant of the samples in the source domain and each semantic memory feature in the semantic memory feature library; constructing a cross - entropy loss function according to the first semantic similarity distribution and the second semantic similarity distribution to determine the learning loss of the semantic features, so that the first semantic similarity distribution and the second semantic similarity distribution tend to be consistent.

9. The method according to claim 2, wherein, the decoupling of the style features and the semantic features includes: using orthogonal constraints to decouple the style features and the semantic features.

10. The method according to claim 9, wherein, the using of orthogonal constraints to decouple the style features and the semantic features includes: forming a style feature matrix according to the style features of each sample in the source domain; forming a semantic feature matrix according to the semantic features of each sample in the source domain; constructing an orthogonal constraint according to one of the style feature matrix and the semantic feature matrix and the transpose of the other matrix to determine the decoupling loss, so that the style feature matrix and the semantic feature matrix tend to be orthogonal.

11. The method according to claim 2, wherein, classifying and predicting the category of the samples in the source domain according to the semantic features includes: inputting the semantic features of the samples in the source domain into the classifier and predicting and outputting the category of the samples.

12. The method according to claim 11, wherein, constructing a cross - entropy loss function according to the predicted categories and the category labels of each sample in each source domain to determine the classification loss.

13. The method according to claim 2, wherein, during the learning process, the parameters of the encoder are updated using the gradient descent method, or the parameters of the memory encoder at the current moment are jointly updated according to the parameters of the memory encoder at the previous moment and the parameters of the encoder at the previous moment.

14. The method according to claim 2, wherein, further comprising: after the learning of the domain generalization model is completed, using the feature encoder and the semantic encoder of the encoder to extract the semantic features of the samples in the target domain, and using the classifier to classify according to the semantic features of the samples in the target domain to determine the corresponding categories of the samples in the target domain.

15. The method according to any one of claims 1 - 14, wherein, the samples in the source domain include image samples used as training data, the samples in the target domain include image samples in the actual operation scenario.

16. The method according to claim 2, wherein, the feature encoder in the encoder and the memory encoder includes ResNet, the semantic encoder or the style encoder in the encoder and the memory encoder includes a multi - layer perceptron, the classifier includes a multi - layer perceptron.

17. A domain generalization model product for domain generalization, which is learned using the domain generalization method according to any one of claims 1 - 16, Comprising: An encoder, comprising: a feature encoder for extracting sample features, a semantic encoder for extracting semantic features cascaded with the feature encoder, and a style encoder for extracting style features cascaded with the feature encoder; A memory encoder, having the same structure as the encoder, and during the learning process, the parameters of the memory encoder are updated according to the accumulation of the parameters of the encoder over time; A classifier, cascaded with the semantic encoder of the encoder.

18. The model product according to claim 17, wherein, the feature encoder in the encoder and the memory encoder includes ResNet, the semantic encoder or the style encoder in the encoder and the memory encoder includes a multi-layer perceptron, and the classifier includes a multi-layer perceptron.

19. A domain generalization device, comprising: A memory; and A processor coupled to the memory, the processor being configured to execute the domain generalization method according to any one of claims 1-16 based on instructions stored in the memory.

20. A non-transitory computer-readable storage medium, having stored thereon a computer program, which when executed by a processor implements the steps of the domain generalization method according to any one of claims 1-16.

Citation Information

Patent Citations

  • A hand-painted clothing commodity image retrieval method based on a dual-path deep semantic network

    CN109670066A

  • Zero-sample image classification method of adversarial network based on meta-learning

    CN112364894A